MTCMB
收藏资源简介:
MTCMB是一个多任务基准框架,用于评估大型语言模型(LLM)在中医药知识、推理和安全方面的能力。它包含12个子数据集,涵盖了五个主要类别:知识问答、语言理解、诊断推理、处方生成和安全评估。该基准集整合了真实世界案例记录、国家执业医师资格考试和经典文本,为中医药能力模型提供了一个真实和全面的测试平台。初步结果表明,当前的大型语言模型在基础知识方面表现良好,但在临床推理、处方规划和安全合规方面仍存在不足。这些发现突出了迫切需要像MTCMB这样的领域对齐基准来指导更胜任和可靠的医疗人工智能系统的开发。
MTCMB is a multi-task benchmark framework designed to evaluate the capabilities of Large Language Models (LLMs) in traditional Chinese medicine (TCM) knowledge, reasoning, and safety. It comprises 12 sub-datasets covering five core categories: knowledge Q&A, language understanding, diagnostic reasoning, prescription generation, and safety assessment. This benchmark integrates real-world case records, the National Licensed Physician Qualification Examination, and classic texts, providing a realistic and comprehensive testbed for TCM-capable language models. Preliminary evaluation results show that current LLMs perform well in basic knowledge, but still have deficiencies in clinical reasoning, prescription planning, and safety compliance. These findings highlight the urgent need for domain-aligned benchmarks like MTCMB to guide the development of more competent and reliable medical artificial intelligence systems.




