KorMedMCQA-V
收藏资源简介:
KorMedMCQA-V是一个韩国医学执照考试风格的多模态多项选择问答基准数据集,用于评估视觉-语言模型(VLMs)。该数据集包含1,534个问题和2,043张相关图片,来自2012年至2023年的韩国医学执照考试,其中约30%的问题包含多张图片,需要跨图片证据整合。图片涵盖多种临床模态,包括X光、计算机断层扫描(CT)、心电图(ECG)、超声波、内窥镜和其他医学视觉内容。
KorMedMCQA-V is a multimodal multiple-choice question answering (MCQA) benchmark dataset modeled after the Korean medical licensing examination, designed for evaluating vision-language models (VLMs). This dataset comprises 1,534 questions and 2,043 associated images sourced from Korean medical licensing examinations administered between 2012 and 2023. Approximately 30% of the questions involve multiple images, necessitating the integration of evidence across different images. The images cover a wide range of clinical modalities, including X-rays, computed tomography (CT), electrocardiograms (ECG), ultrasound scans, endoscopy imagery, and other types of medical visual content.
KorMedMCQA-V 数据集概述
数据集基本信息
- 数据集名称:KorMedMCQA-V
- 数据集类型:多模态基准测试数据集
- 核心用途:用于评估视觉-语言模型在韩国医师资格考试风格的多模态多项选择题上的表现
- 数据来源:韩国医师资格考试(2012-2023年)
- 公开状态:已公开发布
数据构成
- 问题数量:1,534 个问题
- 关联图像数量:2,043 张图像
- 多图像问题比例:约 30%(需要跨图像证据整合)
- 图像覆盖的临床模态:X光、计算机断层扫描(CT)、心电图(ECG)、超声、内窥镜及其他医学视觉图像
评估基准与结果
- 评估模型数量:超过 50 个视觉-语言模型
- 模型类别:涵盖专有模型和开源模型,包括通用模型、医学专用模型和韩语专用模型
- 评估协议:统一的零样本评估协议
- 关键性能指标:
- 最佳专有模型(Gemini-3.0-Pro)准确率:96.9%
- 最佳开源模型(Qwen3-VL-32B-Thinking)准确率:83.7%
- 最佳韩语专用模型(VARCO-VISION-2.0-14B)准确率:43.2%
- 主要发现:
- 面向推理的模型变体比指令调优的对应模型有高达 +20 个百分点的提升。
- 医学领域专业化相对于强大的通用基线模型带来的增益不一致。
- 所有模型在多图像问题上的表现均下降。
- 不同成像模态间的性能存在显著差异。
数据集定位与关联
- 补充基准:是对纯文本基准 KorMedMCQA 的补充。
- 统一评估套件:与 KorMedMCQA 共同构成了一个用于评估韩语医学推理(涵盖纯文本和多模态条件)的统一评估套件。
获取与使用
- 数据集地址:https://huggingface.co/datasets/seongsubae/KorMedMCQA-V
- 代码仓库:https://github.com/baeseongsu/kormedmcqa_v
- 论文地址:https://arxiv.org/abs/2602.13650
- 排行榜地址:https://kormedmcqa-v.github.io/
- 许可证:数据集采用 CC BY-NC-SA 4.0 许可证(https://creativecommons.org/licenses/by-nc-sa/4.0/)
引用信息
如果研究中使用此基准,请引用论文:
@article{choi2026kormedmcqav, title={KorMedMCQA-V: A Multimodal Benchmark for Evaluating Vision-Language Models on the Korean Medical Licensing Examination}, author={Choi, Byungjin and Bae, Seongsu and Kweon, Sunjun and Choi, Edward}, journal={arXiv preprint arXiv:2602.13650}, year={2026} }



