GMAI-MMBench
收藏资源简介:
GMAI-MMBench是由上海人工智能实验室等机构创建的一个综合性的多模态评估基准,专门用于测试大型视觉语言模型在真实临床场景中的能力。该数据集包含了来自全球的285个多样化的临床相关数据集,覆盖39种模态,涉及18个临床VQA任务和18个临床部门。数据集的创建过程包括从公共和医院来源收集数据,标准化图像和标签,以及构建一个词法树结构以方便用户定制评估任务。GMAI-MMBench旨在解决医学领域中大型视觉语言模型的评估问题,特别是在诊断和治疗方面的应用。
GMAI-MMBench is a comprehensive multimodal evaluation benchmark developed by institutions including the Shanghai AI Laboratory, which is specifically designed to test the capabilities of large vision-language models in real-world clinical scenarios. This benchmark encompasses 285 diverse clinical-related datasets from across the globe, covering 39 modalities, and involving 18 clinical VQA tasks and 18 clinical departments. The development process of this benchmark includes collecting data from public and hospital sources, standardizing images and labels, and constructing a lexical tree structure to facilitate users in customizing evaluation tasks. GMAI-MMBench aims to address the evaluation challenges of large vision-language models in the medical field, particularly their applications in diagnosis and treatment.

- 1GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI上海人工智能实验室 · 2024年



