MedSPOT
收藏资源简介:
MedSPOT是由印度国立技术学院等机构联合构建的临床GUI工作流感知序列标注基准,涵盖10种医疗影像软件平台的216个任务驱动视频与597个标注关键帧。该数据集通过分层标注捕捉动态界面状态下的空间精度与上下文依赖关系,每个任务包含2-3个互相关联的临床工作流步骤。数据采集过程模拟真实医疗操作场景,采用严格的多级标注协议确保决策帧的因果一致性。该基准旨在评估多模态大模型在安全关键医疗环境中的序列推理能力,解决传统单步标注方法无法反映临床工作流错误传播的核心问题。
MedSPOT is a clinical GUI workflow-aware sequence labeling benchmark jointly developed by the National Institute of Technology India and other collaborating institutions. It encompasses 216 task-driven videos and 597 annotated key frames spanning 10 medical imaging software platforms. Through hierarchical annotation, this dataset captures spatial accuracy and contextual dependencies within dynamic interface states, with each task consisting of 2 to 3 mutually correlated clinical workflow steps. The data collection process simulates real-world medical operational scenarios and adopts a strict multi-level annotation protocol to ensure the causal consistency of decision-making frames. This benchmark aims to evaluate the sequential reasoning capabilities of multimodal large language models (LLMs) in safety-critical medical environments, addressing the core limitation of traditional single-step annotation methods that fail to reflect error propagation in clinical workflows.
MedSPOT 数据集概述
数据集基本信息
- 数据集名称:MedSPOT
- 核心用途:评估多模态大语言模型在医学影像软件图形用户界面上的定位与交互能力。
- 发布状态:已发布。
- 相关论文:MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI
- 数据集访问地址:https://huggingface.co/datasets/Tajamul21/MedSPOT
- 项目代码仓库:https://github.com/Tajamul21/MedSPOT
- 官方网站:https://rozainmalik.github.io/MedSPOT_web/
数据集内容与特点
- 评估对象:涵盖10款医学影像应用程序的GUI,包括3DSlicer、DICOMscope、Weasis、MITK等。
- 任务性质:工作流程感知的顺序性基础任务。
- 评估协议:采用顺序评估,若模型在某一步失败,则任务提前终止,以模拟真实GUI交互中错误累积的情况。
评估指标
| 指标名称 | 全称 | 描述 |
|---|---|---|
| TCA | 任务完成准确率 | 所有步骤均按顺序正确完成的任务比例。 |
| SHR | 步骤命中率 | 所有被评估步骤中,每一步的准确率。 |
| S1A | 第一步准确率 | 每个任务中第一步的准确率。 |
数据集结构
MedSPOT-Bench/ Annotations/ 3DSlicer_Annotation.json DICOMscope_Annotation.json Weasis_Annotation.json ... Images/ 3DSlicer/ DICOMscope/ Weasis/ ...
标注格式
标注文件为JSON格式,每个文件包含一个tasks列表。每个任务包含task_overview和steps。每个步骤包含:
step_id:步骤序号。image_path:对应图像路径。instruction:操作指令。actions:一个动作列表,每个动作包含type(如“click”)、target(目标描述)和bbox(边界框坐标)。
评估与结果
- 评估脚本:提供针对多个模型的独立评估脚本,包括GUI-Actor、GPT-5、GPT-4o-mini、CogAgent-9B、Qwen2-VL、Gemma3-27B、Llama-3.2-11B等。
- 结果保存路径:
results/ ModelName/ SoftwareName/ task_results.json task_metrics.json failure_statistics.json overall_dataset_metrics.json
使用依赖
- 通用依赖:
torch>=2.0,transformers>=4.40,pillow,tqdm。 - 模型加载:部分模型从Hugging Face加载,需提前登录并获取访问权限。
- API模型:评估GPT-5、GPT-4o-mini等模型需预先设置
OPENAI_API_KEY环境变量。
参考文献
若在研究中使用本数据集,请引用相关论文。

- 1MedSPOT: A Workflow-Aware Sequential Grounding Benchmark for Clinical GUI印度国立技术学院·Srinagar分校·Gaash研究实验室; e&集团; 穆罕默德·本·扎耶德人工智能大学; 阿卜杜拉国王科技大学 · 2026年



