MMAT-1M
收藏资源简介:
MMAT-1M是一个百万规模的多模态代理调优数据集,旨在释放多模态大型语言模型在思维链推理、反思和动态工具利用方面的全部潜力。该数据集通过一个新颖的四阶段数据引擎构建:首先,策划带有问答对的公开多模态数据集;其次,利用GPT-4o为这些问答对生成理由,并通过多轮范式动态整合API调用和检索增强生成(RAG)信息;第三,通过反思精炼理由,确保逻辑一致性和准确性,形成带有理由和反思(RR)的多轮对话数据集;最后,可选择将多轮对话压缩为单轮理由和反思格式(ORR)以提高效率。
MMAT-1M is a multimodal agent tuning dataset of a million-scale, designed to unleash the full potential of multimodal large language models in chain-of-thought reasoning, reflection, and dynamic tool utilization. The dataset is constructed through a novel four-stage data engine: initially, planning public multimodal datasets with question-answer pairs; secondly, employing GPT-4o to generate rationales for these question-answer pairs and dynamically integrating API calls and retrieval-augmented generation (RAG) information through multi-round paradigms; thirdly, refining rationales through reflection to ensure logical consistency and accuracy, forming a multi-round dialogue dataset with rationales and reflection (RR); finally, optionally compressing multi-round dialogues into a single-round rationale and reflection format (ORR) to enhance efficiency.
MMAT-1M数据集概述
基本信息
- 名称: MMAT-1M
- 类型: 百万规模多模态代理调优数据集
- 设计目的: 释放多模态大语言模型在思维链推理、反思和动态工具利用方面的潜力
- 发布状态: 已发布(2025-07-17)
- 论文状态: 被ICCV 2025接收(2025-07-24)
- 论文标题: "A Large Reasoning Dataset for Multimodal Agent Tuning"
- arXiv版本: 已更新(2025-07-30)
数据集特点
- 规模: 百万级
- 数据构建方法: 四阶段数据引擎
- 从公开多模态数据集中筛选问答对
- 使用GPT-4o生成推理依据并动态整合API调用和RAG信息
- 通过反思优化推理依据确保逻辑一致性
- 可选将多轮对话压缩为单轮格式(ORR)
- 输出格式:
- 多轮对话数据集(含Rationale和Reflection,RR)
- 单轮压缩格式(ORR)
许可证信息
- 复合许可证: 整合多个来源数据,需遵守各原始许可证
- Visual CoT: Apache 2.0(需署名)
- LLaVA-CoT: Apache 2.0(需署名)
- The Cauldron: 子集特定许可证 + CC-BY-4.0(商业使用需额外授权)
- TabMWP: CC BY-NC-SA 4.0(仅限非商业用途)
- Infoseek: Apache 2.0(需署名)
使用条款
- 责任限制:
- 不保证单个数据样本的法律状态
- 不对超出原始许可证条款的滥用行为负责
- 违规报告: 提供联系方式供举报许可证违规行为
相关资源
- 主页: https://MMAT-1M.github.io/
- Hugging Face数据集: https://huggingface.co/datasets/VIS-MPU-Agent/MMAT-1M
- arXiv论文: https://arxiv.org/abs/2507.21924
引用格式
bibtex @inproceedings{Gao2025MMAT1M, title={MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning}, author={Tianhong Gao and Yannian Fu and Weiqun Wu and Haixiao Yue and Shanshan Liu and Gang Zhang}, booktitle={Proceedings of ICCV}, year={2025}, }




