MMAD
收藏资源简介:
MMAD数据集是一个全面的多模态大语言模型在工业异常检测领域的基准测试数据集,包含问题、图像和描述文本。所有问题都以多选题形式呈现,并经过人工验证。图像来自多个来源,保留了地面真值的掩码格式,以方便未来对多模态大语言模型的分割性能进行评估。描述文本大部分质量良好,但未经过人工验证,使用时需谨慎。MMAD旨在评估当前多模态大语言模型在工业质量检测中的表现,并识别在工业异常检测中的关键挑战。
The MMAD dataset is a comprehensive benchmark dataset for multimodal large language models in the field of industrial anomaly detection. It contains questions, images and descriptive texts. All questions are presented in multiple-choice format and have been manually verified. The images are sourced from multiple origins, with their ground-truth mask formats retained to facilitate future evaluation of the segmentation performance of multimodal large language models. Most of the descriptive texts are of good quality, but they have not been manually verified, so caution should be exercised when using them. The MMAD dataset aims to evaluate the performance of current multimodal large language models in industrial quality inspection and identify key challenges in industrial anomaly detection.
MMAD: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
数据集概述
- 任务类别: 问答
- 标签:
- 异常检测
- 多模态大语言模型 (MLLM)
- 规模: 10K<n<100K
- 许可证: MIT
数据集内容
- 内容: 包含问题、图像和描述文本。
- 问题: 所有问题均为多选题格式,并经过人工验证,包括选项和答案。
- 图像: 图像来源包括以下数据集:
- DS-MVTec
- MVTec-AD
- MVTec-LOCO
- VisA
- GoodsAD 图像保留了地面真值的掩码格式,以方便未来对多模态大语言模型分割性能的评估。
- 描述文本: 大多数图像都有对应的文本文件,位于同一文件夹中,包含相关描述。由于这不是该基准的主要关注点,因此未进行人工验证。尽管大多数描述质量良好,但请谨慎使用。
数据集目标
- 评估当前多模态大语言模型在工业质量检测中的表现。
- 确定在工业异常检测中表现最佳的多模态大语言模型。
- 识别多模态大语言模型在工业异常检测中的关键挑战。
评估方法
- 请参考GitHub仓库中的evaluation/examples文件夹。
引用
bibtex @inproceedings{Jiang2024MMADTF, title={MMAD: The First-Ever Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection}, author={Xi Jiang and Jian Li and Hanqiu Deng and Yong Liu and Bin-Bin Gao and Yifeng Zhou and Jialin Li and Chengjie Wang and Feng Zheng}, year={2024}, journal={arXiv preprint arXiv:2410.09453}, }




