IllusionMNIST, IllusionFashionMNIST, IllusionAnimals, IllusionChar
收藏资源简介:
本研究引入了四个专门用于评估多模态模型在视觉错觉识别和解释能力的数据集:IllusionMNIST、IllusionFashionMNIST、IllusionAnimals和IllusionChar。这些数据集包含训练集和测试集,旨在全面评估模型的性能。数据集通过结合LLM生成的描述和ControlNet模型生成,确保了数据集的多样性和质量。数据集的创建过程包括生成场景描述、合成图像以及通过人工审核确保数据集的可靠性。这些数据集主要应用于视觉问答任务,旨在提高多模态模型对视觉错觉的理解和解释能力,从而增强模型的鲁棒性和人类类似的视觉理解能力。
This study introduces four datasets specifically tailored to evaluate the visual illusion recognition and interpretation capabilities of multimodal models: IllusionMNIST, IllusionFashionMNIST, IllusionAnimals, and IllusionChar. Each dataset comprises training and test subsets, enabling comprehensive performance assessment of models. These datasets are generated by combining descriptions produced by Large Language Models (LLMs) and the ControlNet model, which ensures their diversity and quality. The dataset creation process includes generating scene descriptions, synthesizing images, and conducting manual reviews to guarantee the reliability of the datasets. Primarily applied to visual question answering (VQA) tasks, these datasets aim to enhance multimodal models' understanding and interpretation of visual illusions, thereby improving the models' robustness and human-like visual comprehension abilities.




