MisCaption This!
收藏资源简介:
‘MisCaption This!’是一个由大型视觉语言模型(LVLM)生成的误标注图像训练数据集。该数据集通过操纵真实图像标题对,使用LVLM生成虚假标题,创建出具有误导性的图像标题对。数据集的构建目的是为了提高检测模型对现实世界误信息的泛化能力,同时探索不同的训练策略和集成方法对模型性能的影响。该数据集可应用于多模态误信息检测领域,旨在解决图像与文本之间的不实信息检测问题。
'MisCaption This!' is a mislabeled image training dataset generated by Large Vision-Language Models (LVLMs). It is constructed by manipulating real image-caption pairs and utilizing LVLMs to generate deceptive captions, thereby creating misleading image-caption pairs. The dataset is developed to improve the generalization ability of detection models against real-world misinformation, and to investigate the impacts of different training strategies and ensemble methods on model performance. This dataset can be applied in the domain of multimodal misinformation detection, with the goal of addressing the task of detecting false information between images and their associated captions.

- 1Latent Multimodal Reconstruction for Misinformation Detection希腊信息技术研究所, 雅典 · 2025年



