HAVEN
收藏资源简介:
HAVEN数据集是由中国科学院大学等机构创建的,用于评估大型多模态模型在视频理解任务中产生虚构信息的问题。该数据集基于三个维度构建,包括虚构信息产生的原因、虚构信息的方面和问题格式,共包含6497个问题。数据来源于公共视频数据集和手动收集的YouTube视频。该数据集旨在解决视频理解中的虚构信息问题,为大型多模态模型的评估提供了基准。
The HAVEN dataset was developed by institutions including the University of Chinese Academy of Sciences to evaluate hallucination issues of large multimodal models in video understanding tasks. Constructed based on three dimensions, namely the causes of hallucinatory information generation, the aspects of hallucinatory information, and question formats, the dataset contains a total of 6497 questions. Its data is sourced from public video datasets and manually collected YouTube videos. This dataset aims to address the hallucination problem in video understanding and provides a benchmark for the evaluation of large multimodal models.

- 1Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation中国科学院大学 · 2025年



