H²VU-Benchmark
收藏资源简介:
H²VU-Benchmark是一个全面评估视频理解能力的基准,包含离线通用视频和在线流视频两大类。数据集涵盖了从几秒钟到1.5小时的视频,以桥接当前基准中的时间差距。评估任务不仅包括传统的感知和推理任务,还引入了反常识理解和轨迹状态跟踪模块,以测试模型在视频内容方面的深度理解能力。数据集的构建经过精心设计,包括静态场景过滤、对话内容识别和先验知识依赖性净化等步骤,以保持数据集质量和评估的有效性。
H²VU-Benchmark is a comprehensive benchmark for evaluating video understanding capabilities. It consists of two main categories: offline general-purpose videos and online streaming videos. The dataset covers videos ranging from a few seconds to 1.5 hours, aiming to bridge the temporal gaps existing in current benchmarks. Its evaluation tasks not only include traditional perception and reasoning tasks, but also introduce counter-intuitive understanding and trajectory state tracking modules to test the deep comprehension capabilities of models regarding video content. The construction of the dataset is meticulously designed, including steps such as static scene filtering, dialogue content recognition and prior knowledge dependence purification, to maintain the dataset quality and the effectiveness of the evaluation.




