FineVidBench
收藏资源简介:
FineVidBench是一个由华中科技大学提出的细粒度视频理解评估框架,旨在严格评估视频大型语言模型在场景和片段层面的性能。该数据集包含910个视频和22718个问题-答案对,视频来源于多个公共数据集,如Something-Something V2和Moments in Time等。数据集通过自动化流程和人工审核相结合的方式生成,涵盖了多种动作类别,包括易于识别的独特动作、无明确特征的灵活动作以及难以用肉眼检测的轻微动作。FineVidBench能够全面评估视频大型语言模型捕捉和解释时间细节的能力。
FineVidBench is a fine-grained video understanding evaluation framework proposed by Huazhong University of Science and Technology, which aims to rigorously evaluate the performance of video large language models at the scene and clip levels. This dataset contains 910 videos and 22718 question-answer pairs, with videos sourced from multiple public datasets such as Something-Something V2 and Moments in Time. The dataset is generated through a combination of automated workflows and manual review, covering a diverse range of action categories, including easily recognizable distinct actions, flexible actions without clear features, and subtle actions that are difficult to detect with the naked eye. FineVidBench can comprehensively evaluate the ability of video large language models to capture and interpret temporal details.




