AutoEval-Video
收藏资源简介:
AutoEval-Video是由上海交通大学电子信息与电气工程学院清远研究院和字节跳动AI实验室共同创建的一个创新且具有挑战性的基准数据集,旨在全面评估大型视觉语言模型在开放式视频问答中的表现。该数据集包含327个复杂的开放式视频问答实例,覆盖9个技能维度,涉及感知、理解和生成能力。数据集中的视频来自YouTube,涵盖超过40个不同的主题。通过使用基于大型语言模型的评估方法,AutoEval-Video能够高效地评估对开放式问题的响应,特别开发了对抗性标注机制以提高规则的鲁棒性。该数据集的应用领域包括视频理解、时间动态理解等,旨在解决当前模型在这些领域的局限性。
AutoEval-Video is an innovative and challenging benchmark dataset co-created by the Qingyuan Research Institute, School of Electronic Information and Electrical Engineering, Shanghai Jiao Tong University, and ByteDance AI Lab. It aims to comprehensively evaluate the performance of large vision-language models in open-ended video question answering. This dataset contains 327 complex open-ended video question answering instances, covering 9 skill dimensions involving perception, understanding and generation capabilities. The videos in the dataset are sourced from YouTube, spanning over 40 distinct topics. By adopting evaluation methods based on large language models, AutoEval-Video can efficiently assess responses to open-ended questions. A dedicated adversarial annotation mechanism was developed to enhance the robustness of the annotation rules. Its application fields include video understanding, temporal dynamic understanding and other related areas, aiming to address the limitations of current models in these domains.




