AGQA 2.0
收藏资源简介:
AGQA 2.0是一个用于评估视频时空组合推理能力的数据集,由华盛顿大学和斯坦福大学联合开发。该数据集包含9685万个问题答案对,主要通过自然语言模板和场景图注释生成关于视频的问题,旨在测试模型的组合推理能力。数据集创建过程中,采用了严格的平衡算法来减少语言偏差,确保问题答案分布的平衡性。AGQA 2.0主要应用于视频理解和视觉推理领域,特别是在评估模型对复杂视觉场景的理解能力方面具有重要价值。
AGQA 2.0 is a dataset designed to evaluate spatiotemporal compositional reasoning capabilities for videos, jointly developed by the University of Washington and Stanford University. Comprising 96.85 million question-answer pairs, it primarily generates questions about video content via natural language templates and scene graph annotations, with the goal of testing models' compositional reasoning abilities. During its construction, strict balancing algorithms are utilized to reduce language bias and ensure a balanced distribution of question-answer pairs. AGQA 2.0 is mainly applied in the domains of video understanding and visual reasoning, and it holds substantial value particularly for assessing models' capacity to understand complex visual scenes.

- 1AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning华盛顿大学 · 2022年



