EgoGazeVQA
收藏资源简介:
EgoGazeVQA数据集是首个以第一人称视角视频为基础,并结合眼动追踪数据,用于评估大型多模态语言模型(MLLMs)在理解用户意图方面的性能的基准数据集。该数据集包含了从Ego4D、EgoExo4D和EGTEA Gaze+三个主要的第一人称视频数据集中提取的900个视频片段,以及由MLLMs生成的1757个基于眼动和文本描述的问答对。每个问答对都经过了人工审核,以确保其相关性和准确性。EgoGazeVQA数据集旨在帮助MLLMs更好地理解用户在日常生活场景中的意图和活动,从而提升人工智能助手的个性化和主动性。
EgoGazeVQA is the first benchmark dataset based on first-person perspective videos combined with eye-tracking data, dedicated to evaluating the performance of large multimodal language models (MLLMs) in understanding user intentions. The dataset contains 900 video clips extracted from three major first-person video datasets, namely Ego4D, EgoExo4D, and EGTEA Gaze+, along with 1757 question-answer pairs based on eye-tracking and text descriptions generated by MLLMs. Each question-answer pair has undergone manual review to ensure its relevance and accuracy. The EgoGazeVQA dataset aims to help MLLMs better understand user intentions and activities in daily life scenarios, thereby enhancing the personalization and initiative of AI assistants.

- 1通过北京航空航天大学计算机科学与工程学院虚拟现实技术与系统国家重点实验室; 清华大学人工智能学院 · 2025年



