Egocentric2Embodiment dataset (E2E-3M)
收藏资源简介:
E2E-3M数据集是由香港科技大学(广州)和中关村研究院等机构联合构建的大规模第一人称视角视频问答数据集,旨在通过人类第一人称视频数据提升机器人的物理智能。该数据集包含300万条结构化标注数据,数据来源于Ego4D、BuildAI和EgoDex等多源人类第一人称视频,通过Egocentric2Embodiment翻译流程将原始视频转化为多层次的视觉问答监督信号。数据集标注过程采用模式驱动的自动化流程,涵盖时间、空间、力学等七种互补的问答模式,并通过规则验证确保标注质量。该数据集主要用于训练和评估第一人称视角下的视觉语言动作(VLA)模型,解决机器人领域因缺乏大规模第一人称数据导致的规划与交互推理能力不足问题。
The E2E-3M dataset is a large-scale first-person video question answering (QA) dataset jointly developed by institutions including Hong Kong University of Science and Technology (Guangzhou) and Zhongguancun Institute. It aims to enhance the physical intelligence of robots via human first-person video data. This dataset contains 3 million structured annotated samples, sourced from multiple human first-person video datasets such as Ego4D, BuildAI, and EgoDex. Raw videos are converted into multi-level visual QA supervision signals via the Egocentric2Embodiment pipeline. The dataset’s annotation process adopts a pattern-driven automated workflow, covering seven complementary QA modes including temporal, spatial, mechanical and other categories, and ensures annotation quality through rule-based validation. It is primarily used to train and evaluate first-person visual language action (VLA) models, addressing the insufficient planning and interactive reasoning capabilities in robotics caused by the lack of large-scale first-person video datasets.

- 1PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence香港科技大学(广州)、中关村研究院、中关村人工智能研究所、哈尔滨工业大学、华中科技大学 · 2025年



