EgoMAN
收藏资源简介:
EgoMAN是由Meta与华盛顿大学联合构建的大规模自我中心视角交互数据集,包含300+小时视频、1500+场景的219,000条6自由度手腕轨迹数据及300万结构化视觉-语言-运动问答对。数据集通过Aria眼镜采集,整合了EgoExo4D、Nymeria和HOT3D-Aria等多源数据,涵盖烹饪、自行车维修等日常活动,每条轨迹均标注了接近(approach)和操控(manipulation)双阶段时空信息。其创新性地采用GPT-4自动生成语义推理(21.6%)、空间推理(42.6%)和运动推理(35.8%)三类问答对,通过轨迹令牌(<START>/<CONTACT>/<END>)实现意图与运动的显式关联,为3D手部轨迹预测、人机交互及具身智能研究提供了首个融合多模态推理的基准平台。
EgoMAN is a large-scale egocentric interaction dataset jointly constructed by Meta and the University of Washington. It contains over 300 hours of video, 219,000 6-degree-of-freedom wrist trajectory entries across 1,500+ scenes, and 3 million structured visual-language-motion question-answering pairs. The dataset is collected via Aria glasses, and integrates multi-source data including EgoExo4D, Nymeria, and HOT3D-Aria. It covers daily activities such as cooking and bicycle repair. Each trajectory is annotated with two-stage spatio-temporal information for the approach and manipulation phases. Notably, this dataset innovatively uses GPT-4 to automatically generate three categories of question-answering pairs: semantic reasoning (21.6%), spatial reasoning (42.6%), and motion reasoning (35.8%). Explicit association between movement intent and motion is achieved through trajectory tokens (<START>/<CONTACT>/<END>). It serves as the first benchmark platform integrating multimodal reasoning for research in 3D hand trajectory prediction, human-computer interaction, and embodied intelligence.
EgoMAN 数据集概述
数据集名称
EgoMAN 数据集
核心描述
EgoMAN 是一个大规模的第一人称视角(Egocentric)数据集,用于支持交互阶段感知的3D手部轨迹预测。它包含219K个6自由度(6-DoF)手部轨迹和3M个结构化问答对,用于语义、空间和运动推理。
关键数据构成
- 轨迹数据:219,000个6自由度(6-DoF)手部轨迹。
- 问答对数据:3,000,000个结构化问答对,涵盖语义、空间和运动推理。
数据集特点与应用
- 主要任务:用于交互阶段感知的3D手部轨迹预测。
- 数据来源:源自第一人称视角的人类交互视频。
- 关联模型:与EgoMAN模型(一个从推理到运动的框架)共同提出,该框架通过轨迹-令牌接口连接视觉-语言推理和运动生成。
外部资源链接
- 论文:https://arxiv.org/abs/2512.16907
- 代码:标注为“Under Internal Review”。

- 1Flowing from Reasoning to Motion: Learning 3D Hand Trajectory Prediction from Egocentric Human Interaction VideosMeta, 华盛顿大学 · 2025年



