面向示教视频理解的综合性数据集
收藏资源简介:
本数据集面向机器人示教视频理解任务构建,基于大规模机器人操作任务室内场景操控数据集 DROID 筛选加工得到(原始DROID数据集的大小为8.7TB),总大小35GB,主要包含以下两部分: 1. 示教视频多视角关键帧图像:从 DROID 中筛选得到约 49,935 段独立示教视频,对每段示教视频从不同视角均匀采样裁剪为最多 16 帧关键帧,用于视频理解/动作计划推断等任务。 2. 文本与结构化标注:对齐每段示教视频的人类指令,并基于关键帧参考生成简要动作序列标注(“plan”和“code”两类信息),以 JSON形式提供。
This dataset is constructed for robot teaching video understanding tasks, and is screened and processed based on the large-scale indoor scene manipulation dataset DROID for robotic manipulation tasks (the original DROID dataset has a size of 8.7 TB). The total size of this dataset is 35 GB, which mainly includes the following two parts: 1. Multi-view keyframe images of teaching videos: Approximately 49,935 independent teaching video clips are screened from DROID. For each teaching video clip, up to 16 keyframes are uniformly sampled and cropped from different perspectives, which are suitable for tasks such as video understanding and action plan inference. 2. Text and structured annotations: Human instructions corresponding to each teaching video clip are aligned, and concise action sequence annotations with two types of information, "plan" and "code", are generated based on the keyframe references. All annotations are provided in JSON format.




