Plex RoboSuite
收藏资源简介:
PLEX 是由微软研究院等机构开发的用于机器人操作预训练的数据集和模型架构。该数据集旨在通过少量任务无关的视触觉运动轨迹和大量任务相关的物体操作视频来学习丰富的机器人操作表示。数据集包含多任务视频演示(Dmtvd)、视触觉轨迹(Dvmt)和目标任务演示(Dttd)三种类型的数据。Dmtvd 数据丰富多样,涵盖各种任务的高质量视频演示;Dvmt 数据则包含机器人感知与动作的匹配序列;Dttd 数据虽稀缺但质量高,针对特定任务。PLEX 通过在这些数据上进行预训练和微调,能够在 Meta-World 和 Robosuite 等基准测试中展现出卓越的零样本性能和微调能力。该数据集的创建过程充分利用了现有的多模态数据资源,并通过巧妙的模型设计实现了高效的数据利用。PLEX 的应用领域主要集中在机器人操作任务的预训练和泛化能力提升上,旨在解决机器人在面对未见过的任务时的适应性问题。
PLEX is a dataset and model architecture developed for robot manipulation pre-training by institutions such as Microsoft Research. This dataset aims to learn rich robot manipulation representations using a small number of task-agnostic visuo-tactile motion trajectories and a large volume of task-specific object manipulation videos. The dataset contains three types of data: multi-task video demonstrations (Dmtvd), visuo-tactile motion trajectories (Dvmt), and target task demonstrations (Dttd). Dmtvd data is diverse and rich, covering high-quality video demonstrations across various tasks; Dvmt data includes matched sequences of robot perception and action; Dttd data, although scarce, is of high quality and targets specific tasks. By pre-training and fine-tuning on these data, PLEX exhibits excellent zero-shot performance and fine-tuning capabilities across benchmark tests such as Meta-World and Robosuite. The development of this dataset fully leverages existing multimodal data resources, and achieves efficient data utilization through ingenious model design. The application scenarios of PLEX mainly focus on robot manipulation pre-training and improving generalization capabilities, aiming to address the adaptability problems of robots when facing unseen tasks.




