VIMA
收藏资源简介:
VIMA-BENCH 是由斯坦福大学、NVIDIA 等机构联合开发的多模态机器人操作任务基准数据集,旨在推动机器人在多样化任务中的泛化能力研究。该数据集基于 Ravens 模拟器构建,包含 17 个任务模板,涵盖简单物体操作、视觉目标达成、新概念理解、单次视频模仿、视觉约束满足和视觉推理等六类任务。通过程序化生成,可扩展出数千个任务实例,搭配多模态提示,包括文本、图像和视频帧,为机器人提供丰富的任务描述形式。数据集包含 600K+ 专家轨迹,用于模仿学习,涵盖多种任务场景和操作方式。其创建过程充分考虑了任务多样性和泛化能力测试,设计了从随机化物体放置到全新任务的四层评估协议,系统性地衡量机器人在不同难度下的零样本泛化能力。VIMA-BENCH 主要应用于机器人学习领域,致力于解决机器人在面对多样化任务时的泛化难题,使机器人能够通过少量示例快速适应新任务,提升其在复杂环境中的适应性和灵活性。
VIMA-BENCH is a multimodal robotic manipulation task benchmark dataset jointly developed by Stanford University, NVIDIA and other institutions, aiming to advance research on the generalization capability of robots across diverse tasks. Built on the Ravens simulator, this dataset includes 17 task templates covering six categories of tasks: simple object manipulation, visual goal achievement, novel concept understanding, one-shot video imitation, visual constraint satisfaction, and visual reasoning. Generated through procedural methods, it can be scaled to produce thousands of task instances, and is equipped with multimodal prompts including text, images and video frames to provide rich task description modalities for robotic systems. The dataset contains over 600,000 expert trajectories for imitation learning, covering a wide range of task scenarios and manipulation approaches. Its development fully accounts for task diversity and generalization capability testing, and designs a four-layer evaluation protocol spanning from randomized object placement to entirely novel tasks, which systematically evaluates the zero-shot generalization ability of robots across varying difficulty levels. VIMA-BENCH is primarily applied in the field of robotic learning, dedicated to addressing the generalization challenges of robots when facing diverse tasks, enabling robots to rapidly adapt to new tasks with limited demonstrations, and enhancing their adaptability and flexibility in complex environments.




