CapNav
收藏资源简介:
CapNav是由华盛顿大学团队构建的能力条件导航基准数据集,旨在评估视觉语言模型在复杂室内环境中考虑不同代理移动约束时的导航性能。该数据集包含45个真实3D扫描室内场景、473个导航任务和2365个问答对,总计5075条可遍历性标注。数据通过人工标注3D场景导航图和代理移动能力参数构建,涵盖五种典型人类和机器人代理配置。该数据集主要应用于具身智能和辅助机器人领域,解决现有导航系统忽视代理物理约束的关键问题,推动能力感知的智能导航技术发展。
CapNav is a capability-conditioned navigation benchmark dataset developed by the University of Washington team, which aims to evaluate the navigation performance of vision-language models when considering different agent movement constraints in complex indoor environments. This dataset contains 45 real 3D-scanned indoor scenes, 473 navigation tasks, and 2365 question-answer pairs, with a total of 5075 traversability annotations. The dataset is constructed via manual annotation of 3D scene navigation graphs and agent mobility parameters, covering five typical human and robotic agent configurations. This dataset is mainly applied in the fields of embodied intelligence and assistive robotics, addressing the key issue that existing navigation systems neglect agent physical constraints, and promoting the development of capability-aware intelligent navigation technologies.
- 1CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation华盛顿大学; 加州大学圣克鲁兹分校 · 2026年



