ViewSpatial-Bench
收藏资源简介:
ViewSpatial-Bench是一个针对多视角空间定位识别的综合基准,包含超过5700个精心挑选的样本和五种任务类型,用于评估视觉语言模型(VLMs)在3D环境中的空间定位能力。数据集包括从相机和人两个视角出发的任务,涵盖了物体相对方向识别、物体视角方向识别、场景模拟相对方向识别等。数据集来源于ScanNet和MS-CoCo的验证集,通过自动化的3D标注流程生成精确的方向标签,为VLMs的训练提供了丰富的空间关系数据。该数据集旨在解决当前VLMs在跨视角空间理解任务中的局限性,通过建模3D空间关系来增强VLMs的空间理解能力。
ViewSpatial-Bench is a comprehensive benchmark for multi-view spatial localization and recognition. It contains over 5,700 carefully curated samples and five task types, designed to evaluate the spatial localization capabilities of Vision-Language Models (VLMs) in 3D environments. The dataset includes tasks from both camera and human perspectives, covering relative object orientation recognition, object viewpoint direction recognition, simulated scene relative orientation recognition, and other related tasks. Derived from the validation splits of ScanNet and MS-COCO, the dataset generates precise orientation labels through an automated 3D annotation pipeline, providing rich spatial relationship data for the training of VLMs. This benchmark aims to address the current limitations of VLMs in cross-view spatial understanding tasks, and enhance the spatial comprehension ability of VLMs by modeling 3D spatial relationships.
ViewSpatial-Bench 数据集概述
基本信息
- 数据集名称: ViewSpatial-Bench
- 研究领域: 视觉-语言模型(VLMs)的多视角空间定位能力评估
- 作者: Dingming Li, Hongxing Li, Zixuan Wang, Yuchen Yan, Hang Zhang, Siqi Chen, Guiyang Hou, Shengpei Jiang, Wenqi Zhang, Yongliang Shen, Weiming Lu, Yueting Zhuang
- 机构: 浙江大学, 电子科技大学, 香港中文大学
- 状态: 预印本,正在评审中
- 相关资源: 论文 | 代码 | 🤗 数据集
数据集简介
ViewSpatial-Bench 是首个专注于评估视觉-语言模型在多视角空间定位任务中表现的综合性基准测试。该数据集通过自动化的3D方向标注流程生成,包含精确的方向标签,支持五种不同的任务类型,涵盖相机中心视角和人类中心视角。
任务类型
-
相机视角任务:
- Cam-Rel. Dir.: 从图像中直接确定物体之间的空间关系。
- Cam-Obj. Oir.: 从自我中心视角识别个体相对于相机的注视方向。
-
人类视角任务:
- Per-Rel. Dir.: 从图像中角色的视角确定其他物体的空间关系。
- Per-Obj. Oir.: 从图像中角色的位置确定其注视方向。
- Per-Sce. Sim.: 在连续帧中模拟自己在空间场景中的位置,确定其他物体的相对位置。
数据集特点
- 多样性: 包含约43K个多样化的空间关系样本。
- 自动化标注: 利用ScanNet和MS-COCO数据自动生成空间标注。
- 多视角支持: 同时支持相机视角和人类视角的空间推理任务。
性能评估
- 基线模型: 包括Qwen2.5-VL (3B)、GPT-4o和Gemini-2.0-Flash等。
- 改进模型: Multi-View Spatial Model (MVSM) 通过多视角微调策略,在Qwen2.5-VL (3B)上实现了46.24%的整体性能提升。
示例问题
- Per-Sce. Sim.: "站在桌子旁,凝视枕头,架子应该在哪个方向?"
- Cam-Rel. Dir.: "椅子相对于枕头的位置如何?"
- Cam-Obj. Dir.: "以相机镜头为前方,男人朝哪个方向看?"
- Per-Rel. Dir.: "从穿白衣服的男人的视角看,穿红衣服的男人在哪里?"
- Per-Obj. Dir.: "作为照片中穿黑衣服的男人,你面向哪个方向?"
引用
bibtex @misc{li2025viewspatialbenchevaluatingmultiperspectivespatial, title={ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models}, author={Dingming Li and Hongxing Li and Zixuan Wang and Yuchen Yan and Hang Zhang and Siqi Chen and Guiyang Hou and Shengpei Jiang and Wenqi Zhang and Yongliang Shen and Weiming Lu and Yueting Zhuang}, year={2025}, eprint={2505.21500}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2505.21500}, }




