CoSpace
收藏资源简介:
CoSpace是由清华大学团队创建的多图像视觉理解基准数据集,旨在评估视觉语言模型在连续空间感知方面的能力。该数据集包含2918张图像和1626个问答对,涵盖了七种类型的任务。数据集通过提供从静态视点拍摄的一系列空间连续的图像来模拟人类对环境的感知,图像之间存在空间重叠,以便模型可以更好地理解图像之间的关联。CoSpace的应用领域主要集中在视觉语言模型的空间理解能力评估,目的是解决模型在处理真实世界场景时遇到的空间信息整合问题。
CoSpace is a multi-image visual understanding benchmark dataset developed by the Tsinghua University team, which is designed to evaluate the continuous spatial perception capabilities of vision-language models. It contains 2918 images and 1626 question-answer pairs, covering seven types of tasks. To simulate human environmental perception, the dataset provides a series of spatially continuous images captured from static viewpoints, with spatial overlaps between the images, enabling models to better comprehend the associative correlations among these images. The primary application scope of CoSpace focuses on the evaluation of spatial understanding capabilities for vision-language models, with the core goal of solving the spatial information integration problems encountered by models when processing real-world scenes.

- 1CoSpace: Benchmarking Continuous Space Perception Ability for Vision-Language Models清华大学计算机科学与技术系, 清华大学人工智能研究院, 北京科技大学计算机与通信工程学院, 上海人工智能实验室, 江苏省语言能力协同创新中心 · 2025年



