ScanAlign
收藏资源简介:
ScanAlign数据集由上海人工智能实验室和香港大学的研究团队创建,旨在通过视频帧、BEV图像和文本注释来增强视觉语言模型对3D场景的理解。该数据集包含165,000条文本注释,数据来源于室内场景的视频输入,并通过3D重建技术生成BEV图像。数据集的创建过程包括从视频中提取帧、生成BEV图像,并在图像和视频帧中添加空间-时间对象标记(STO标记)。该数据集主要用于3D场景理解任务,如3D问答、密集描述和视觉定位,旨在解决视觉语言模型在3D空间理解中的局限性问题。
The ScanAlign dataset was created by research teams from the Shanghai AI Laboratory and The University of Hong Kong. It aims to enhance the 3D scene understanding capabilities of vision-language models via video frames, BEV images and textual annotations. This dataset contains 165,000 textual annotations, with data sourced from video inputs of indoor scenes, and BEV images generated through 3D reconstruction techniques. The dataset creation process includes extracting frames from videos, generating BEV images, and adding Spatial-Temporal Object (STO) tags to both images and video frames. It is primarily used for 3D scene understanding tasks such as 3D question answering, dense captioning and visual grounding, targeting the limitations of vision-language models in 3D spatial understanding.




