遇见数据集

Visual Spatial Reasoning

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

视觉空间推理 (VSR) 语料库是具有真/假标签的字幕图像对的集合。每个标题描述了图像中两个单独对象的空间关系,视觉语言模型 (VLM) 需要判断标题是否正确地描述了图像 (True) 或不正确 (False)。下面是几个例子。

The Visual Spatial Reasoning (VSR) corpus is a collection of image-caption pairs with true/false labels. Each caption describes the spatial relationship between two distinct objects in the corresponding image, and the vision-language model (VLM) is required to determine whether the caption correctly describes the image (True) or not (False). Several examples are provided below.

提供机构:
OpenDataLab
创建时间:
2023-10-20
搜集汇总
数据集介绍
Visual Spatial Reasoning 数据集图片
背景与挑战
背景概述
Visual Spatial Reasoning是一个视觉空间推理数据集,由剑桥大学于2023年发布,包含图文对及其真/假标签,用于测试视觉语言模型对图像中对象空间关系的判断能力。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务