vcr-org/VCR-wiki-en-easy-test-100
收藏资源简介:
VCR-Wiki数据集是一个用于视觉字幕恢复(VCR)任务的数据集,旨在评估视觉语言模型在图像中恢复部分遮挡文本的能力。数据集包含图像、堆叠图像、仅包含文本的图像、字幕和交叉文本等字段。数据集的构建过程包括数据收集、初始过滤、N-gram选择、创建嵌入文本的图像、图像拼接和第二轮过滤。数据集的使用受CC BY-SA 4.0许可证约束,适用于研究和教育目的。
The VCR-Wiki dataset is designed for the Visual Caption Restoration (VCR) task, aiming to evaluate the capability of vision-language models to restore partially obscured texts within images. The dataset includes fields such as images, stacked images, images with only text, captions, and crossed texts. The dataset construction process involves data collection, initial filtering, N-gram selection, creating text embedded in images, image concatenation, and second-round filtering. The dataset is licensed under CC BY-SA 4.0 and is intended for research and educational purposes.
VCR-Wiki 数据集概述
数据集描述
VCR-Wiki 数据集是为视觉字幕恢复(Visual Caption Restoration, VCR)任务设计的,旨在评估视觉语言模型在图像中恢复部分遮挡文本的能力。数据集包含图像、字幕以及用于任务的合成图像。
数据集特征
- question_id:
int64,当前分区的实例ID。 - image:
image,原始视觉图像(VI)。 - caption:
string,TEI图像中未遮挡的原始文本。 - stacked_image:
image,包含原始视觉图像和遮挡文本嵌入图像的堆叠图像。 - only_it_image:
image,遮挡的TEI图像。 - only_it_image_small:
image,小尺寸的遮挡TEI图像。 - crossed_text:
List[string],当前实例中遮挡的n-gram。
数据集分割
- test: 包含100个样本,总字节数为19073565。
数据集大小
- 下载大小: 19047792字节
- 数据集大小: 19073565字节
数据集配置
- default: 包含测试数据文件,路径为
data/test-*。
数据集来源
- wikimedia/wit_base
任务类别
- visual-question-answering
语言
- en
数据集构建
- 数据收集和初步过滤: 从
wikimedia/wit_base收集数据,并过滤掉包含敏感内容的实例。 - N-gram选择: 截断描述并使用spaCy进行分词,随机遮挡5-gram。
- 创建嵌入文本的图像: 将文本嵌入图像中,并根据任务难度调整遮挡矩形的大小。
- 图像拼接: 将TEI与VI拼接成堆叠图像。
- 第二轮过滤: 过滤掉没有遮挡n-gram或高度超过900像素的实例。
数据集声明
VCR-Wiki数据集及其子集在CC BY-SA 4.0许可下提供,仅用于视觉字幕恢复及相关视觉语言任务的研究和教育目的。用户需确保其使用符合伦理指南,并遵守许可条款。
引用
bibtex @article{zhang2024vcr, title = {VCR: Visual Caption Restoration}, author = {Tianyu Zhang and Suyuchen Wang and Lu Li and Ge Zhang and Perouz Taslakian and Sai Rajeswar and Jie Fu and Bang Liu and Yoshua Bengio}, year = {2024}, journal = {arXiv preprint arXiv: 2406.06462} }




