eliolio/docvqa
收藏资源简介:
DocVQA数据集是一个用于文档图像视觉问答(VQA)的数据集,包含50,000个问题,这些问题基于12,767张图像。数据集按80−10−10的比例随机划分为训练集、验证集和测试集,其中训练集包含39,463个问题和10,194张图像,验证集包含5,349个问题和1,286张图像,测试集包含5,188个问题和1,287张图像。文档图像来源于UCSF Industry Documents Library,包含打印、打字和手写内容,涵盖了信件、备忘录、笔记、报告等多种文档类型。
The DocVQA dataset is a document image visual question answering (VQA) benchmark dataset, consisting of 50,000 questions grounded in 12,767 images. It is randomly partitioned into training, validation, and test sets at an 80−10−10 ratio. Specifically, the training set holds 39,463 questions and 10,194 images, the validation set contains 5,349 questions and 1,286 images, and the test set includes 5,188 questions and 1,287 images. The document images are sourced from the UCSF Industry Documents Library, covering printed, typed, and handwritten content, and spanning a variety of document types such as letters, memoranda, notes, reports, and more.
DocVQA - A Dataset for VQA on Document Images
数据集概述
- 名称: DocVQA
- 任务类型: 文档图像问答(Document-Question-Answering)
- 数据来源: 文档图像来自UCSF Industry Documents Library,包含打印、打字和手写内容,涵盖信件、备忘录、笔记、报告等多种文档类型。
数据集结构
- 总问题数: 50,000
- 总图像数: 12,767
- 数据分割: 随机分为80-10-10的训练、验证和测试集。
- 训练集: 39,463个问题,10,194张图像
- 验证集: 5,349个问题,1,286张图像
- 测试集: 5,188个问题,1,287张图像
获取方式
- 数据集可从RRC挑战页面的“Downloads”标签下载。
引用信息
@InProceedings{mathew2021docvqa, author = {Mathew, Minesh and Karatzas, Dimosthenis and Jawahar, CV}, title = {Docvqa: A dataset for vqa on document images}, booktitle = {Proceedings of the IEEE/CVF winter conference on applications of computer vision}, year = {2021}, pages = {2200--2209}, }




