遇见数据集

Data set of the paper "Publishing an OCR ground truth data set for reuse in an unclear copyright setting"

收藏
NIAID Data Ecosystem2026-03-12 收录
数据链接:
官方服务:

资源简介:

The data set consists of a METS file for each of the PDFs that were used for transcription and a directory data/page_xml that contains the transcriptions of the ground truth in PAGE-XML format. In parallel to the data set publication, a data paper will be published that contains a detailed description of the data set. As soon as it is published, we will link to it. The corresponding source code can be found here https://github.com/millawell/ocr-data/tree/1.1

创建时间:
2021-05-12
二维码
社区交流群
二维码
科研交流群
商业服务