图片版面分析标注数据
收藏资源简介:
本数据集主要应用于光学图像识别(OCR)领域的模型训练。数据内容是对各行业搜集到的开源图片进行专业规范的人工标注,从图片中标注出标题、正文、图片、页眉、页脚、表格、目录、公式等。图片涉及到的领域众多,包括报纸、政府公开文件、各类报告、海报、手稿、表格等。 致力于提供高质量的人工标注数据,服务于人工智能领域的图像训练。
This dataset is primarily intended for model training in the field of Optical Character Recognition (OCR). Its content consists of professionally standardized manual annotations performed on open-source images collected from various industries, where the annotations label various elements within the images including titles, main body text, embedded images, headers, footers, tables, table of contents, and mathematical formulas, among others. The included images span a diverse range of domains, such as newspapers, official government documents, various types of reports, posters, manuscripts, tables, and more. It is dedicated to providing high-quality manually annotated data to support image training tasks in the field of artificial intelligence.




