youssefkhalil320/urdu_images_doc_tags_all_v10
收藏官方服务:
资源简介:
这是一个包含文档图像和相关元数据的数据集,其中包括文档的ID、源PDF文件路径、图像、图像预览、HTML和OTS标签格式、Markdown格式文本、语言及其置信度、难度评分、页面文本长度、是否缺失边界框、是否含有非全宽文本、页面宽度和高度等信息。数据集分为训练集,共有6741个样本。
This dataset consists of document images and associated metadata, including document ID, source PDF path, image, image preview, HTML and OTSL tag formats, Markdown text, language and its confidence, difficulty score, page text length, whether missing bounding box, whether containing non-full width text, page width, and height. The dataset is split into a training set with a total of 6741 samples.
提供机构:
youssefkhalil320


