遇见数据集

Dongba1800

收藏
科学数据银行2024-09-12 更新2026-04-23 收录
官方服务:

资源简介:

Dataset for Single Character Detection in Dongba Manuscripts. It includes 1,800 curated JPEG image files and 1,800 text annotation files in TXT format. All files are named in a consistent format to ensure easy indexing and association between images and their corresponding annotations: JPEG images are named 'image_<number>.jpg' (e.g., 'image_1.jpg'), and TXT files are named 'gt_image_<number>.txt' (e.g., 'gt_image_1.txt'). In these TXT files, annotations of Dongba characters include a verified total of 111,702 characters, ensuring the accuracy and reliability of the data. Each character's spatial position is identified by a series of coordinate pairs that define the polygonal boundaries of the text boxes. For example, the coordinate sequence "161, 59, 202, 57, 256, 85, 239, 154, 182, 147, 163, 107" represents the vertices of a polygon, with each pair like "161, 59" indicating the x and y coordinates of a vertex. Coordinates are typically listed in a clockwise direction to comprehensively outline the full contour of the polygon. To differentiate between records, the annotation files use "###" as a delimiter to signify the end of a record.

创建时间:
2024-09-09
二维码
社区交流群
二维码
科研交流群
商业服务