HAUR-5
收藏资源简介:
HAUR-5数据集由厦门大学的研究团队创建,旨在解决多模态模型在理解文本密集图像中人类注释时的局限性。该数据集包含37,702张图像,涵盖了五种常见的人类注释类型,包括高亮、下划线、波浪线、矩形框和段落标记。数据集的创建过程包括从无版权的经典小说中提取文本块,并对其进行注释,最终将注释后的文本转换为图像。该数据集的应用领域主要集中在视觉问答任务,旨在帮助模型更准确地理解和回答与人类注释相关的问题,提升多模态模型在实际场景中的应用效果。
The HAUR-5 dataset was developed by a research team from Xiamen University to address the limitations of multimodal models in understanding human annotations on text-dense images. This dataset comprises 37,702 images covering five common categories of human annotations: highlights, underlines, wavy lines, rectangular boxes, and paragraph markers. The dataset creation workflow includes extracting text blocks from copyright-free classic novels, annotating these text blocks, and subsequently converting the annotated text into image formats. The primary application scope of this dataset lies in visual question answering (VQA) tasks, where it is intended to assist models in more accurately comprehending and answering questions related to human annotations, thereby enhancing the practical application efficacy of multimodal models.

- 1HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images厦门大学 · 2024年



