ArcMMLU
收藏资源简介:
ArcMMLU是一个专为中文图书馆与信息科学领域设计的大型语言模型评估基准。该数据集由武汉大学信息管理学院创建,包含6210个高质量的单选题,覆盖档案学、数据科学、图书馆学和信息科学四个主要子领域。数据集的构建过程包括从实际的研究生入学考试、专业考试、课程测验和学术竞赛中收集原始数据,通过光学字符识别(OCR)技术提取文本内容,并进行人工检查和质量过滤。ArcMMLU旨在通过这些精心设计的题目,全面评估和提升大型语言模型在特定领域的应用能力,特别是在复杂信息检索、数据组织和档案图书馆环境中的用户交互等方面。
ArcMMLU is a large language model (LLM) evaluation benchmark specifically designed for the Chinese library and information science domain. This dataset, created by the School of Information Management, Wuhan University, contains 6,210 high-quality multiple-choice questions covering four core subfields: archival science, data science, library science, and information science. The dataset construction process involves collecting raw data from actual graduate entrance examinations, professional exams, course quizzes, and academic competitions, extracting text content via Optical Character Recognition (OCR) technology, and conducting manual inspection and quality filtering. ArcMMLU aims to comprehensively evaluate and enhance the application capabilities of LLMs in targeted domains through these meticulously designed questions, particularly in areas such as complex information retrieval, data organization, and user interaction in archival and library environments.




