Native Chinese Reader (NCR)
收藏资源简介:
Native Chinese Reader (NCR) 是一个专为机器阅读理解设计的数据集,包含8390个文档,平均长度为1024字,涵盖现代和古典中文多种文体。该数据集源自中国高中语文课程的考试题目,旨在评估母语为中文的青少年的语言能力。与现有中文MRC数据集相比,NCR不仅文档长度更长,问题也更具有挑战性,需要较强的推理能力和常识知识来解答。NCR的应用领域主要集中在提升中文自然语言理解能力,特别是在古典文学和诗歌的理解上,旨在缩小当前MRC模型与母语使用者之间的性能差距。
Native Chinese Reader (NCR) is a machine reading comprehension (MRC) dataset consisting of 8,390 documents, with an average length of 1,024 Chinese characters, covering various literary styles of both modern and classical Chinese. Derived from Chinese high school Chinese language curriculum exam questions, this dataset is designed to evaluate the language proficiency of adolescent native Chinese speakers. Compared with existing Chinese MRC datasets, NCR not only features longer documents but also poses more challenging questions that require robust reasoning abilities and commonsense knowledge to solve. The main application domains of NCR focus on enhancing Chinese natural language understanding capabilities, particularly for the comprehension of classical Chinese literature and poetry, with the aim of narrowing the performance gap between current MRC models and native Chinese speakers.




