CMRC 2017
收藏资源简介:
CMRC 2017是由科大讯飞研究院与哈尔滨工业大学联合实验室创建的中文阅读理解数据集,旨在推动中文机器阅读理解研究。该数据集包含两种类型:填空式阅读理解和用户查询阅读理解,涵盖大规模训练数据及人工标注的验证和测试集。数据集内容主要来源于儿童阅读材料,通过自动生成和人工标注相结合的方式创建,适用于机器阅读理解模型的训练与评估,特别是针对中文语言的处理能力提升。
CMRC 2017 is a Chinese machine reading comprehension dataset developed by the joint laboratory of iFLYTEK Research and Harbin Institute of Technology, aiming to advance research in Chinese machine reading comprehension. This dataset covers two types: cloze-style reading comprehension and user query-based reading comprehension, and includes large-scale training data as well as manually annotated validation and test sets. The content of the dataset is mainly derived from children's reading materials, and it is constructed through a combination of automatic generation and manual annotation. It is suitable for training and evaluating machine reading comprehension models, especially for improving the Chinese language processing capabilities of such models.




