CMRC 2019
收藏资源简介:
CMRC 2019是由哈尔滨工业大学社会计算与信息检索研究中心和科大讯飞认知智能国家重点实验室联合创建的中文机器阅读理解数据集,包含超过10万个空白(问题)分布在1万多篇源自中文叙事故事的文本中。数据集通过精心设计,引入了假候选句子以增加难度,要求机器在填充空白时进行句子级别的推理并识别真假句子。该数据集主要用于评估和提升机器在处理复杂语言理解任务时的能力,特别是在需要综合多线索进行推理的场景中。
CMRC 2019 is a Chinese machine reading comprehension dataset jointly created by the Social Computing and Information Retrieval Research Center of Harbin Institute of Technology and iFLYTEK State Key Laboratory of Cognitive Intelligence. It contains over 100,000 blanks (questions) distributed across more than 10,000 texts sourced from Chinese narrative stories. The dataset is meticulously designed to introduce false candidate sentences to increase task difficulty, requiring machines to conduct sentence-level reasoning and distinguish between genuine and fake sentences when filling in the blanks. This dataset is primarily used to evaluate and enhance the capabilities of machines in handling complex language understanding tasks, especially in scenarios that require comprehensive multi-clue reasoning.

- 1A Sentence Cloze Dataset for Chinese Machine Reading Comprehension哈尔滨工业大学社会计算与信息检索研究中心 · 2020年



