Chinese Literature NER-RE Dataset
收藏资源简介:
本数据集名为‘Chinese Literature NER-RE Dataset’,由北京大学电子工程与计算机科学学院创建,包含726篇文章,总计超过100,000个字符。该数据集从数百篇中文文学文章中构建,旨在提供一个话语级别的资源,以增强中文文学文本中的命名实体识别和关系抽取任务。创建过程中采用了启发式标记方法和机器辅助标记方法,确保数据的一致性和质量。该数据集特别适用于研究中文文学文本中的实体和关系,为相关研究提供了重要的基准和资源。
This dataset, named 'Chinese Literature NER-RE Dataset', was created by the School of Electronic Engineering and Computer Science, Peking University. It contains 726 articles with a total of over 100,000 characters. Constructed from hundreds of Chinese literary articles, this dataset aims to provide a discourse-level resource to advance Named Entity Recognition (NER) and Relation Extraction (RE) tasks on Chinese literary texts. Both heuristic annotation and machine-assisted annotation methods were employed during its creation to ensure data consistency and quality. This dataset is particularly suitable for research on entities and relations in Chinese literary texts, serving as an important benchmark and resource for relevant academic studies.




