Corpus of Chinese Dynastic Histories
收藏资源简介:
Corpus of Chinese Dynastic Histories(CCDH)是一个包含超过2300万字符的开放源数据集,涵盖了从公元前3世纪至公元18世纪的中国历代历史文献。该数据集由马萨诸塞大学阿默斯特分校和多伦多大学的研究团队创建,主要用于计算分析历史词汇和语义变化。数据集内容丰富,包括各朝代的官方历史记录,如《史记》、《汉书》等,这些文献以古典汉语书写,为研究提供了丰富的语言材料。创建过程中,研究团队从Wikisource获取文本,并进行了格式化和处理,以适应各种古典汉语研究的需求。该数据集的应用领域广泛,特别适用于历史语言学、语义学及性别研究等领域,旨在解决古典汉语资源稀缺的问题,推动相关学术研究的发展。
The Corpus of Chinese Dynastic Histories (CCDH) is an open-source dataset containing over 23 million characters, covering historical documents of successive Chinese dynasties spanning from the 3rd century BCE to the 18th century CE. Developed by a research team from the University of Massachusetts Amherst and the University of Toronto, this dataset is primarily intended for computational analysis of historical lexical and semantic changes. It encompasses a rich array of official historical records from various dynasties, such as the Records of the Grand Historian and the Book of Han. Written in Classical Chinese, these documents provide ample linguistic resources for academic research. During its development, the research team sourced texts from Wikisource and conducted formatting and processing to accommodate the requirements of diverse Classical Chinese research endeavors. This dataset has a wide range of application domains, and is particularly applicable to fields including historical linguistics, semantics and gender studies. It aims to address the scarcity of available Classical Chinese resources and advance the development of relevant academic research.




