The Historical Uyghur-Chinese Corpus
收藏资源简介:
该数据集包含自晚清至今的Uyghur和Chinese双语文档样本,涵盖法律、公告、官方期刊、报纸和书籍等公共文档。旨在记录和展示这一时期内亚地区关于Chinese和Uyghur语言的官方翻译实践和政策。数据集包含超过200份Uyghur语言文档,每份文档均附有对应的Chinese版本。
This dataset comprises bilingual document samples in Uyghur and Chinese from the late Qing Dynasty to the present, encompassing public documents such as laws, announcements, official journals, newspapers, and books. It aims to document and showcase the official translation practices and policies regarding the Chinese and Uyghur languages in the Inner Asian region during this period. The dataset includes over 200 Uyghur language documents, each accompanied by its corresponding Chinese version.
数据集概述
数据集名称
The Historical Uyghur-Chinese Corpus
数据集内容
该数据集包含自晚清时期至现代的公共文档样本,这些文档以维吾尔语和汉语发布。公共文档包括法律、公告、官方期刊、报纸和书籍等。
数据集目的
记录和展示这一时期内亚地区关于汉语和维吾尔语的官方翻译实践和政策。
数据集组成
数据集分为六个子集合,每个子集合代表不同行政时期或不同类型的文档:
- QA - 12份清代档案文档,全部为双语。
- QB - 50份清代档案文档,全部为双语。
- RA - 5份民国时期档案文档,全部为双语。
- RB - 112篇民国时期的中文文章,其中34篇配有维吾尔语翻译。
- PA - 50份中华人民共和国时期的法律文档,全部为双语。
- PC - 50篇中华人民共和国时期的政策讨论文章,全部为双语。
语言和文字特点
- 维吾尔语:
- 1949年前的文档使用“Turki”,与现代维吾尔语在拼写和形态上有差异。
- 1949年后的文档使用标准现代维吾尔语。
- 汉语:
- QA和QB子集合使用传统文言文,与现代汉语有差异。
- RB和RA子集合使用民国时期的汉语,与现代汉语在词汇上有所不同。
- PA和PC子集合使用现代标准汉语。
项目团队和资金
该项目由香港浸会大学资助,项目名为“History and Politics of Translation: Chinese Public Documents in Inner Asia – Uyghur-Language Module"。
引用信息
Yeung, Jessica; Christian Faggionato; Merhaba Eli; Ahmet Hojam; Robert Barnett; Jenny Li; Phoebe Shing; Sezen Özkan; and Nathan Hill (2021). "The Historical Uyghur-Chinese Corpus", Github repository, https://github.com/HKBUproject/historical-uyghur-chinese-corpus.




