MC2
收藏资源简介:
MC2是由北京大学创建的多语种少数民族语言数据集,是目前最大的开源数据集,涵盖了藏语、维吾尔语、哈萨克语和蒙古语四种语言。数据集通过高质量的网络爬虫技术收集,确保了数据的准确性和多样性。MC2特别关注少数民族语言的书写系统,首次收集了哈萨克语阿拉伯文和传统蒙古文的数据。该数据集旨在提升少数民族语言在人工智能服务中的平等性,为低资源语言的研究提供可靠的数据基础。
MC2 is a multilingual minority language dataset developed by Peking University. It is currently the largest open-source dataset covering four minority languages: Tibetan, Uyghur, Kazakh, and Mongolian. The dataset is collected using high-quality web crawling technologies, which ensures the accuracy and diversity of the data. MC2 places special emphasis on the writing systems of minority languages, and it is the first to collect data in Arabic-script Kazakh and traditional Mongolian scripts. The dataset aims to promote linguistic equality for minority languages in AI services, providing a reliable data foundation for research on low-resource languages.




