遇见数据集

Taxi1500

收藏
arXiv2023-05-15 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

Taxi1500是一个大规模多语言文本分类数据集,由信息与语言处理中心,慕尼黑大学创建。该数据集包含超过1500种语言的1077条圣经经文,主要通过利用圣经的平行翻译来构建。数据集的创建过程涉及开发适用的话题,并通过众包工具收集标注数据。Taxi1500的应用领域主要集中在自然语言处理中的多语言和低资源语言的文本分类问题,旨在通过提供广泛的语言覆盖来解决现有数据集在低资源和濒危语言上的不足。

Taxi1500 is a large-scale multilingual text classification dataset developed by the Center for Information and Language Processing at Ludwig Maximilian University of Munich. This dataset comprises 1077 biblical verses spanning over 1500 languages, and is primarily constructed by utilizing parallel translations of the Bible. The dataset development process involved designing suitable topics and collecting annotated data via crowdsourcing tools. The main application areas of Taxi1500 focus on multilingual and low-resource language text classification tasks in natural language processing, aiming to address the shortcomings of existing datasets in covering low-resource and endangered languages by providing extensive language coverage.

创建时间:
2023-05-15
二维码
社区交流群
二维码
科研交流群
商业服务