PRALEKHA
收藏资源简介:
PRALEKHA是由AI4Bharat的尼勒卡尼中心创建的一个大规模文档对齐评估基准数据集,涵盖11种印度语言和英语,包含超过200万份文档,其中对齐与未对齐文档的比例为1:2。数据集内容包括新闻简报和播客脚本,涵盖书面和口语形式。数据集的创建过程包括从印度新闻信息局和曼尼·基巴特广播节目等可靠平台收集和校对数据。PRALEKHA旨在评估和提升多语言文档对齐技术,特别是在印度语言中的应用,以支持文档级神经机器翻译等应用。
PRALEKHA is a large-scale document alignment evaluation benchmark dataset developed by the Neelakantan Center at AI4Bharat. It covers 11 Indian languages and English, containing over 2 million documents with an aligned-to-unaligned document ratio of 1:2. The dataset encompasses news briefs and podcast scripts, spanning both written and spoken modalities. The dataset construction process involved collecting and proofreading data from trusted platforms including the Press Information Bureau and the Mann Ki Baat radio program. PRALEKHA is designed to evaluate and enhance multilingual document alignment technologies, especially their deployment in Indian languages, to facilitate applications such as document-level neural machine translation.

- 1Pralekha: An Indic Document Alignment Evaluation BenchmarkAI4Bharat的尼勒卡尼中心 · 2024年



