MURI-IT
收藏资源简介:
MURI-IT数据集由慕尼黑大学语言技术实验室创建,包含2,228,499条指令-输出对,涵盖200种语言。该数据集通过多语言反向指令方法生成,利用现有高质量的多语言文本数据,确保文化相关性和多样性。数据集内容丰富,包括来自Wikipedia、WikiHow等多个来源的文本,保留了原始语言的文化和语言细节。创建过程中,通过机器翻译和反向指令生成技术,确保指令与输出对的高质量匹配。MURI-IT数据集主要应用于低资源语言的自然语言处理任务,旨在提升大型语言模型在这些语言上的表现。
The MURI-IT dataset was developed by the Language Technology Laboratory of Ludwig Maximilian University of Munich (LMU Munich). It contains 2,228,499 instruction-output pairs spanning 200 languages. This dataset is constructed using multilingual reverse instruction methodologies, leveraging existing high-quality multilingual textual data to guarantee cultural relevance and diversity. The dataset encompasses rich content, including texts from multiple sources such as Wikipedia and WikiHow, and retains the cultural and linguistic specifics of the original languages. During the development process, machine translation and reverse instruction generation technologies are utilized to ensure high-quality alignment between the instruction and output pairs. The MURI-IT dataset is primarily applied to natural language processing tasks for low-resource languages, with the goal of enhancing the performance of large language models (LLMs) on these languages.

- 1MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions慕尼黑大学语言技术实验室 · 2024年



