AMALIA
收藏资源简介:
AMALIA数据集是一个专注于欧洲葡萄牙语(pt-PT)的大型语言模型数据集,由新里斯本大学·科学技术学院、NOVA LINCS等机构联合创建。数据集包含5.8亿个tokens,主要来源于葡萄牙网络档案Arquivo.pt,经过严格的URL过滤、语言识别、去重和质量分类处理。数据集创建过程包括数据收集、过滤、去重和质量分类,最终形成高、中、低三个质量等级的数据。该数据集旨在解决欧洲葡萄牙语在大型语言模型中的不足问题,应用于自然语言处理领域,特别是语言模型训练和评估。
The AMALIA dataset is a large language model dataset dedicated to European Portuguese (pt-PT), jointly created by institutions such as the NOVA School of Science and Technology of NOVA University Lisbon and NOVA LINCS. It contains 580 million tokens, primarily sourced from the Portuguese web archive Arquivo.pt, and has undergone rigorous processing including URL filtering, language identification, deduplication and quality classification. The dataset creation workflow encompasses data collection, filtering, deduplication and quality classification, ultimately yielding data categorized into three quality tiers: high, medium and low. This dataset is designed to address the shortage of European Portuguese resources for large language models, and is applicable to the field of natural language processing, particularly for language model training and evaluation.



