NaijaVoices
收藏资源简介:
NaijaVoices数据集是一个大规模、高质量、文化丰富的语音文本数据集,专为非洲语言而设计。该数据集由Lanfrica、NaijaVoices社区和Mila - 魁北克人工智能研究所等机构合作创建,旨在弥补非洲语言在语音技术中的数据不足问题。数据集包含超过1800小时的语音数据,来自5000多名演讲者,覆盖了伊博语、豪萨语和约鲁巴语等语言。数据集的创建过程采用了独特的“数据农业”方法,确保数据提供社区在数据收集过程中得到参与、赋权和互利。NaijaVoices数据集在自动语音识别方面进行了微调实验,平均实现了75.86%(Whisper)、52.06%(MMS)和42.33%(XLSR)的词错误率(WER)改进,展示了其在多语言语音处理方面的潜力。
The NaijaVoices dataset is a large-scale, high-quality, culturally rich speech-text dataset designed specifically for African languages. Co-created by institutions including Lanfrica, the NaijaVoices Community, and Mila – Quebec Artificial Intelligence Institute, this dataset aims to address the shortage of data for African languages in speech technology. It contains over 1,800 hours of speech data from more than 5,000 speakers, covering languages such as Igbo, Hausa, and Yoruba. The dataset was developed using a unique "data farming" approach, which ensures that the data-providing communities are engaged, empowered, and mutually beneficial throughout the data collection process. Fine-tuning experiments on automatic speech recognition (ASR) using the NaijaVoices dataset achieved average Word Error Rate (WER) improvements of 75.86% (Whisper), 52.06% (MMS), and 42.33% (XLSR), demonstrating its potential for multilingual speech processing.
- 1The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African LanguagesLanfrica, NaijaVoices, Mila - Quebec AI Institute, MLCollective, Ohio State University, INRIA, France, Alex Ekwueme Federal University Ndufu Alike Ikwo, Nigeria, Obafemi Awolowo University, Nigeria, Polytechnique Montreal, Canada, University of Montreal, Canada, McGill University, Canada, Canada CIFAR AI Chair · 2025年



