CN-Celeb
收藏资源简介:
CN-Celeb是由清华大学语音与语言技术中心创建的大型中文语音识别数据集,专注于中国名人语音数据。该数据集包含超过130,000条来自1,000位中国名人的语音记录,覆盖11种不同的语音类型,如娱乐、采访、歌唱等。数据集的创建过程包括自动化数据提取和人工审核,确保数据质量。CN-Celeb主要用于研究在不受限制的环境下的语音识别技术,旨在解决现有技术在复杂环境中的性能问题。
CN-Celeb is a large-scale Chinese speech recognition dataset created by the Center for Speech and Language Technology at Tsinghua University, focusing on speech data of Chinese celebrities. This dataset contains over 130,000 speech recordings from 1,000 Chinese celebrities, covering 11 distinct speech types such as entertainment content, interviews, singing performances and more. The development process of the dataset includes automated data extraction and manual review to ensure high data quality. CN-Celeb is primarily used for researching speech recognition technologies in unrestricted environments, aiming to address the performance issues of existing technologies in complex scenarios.

- 1CN-CELEB: a challenging Chinese speaker recognition dataset清华大学语音与语言技术中心 · 2019年



