NVSpeech
收藏资源简介:
NVSpeech是一个用于中文语音中副语言发声识别和合成的集成和可扩展的流程。该数据集包含48,430个人类语音的句子,带有18个词级副语言类别标签。通过使用副语言感知的ASR模型,自动标注了一个包含174,179个句子的大规模语料库,并支持词级对齐和副语言线索。NVSpeech通过统一副语言发声的识别和生成,提供了第一个开放的大型词级标注的流程,以支持普通话中表达性语音建模,并以可扩展和可控的方式进行识别和合成。数据集和音频演示可在https://nvspeech170k.github.io/获取。
NVSpeech is an integrated and extensible pipeline for paralinguistic vocalization recognition and synthesis in Mandarin Chinese. This dataset includes 48,430 human speech sentences annotated with 18 word-level paralinguistic category labels. A large-scale corpus of 174,179 sentences was automatically annotated using a paralinguistics-aware ASR model, and this corpus supports word-level alignment and paralinguistic cues. By unifying the recognition and generation of paralinguistic vocalizations, NVSpeech provides the first open, large-scale word-level annotated pipeline to support expressive speech modeling in Mandarin, enabling scalable and controllable recognition and synthesis. The dataset and audio demos are accessible at https://nvspeech170k.github.io/
NVSpeech数据集概述
数据集基本信息
- 名称: NVSpeech
- 类型: 语音数据集(包含副语言特征标注)
- 语言: 中文普通话为主(含少量英文示例)
- 规模:
- 人工标注部分: 48,430条话语
- 自动标注部分: 174,179条话语(573小时)
- 标注级别: 词级别
- 标注类别: 18种副语言类别
核心特征
-
副语言标注:
- 包含非言语声音(如[Laughter]、[Breathing])
- 词汇化插入语(如[Uhm]、[Oh])
- 情感/态度标记(如[Surprise-oh]、[Dissatisfaction-hnn])
-
多任务支持:
- 自动语音识别(ASR)
- 文本到语音合成(TTS)
- 联合建模能力
技术贡献
-
数据集构建:
- 首个大规模中文词级别副语言标注数据集
- 包含人工标注和自动标注双版本
-
模型创新:
- 副语言感知ASR模型(将副语言标记作为可解码token)
- 可控TTS系统(支持副语言特征的显式控制)
-
完整流程:
- 数据标注 → ASR建模 → 自动标注 → TTS训练
应用示例
-
TTS合成控制: text "还需要[Breathing]...调整" "[Surprise-oh],这才对嘛![Breathing]"
-
ASR转录示例: text "慧星是追求远方的家伙,[Breathing]每隔一段时间..." "[Question-ei]?前面有一个吃饭的地方..."
数据类别示例
| 类型 | 示例标记 |
|---|---|
| 非言语声音 | [Laughter], [Cough], [Sigh] |
| 情感/态度 | [Surprise-wa], [Dissatisfaction-hnn] |
| 话语标记 | [Question-en], [Confirmation-en] |

- 1NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations香港中文大学(深圳) · 2025年



