遇见数据集

In-house Code-Switch ASR Dataset

收藏
arXiv2025-09-30 收录
官方服务:

资源简介:

该数据集是一个内部数据集,包含了大约16万小时的标记英文和中文普通话音频,这些音频是在各种不同的声学环境下收集的。此外,该数据集还包括名为“test-name”和“test-term”的测试集,这些测试集是从实际的内部会议中收集的,并配有特定的语境知识偏置列表。整个数据集的规模达到了16万小时,旨在支持自动语音识别(ASR)任务。

This dataset is an internal dataset containing approximately 160,000 hours of annotated English and Mandarin Chinese audio data collected across a diverse range of acoustic environments. Additionally, the dataset includes two test sets named "test-name" and "test-term". These test sets are sourced from actual internal meetings and accompanied by a specific contextual knowledge bias list. The total size of the dataset is 160,000 hours, and it is intended to support automatic speech recognition (ASR) tasks.

提供机构:
In-house
搜集汇总
数据集介绍
In-house Code-Switch ASR Dataset 数据集图片
背景与挑战
背景概述
该数据集是用于上下文语音识别研究的内部大规模数据集,包含160k小时的中文普通话和英语音频,测试集来自真实会议。数据集分为test-name和test-term两个子集,分别针对人名和术语进行上下文偏置,覆盖开放域场景,并提供了详细的统计指标如罕见比率和覆盖率,适用于ASR定制化和个性化研究。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务