In-house Code-Switch ASR Dataset
收藏arXiv2025-09-30 收录
数据链接:
官方服务:
资源简介:
该数据集是一个内部数据集,包含了大约16万小时的标记英文和中文普通话音频,这些音频是在各种不同的声学环境下收集的。此外,该数据集还包括名为“test-name”和“test-term”的测试集,这些测试集是从实际的内部会议中收集的,并配有特定的语境知识偏置列表。整个数据集的规模达到了16万小时,旨在支持自动语音识别(ASR)任务。
This dataset is an internal dataset containing approximately 160,000 hours of annotated English and Mandarin Chinese audio data collected across a diverse range of acoustic environments. Additionally, the dataset includes two test sets named "test-name" and "test-term". These test sets are sourced from actual internal meetings and accompanied by a specific contextual knowledge bias list. The total size of the dataset is 160,000 hours, and it is intended to support automatic speech recognition (ASR) tasks.
提供机构:
In-house搜集汇总
数据集介绍

背景与挑战
背景概述
该数据集是用于上下文语音识别研究的内部大规模数据集,包含160k小时的中文普通话和英语音频,测试集来自真实会议。数据集分为test-name和test-term两个子集,分别针对人名和术语进行上下文偏置,覆盖开放域场景,并提供了详细的统计指标如罕见比率和覆盖率,适用于ASR定制化和个性化研究。
以上内容由遇见数据集搜集并总结生成



