sample-dataset-test-neural-nopack
收藏资源简介:
该数据集是一个专为语音处理任务设计的音频数据集,特别适用于Whisper模型。数据集由Trelis Studio准备,包含经过语音活动检测(VAD)处理的音频片段,去除了静音部分。数据集统计信息显示,共有6个源文件,包含184个训练样本和20个验证样本,总时长为62.6分钟。数据集的列包括音频片段(16kHz)、纯文本转录、带时间戳的转录、片段起始和结束时间、语音持续时间、词级时间戳以及源文件名。音频片段经过Silero VAD处理,确保仅保留语音区域,时间戳相对于拼接后的语音音频。数据集适用于Whisper时间戳训练,建议使用两桶方法:50%使用纯文本转录,50%使用带时间戳的转录。
This dataset is an audio dataset tailored for speech processing tasks, particularly suitable for the Whisper model. Prepared by Trelis Studio, it includes audio segments processed via Voice Activity Detection (VAD) with silent parts removed. Statistical summary of the dataset shows 6 source files, 184 training samples, 20 validation samples, with a total duration of 62.6 minutes. The dataset contains the following columns: 16kHz audio segments, raw text transcripts, timestamped transcripts, segment start and end times, speech duration, word-level timestamps, and source file names. All audio segments are processed with Silero VAD to retain only speech regions, and the timestamps are relative to the concatenated speech audio. This dataset is intended for Whisper timestamp training, and the two-bucket approach is recommended: 50% of the data uses raw text transcripts, while the other 50% uses timestamped transcripts.




