sample-dataset-test-neural-pack
收藏资源简介:
该数据集名为'sample-dataset-test-neural-pack',是一个专注于音频和语音处理的数据集,特别适用于whisper模型。数据集使用Trelis Studio准备,包含6个源文件,136个训练样本和15个验证样本,总时长为62.6分钟。数据集的主要字段包括:音频段(16kHz,仅语音,通过VAD去除静音)、纯文本转录(无时间戳)、带Whisper时间戳标记的转录、段落在原始音频中的开始和结束时间、语音持续时间(不包括静音)、词级时间戳(相对于仅语音的音频)以及原始音频文件名。数据集经过Silero VAD处理,确保训练数据与推理行为匹配。对于Whisper时间戳训练,建议使用两桶方法:50%使用纯文本转录,50%使用带时间戳标记的转录。
This dataset is named 'sample-dataset-test-neural-pack', a dataset focused on audio and speech processing specifically tailored for Whisper models. Prepared using Trelis Studio, it contains 6 source files, 136 training samples and 15 validation samples, with a total duration of 62.6 minutes. The core fields of the dataset include: audio segments (16kHz, speech-only, with silence removed via VAD), plain text transcriptions (without timestamps), transcriptions with Whisper timestamps, start and end timestamps of speech segments in the original audio, speech duration (excluding silence), word-level timestamps (relative to the speech-only audio), and the original audio filename. The dataset has been processed with Silero VAD to align the training data with inference behavior. For Whisper timestamp training, the two-bucket approach is recommended: 50% of the samples use plain text transcriptions, while the remaining 50% use transcriptions with timestamps.




