官方服务:
资源简介:
Audio Visual Signal Processing Dataset
应用场景:
创建时间:
2026-01-15
相关数据集
hillenbrand_vowels
Hillenbrand Vowel 数据集包含了四种人群(男性、女性、男孩和女孩)产生的美式英语元音录音。每个音频样本都附带有每10毫秒提取一次的帧级别格式跟踪(F1, F2, F3, F4)。该数据集以与Hugging Face datasets库兼容的结构提供,便于加载和处理。
Hugging Face2025-11-17 更新670
MatrixSpeechAI/palki_sharma_stage_3
该数据集包含语音相关的特征数据,主要特征包括文本、语音音高均值、语音音高标准差、信噪比、C50、语速、音素、语音传输指数、SI-SDR、PESQ、噪声、混响、语音单调性、噪声SDR、语音质量PESQ和文本描述等。数据集仅包含一个训练集分割,包含1734个样本,总大小为1395587字节。
Hugging Face2024-12-09 更新180
CSV files of all paper data
The data provided below should be enough to reproduce the primary data figures in the paper. They are organized in the order in which they appear in the paper. Each CSV file contains 11 columns: -
DataONE2015-12-21 更新90
Beep detection dataset
这是一组包含各种频率、持续时间和失真的蜂鸣声(语音信箱音)音频文件。该数据集还包括一些包含静音或无蜂鸣声的语音文件。这些文件是为训练应答机检测模型而生成的,因为作者无法找到足够多的现成蜂鸣声录音。
github2024-08-08 更新660
Synthetic noise dataset
The synthetic noise dataset is divided into 3 subsets: 80,000 noise tracks for training, 2,000 noise tracks for validation, and remaining 2,000 noise tracks for testing. The synthetic noise tracks are
DataCite Commons2025-06-10 更新130



