AudioSet
收藏资源简介:
Audioset 是一个音频事件数据集,由超过 200 万个人工注释的 10 秒视频片段组成。这些剪辑是从 YouTube 收集的,因此其中许多质量很差,并且包含多个声源。使用 632 个事件类的分层本体来注释这些数据,这意味着可以将相同的声音注释为不同的标签。例如,吠叫的声音被注释为 Animal、Pets 和 Dog。所有视频都分为评估/平衡训练/不平衡训练集。
Audioset is an audio event dataset composed of over 2 million manually annotated 10-second video clips. These clips are harvested from YouTube, resulting in many of them having subpar audio quality and containing multiple concurrent sound sources. The data are annotated against a hierarchical ontology of 632 event classes, which permits the same sound to be assigned multiple distinct labels. For instance, the sound of a barking dog is annotated with the labels Animal, Pets, and Dog. All video clips are partitioned into three subsets: the evaluation set, the balanced training set, and the unbalanced training set.

- AudioSet首次发表,由Google AI团队发布,包含约200万个音频片段,涵盖527个声音事件类别。
- AudioSet被广泛应用于音频事件检测和分类任务,成为音频处理领域的重要基准数据集。
- AudioSet的扩展版本发布,增加了更多的音频样本和类别,进一步丰富了数据集的内容。
- AudioSet开始应用于多模态学习研究,特别是在音频与视频数据的联合分析中展现出其独特价值。
- AudioSet的标注质量得到进一步提升,引入了更精细的标签体系,提高了数据集在复杂场景下的应用效果。
- AudioSet被用于开发新一代的音频识别模型,推动了音频技术在智能家居、自动驾驶等领域的应用。
- 1AudioSet: An ontology and human-labeled dataset for audio eventsGoogle · 2017年
- 2Weakly-Supervised Sound Event Detection Using Audiovisual CorrespondenceUniversity of Surrey · 2020年
- 3Sound Event Detection Using Weakly Labeled Data with AudioSetUniversity of Rochester · 2019年
- 4Audio-Visual Scene Analysis with Self-Supervised Multisensory FeaturesUniversity of Oxford · 2018年
- 5Learning to Recognize Sounds with Weak SupervisionUniversity of California, Berkeley · 2021年



