AVE-PM
收藏资源简介:
AVE-PM是一个专门为 portrait 模式短视频设计的音频事件定位数据集,包含25335个10秒视频剪辑,涵盖86个细粒度类别,具有帧级注释。该数据集由抖音平台上的用户生成内容构成,反映了不受约束的用户生成内容的真实情况。数据集的构建过程包括从抖音平台收集原始视频,通过众包方式进行注释,并最终切分成10秒的剪辑。该数据集旨在推动移动-centric视频内容时代的音频事件定位研究。
AVE-PM is an audio event localization dataset specifically designed for portrait-mode short videos. It includes 25,335 10-second video clips spanning 86 fine-grained categories, with frame-level annotations. This dataset is sourced from user-generated content on the Douyin (TikTok) platform, reflecting the authentic real-world scenarios of unconstrained user-generated content. The construction pipeline of this dataset involves collecting raw videos from the Douyin platform, performing annotations via crowdsourcing, and finally splitting the collected content into 10-second video clips. This dataset aims to advance audio event localization research in the era of mobile-centric video content.

- 1Audio-visual Event Localization on Portrait Mode Short Videos武汉大学 · 2025年



