S2Cap
收藏资源简介:
S2Cap是一个用于歌唱风格描述的音频-文本对数据集,由POSTECH和HJ AILAB创建。该数据集包含12,105个音乐音频样本和71,215个描述性文本,涵盖了音高、音量、节奏、情绪、歌手性别和年龄、音乐类型和情感表达等多种声乐和音乐属性。数据集的创建过程包括从Melon播放列表数据集和网络抓取中获取元数据,使用Demucs模型分离声乐部分,并通过Qwen2 Audio生成描述性文本。S2Cap数据集的应用领域主要是歌唱风格描述和语音生成,旨在解决现有数据集在音乐特征捕捉方面的不足。
S2Cap is an audio-text paired dataset dedicated to singing style description, developed by POSTECH and HJ AILAB. This dataset comprises 12,105 musical audio samples and 71,215 descriptive texts, covering a broad spectrum of vocal and musical attributes including pitch, volume, rhythm, emotion, singer's gender and age, music genre, and emotional expression. The development workflow of the S2Cap dataset includes collecting metadata from the Melon playlist dataset and web scraping, isolating vocal segments using the Demucs model, and generating descriptive texts via Qwen2 Audio. The primary application domains of the S2Cap dataset are singing style description and speech generation, with the goal of addressing the shortcomings of existing datasets in capturing musical features.
S2cap 数据集概述
数据集名称
S2cap
数据集描述
S2cap 数据集用于构建歌唱风格描述数据集。该数据集包含歌唱风格的相关数据和生成提示。
数据集状态
数据集和生成提示已可用,但详细和重构后的代码将在 ICASSP 2025 评审后更新。
引用信息
使用该数据集时,请引用相关论文。

- 1Constructing a Singing Style Caption DatasetPOSTECH · 2024年



