TTS-AGI/podcast-dramabox-dacvae-pairs
收藏资源简介:
该数据集名为podcast-dramabox-dacvae-pairs,是一个用于训练DramaBox到DACVAE潜变量翻译器的配对音频编解码器潜变量数据集。数据集中的潜变量共享相同的网格:25 Hz频率、128维、帧对齐(长度相同)。数据来源于TTS-AGI/podcast-tokenized-bg3.5-enj5数据集。构建过程包括从DACVAE潜变量解码生成48 kHz单声道音频,复制为立体声后通过DramaBox/LTX-2.3音频VAE编码并分块,生成DramaBox潜变量作为输入,同时将DACVAE潜变量作为目标,两者长度修剪至最小(差异不超过1帧)。数据集使用.npz格式存储,包含DramaBox潜变量(输入)、DACVAE潜变量(目标)、样本键和元数据(如转录文本和情感分数)。
The dataset named podcast-dramabox-dacvae-pairs is a paired audio-codec latent dataset for training a DramaBox to DACVAE latent translator. Both codecs share an identical grid: 25 Hz, 128-dimensional, frame-aligned (same length). It is derived from the TTS-AGI/podcast-tokenized-bg3.5-enj5 dataset. The construction process involves decoding DACVAE latents to generate 48 kHz mono audio, duplicating to stereo, encoding via DramaBox/LTX-2.3 audio VAE with patchification to produce DramaBox latents as input, while DACVAE latents serve as the target, with both trimmed to the minimum length (differing by ≤1 frame). The dataset is stored in .npz format, containing DramaBox latents (input), DACVAE latents (target), sample keys, and metadata (e.g., transcript and emotion scores).




