UniLSTalkDataset
收藏资源简介:
UniLS-Talk 数据集是一个大规模高质量3D面部运动数据集合,旨在支持统一说话与倾听虚拟人生成的研究。数据集通过精心设计的跟踪流程提取每帧的FLAME参数,包括表情系数、眼球注视、下颌姿态和头部姿态注释。数据集包含两部分互补内容:1) 来自Seamless Interaction数据集的配对对话数据,提供具有自然轮流动态的双说话者同步视频;2) 从CelebV、TalkingHead-1KH、TEDTalk、VFHQ等野外视频中聚合的非配对多场景数据,涵盖不同身份和环境下的多样化面部行为(新闻广播、访谈、随意交谈等)。数据集总计1,204小时,其中配对对话数据657.5小时(含音频和运动数据),非配对多场景数据546.5小时(仅含运动数据)。配对对话数据被划分为622.5小时训练集、4.8小时验证集和30.2小时测试集。所有数据均包含25fps的FLAME表情参数、下颌和头部姿态以及眼球注视注释。
The UniLS-Talk dataset is a large-scale, high-quality 3D facial motion dataset designed to support research on unified speaking and listening virtual avatar generation. The dataset extracts per-frame FLAME parameters via a carefully designed tracking pipeline, including annotations for expression coefficients, eye gaze, jaw pose, and head pose. The dataset consists of two complementary components: 1) Paired conversational data sourced from the Seamless Interaction dataset, which provides synchronized videos of two speakers with natural turn-taking dynamics; 2) Unpaired multi-scenario data aggregated from in-the-wild videos including CelebV, TalkingHead-1KH, TEDTalk, VFHQ and other sources, covering diverse facial behaviors across different identities and environments (e.g., news broadcasting, interviews, casual conversations). In total, the dataset spans 1,204 hours, with 657.5 hours of paired conversational data (including both audio and motion data) and 546.5 hours of unpaired multi-scenario data (containing only motion data). The paired conversational data is split into 622.5 hours for training, 4.8 hours for validation, and 30.2 hours for testing. All data includes 25fps FLAME expression parameters, jaw and head pose annotations, as well as eye gaze annotations.
UniLS-Talk 数据集概述
数据集简介
UniLS-Talk 数据集是一个用于统一说话与倾听虚拟人生成研究的大规模高质量3D面部运动数据集合。该数据集通过精心设计的追踪流程,提取了每帧的FLAME参数,包括表情系数、眼球注视、下颌姿态和头部姿态标注。
数据集构成
数据集由两个互补的部分组成:
1. 配对对话数据
- 来源:Seamless Interaction 数据集。
- 内容:提供同步的双说话者视频,包含说话与倾听之间自然的轮流动态。
- 时长:657.5小时。
- 数据模态:包含音频和运动数据。
2. 非配对多场景数据
- 来源:从CelebV、TalkingHead-1KH、TEDTalk、VFHQ以及其他野外视频中聚合。
- 内容:涵盖不同身份和环境(如新闻广播、访谈、随意交谈等)中的多样化面部行为。
- 时长:546.5小时。
- 数据模态:仅包含运动数据,不包含音频。
数据统计
| 类别 | 来源 | 时长 | 音频 | 运动 |
|---|---|---|---|---|
| 配对对话 | Seamless Interaction 数据集 | 657.5 h | ✅ | ✅ |
| 非配对多场景 | 来自野外视频的不同身份和环境 | 546.5 h | ❌ | ✅ |
| 总计 | 1,204 h |
数据划分与标注
- 配对对话数据划分:
- 训练集:622.5小时。
- 验证集:4.8小时。
- 测试集:30.2小时。
- 标注信息:所有数据均包含FLAME表情参数、下颌与头部姿态以及眼球注视标注,帧率为25 fps。
相关链接
- FLAME模型:https://flame.is.tue.mpg.de/
- Seamless Interaction数据集:https://ai.meta.com/research/seamless-interaction/




