CN-Celeb-AV
收藏资源简介:
CN-Celeb-AV是一个多模态音频-视觉数据集,旨在评估无约束条件下的人识别技术。该数据集由清华大学语音和语言技术中心及北京邮电大学人工智能学院共同创建,包含超过419,663个视频片段,涵盖1,136位来自公共媒体的个体,涉及11种不同的真实世界场景类型。数据集的创建过程中,特别强调了多类型数据和部分信息片段的处理,以更真实地模拟现实复杂性。CN-Celeb-AV适用于研究音频-视觉人识别技术,特别是在无约束环境下的性能评估,为该领域的研究提供了一个新的基准数据集。
CN-Celeb-AV is a multimodal audio-visual dataset designed to evaluate unconstrained person recognition technologies. It was jointly developed by the Center for Speech and Language Technology at Tsinghua University and the School of Artificial Intelligence at Beijing University of Posts and Telecommunications. The dataset contains over 419,663 video clips, covering 1,136 individuals from public media and spanning 11 distinct real-world scenario types. Special emphasis was placed on the processing of multi-type data and partial information segments during dataset construction to more realistically simulate the complexity of real-world environments. CN-Celeb-AV is applicable to research on audio-visual person recognition technologies, particularly performance evaluation in unconstrained environments, providing a new benchmark dataset for relevant research in this field.

- 1CN-Celeb-AV: A Multi-Genre Audio-Visual Dataset for Person Recognition清华大学语音和语言技术中心,北京邮电大学人工智能学院 · 2023年



