遇见数据集

UniTalk

收藏
arXiv2025-09-30 收录
官方服务:

资源简介:

该数据集名为UniTalk,专为活跃说话者检测而设计,特别关注包含较少代表性语言、噪声背景以及拥挤场景等具有挑战性的真实世界情境。该数据集包含超过44.5小时的视频,并在48,693个说话者身份的帧级别提供了活跃说话者的标注。UniTalk为活跃说话者检测设立了新的基准,为在真实条件下开发和评估模型提供了宝贵的资源。该数据集规模超过44.5小时的视频,所涉及的任务是活跃说话者检测(Active Speaker Detection,简称Asd)。

This dataset, named UniTalk, is specifically designed for active speaker detection, with a particular focus on challenging real-world scenarios including underrepresented languages, noisy backgrounds, and crowded environments. The dataset contains over 44.5 hours of video, and provides frame-level annotations of active speakers for 48,693 distinct speaker identities. UniTalk sets a new benchmark for active speaker detection, serving as a valuable resource for developing and evaluating models under real-world conditions. This dataset covers over 44.5 hours of video, and the target task is Active Speaker Detection (abbreviated as Asd).

提供机构:
Hugging Face
搜集汇总
数据集介绍
UniTalk 数据集图片
背景与挑战
背景概述
UniTalk-ASD是一个用于活动说话人检测的数据集,规模在100万到1000万条数据之间,包含视频、音频和CSV文件。数据集分为训练集和验证集,CSV文件提供视频ID、时间戳、人脸边界框坐标、说话状态标签(如SPEAKING_AUDIBLE或NOT_SPEAKING)等信息,适用于多模态说话人检测任务。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务