Short-Clips, Full-Videos
收藏资源简介:
本数据集由纽约州立大学石溪分校计算机科学系与Atmanity Inc.合作构建,旨在帮助对话式AI理解何时以及如何回应。数据集包含真实对话视频中的视觉、听觉和文本流,共分为Short-Clips和Full-Videos两个子集。Short-Clips子集包含4,393个反应片段、2,000个完整回应片段和2,000个沉默片段,用于训练和测试模型对孤立短片段中适当声音反应的预测能力。Full-Videos子集则用于评估模型在连续对话中判断何时说话的能力。这些数据集的创建有助于提升对话式AI的实时响应能力和自然度。
This dataset was co-developed by the Department of Computer Science, Stony Brook University, State University of New York and Atmanity Inc., with the goal of helping conversational AI understand when and how to generate appropriate responses. The dataset contains visual, auditory and textual streams from real conversational videos, and is split into two subsets: Short-Clips and Full-Videos. The Short-Clips subset includes 4,393 reaction clips, 2,000 complete response clips and 2,000 silent clips, which are utilized to train and test models' ability to predict proper vocal responses in isolated short segments. The Full-Videos subset is intended to evaluate models' capacity to judge when to speak during continuous conversations. The creation of this dataset helps enhance the real-time response performance and naturalness of conversational AI.
数据集概述
基本信息
- 数据集名称: Beyond Words: Multimodal LLM Knows When to Speak
- 数据集状态: 待发布(Codes and datasets to be released in the future)
相关研究
- 关联论文: "Beyond Words: Multimodal LLM Knows When to Speak"
备注
- 该数据集目前尚未发布,具体内容和发布时间需关注后续更新。




