TimeChat-Online-139K
收藏资源简介:
TimeChat-Online-139K是一个为流式视频问答(Streaming VideoQA)任务设计的综合数据集,包含多样的交互模式,包括回溯、当前感知和未来响应等场景。该数据集由平均长度为11.1分钟的长视频组成,并使用GPT-4o对视频进行标注,形成包含各种视频问答对的数据集。数据集的创建旨在解决流式视频问答中存在的挑战,如长视频的高冗余问题,以及实时交互的需求。通过引入差异令牌丢弃(DTD)机制,TimeChat-Online-139K能够有效减少视频令牌的数量,提高视频问答的效率。该数据集的创建和应用对于未来视频语言模型的开发具有重要意义。
TimeChat-Online-139K is a comprehensive dataset designed for the Streaming Video Question Answering (Streaming VideoQA) task, featuring diverse interaction scenarios including backward browsing, current perception, and future response. It consists of long-form videos with an average duration of 11.1 minutes, and uses GPT-4o to annotate the videos, forming a dataset containing various video question-answer pairs. This dataset was created to address the core challenges in Streaming VideoQA, such as the high redundancy issue in long videos and the demand for real-time interaction. By introducing the Differentiable Token Dropping (DTD) mechanism, TimeChat-Online-139K can effectively reduce the number of video tokens and improve the efficiency of video QA. The development and application of this dataset hold significant importance for the future development of video-language models.

- 1TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos北京大学, 华南理工大学, 香港大学, 快手科技 · 2025年



