OMNIPRO
收藏资源简介:
OMNIPRO是由中国人民大学与腾讯微信视觉团队联合创建的首个全主动流式视频理解综合基准数据集。该数据集包含2700个人工验证样本,涵盖9个子任务和3个认知层次,全面覆盖了6种基本视频理解能力,其中84%的样本依赖音频信号(语音或非语音)。数据来源于LongVALE和COIN两个公开数据集的测试集,共计1771个源视频,通过Gemini 3 Flash模型生成多模态密集描述与结构化问答对。该数据集旨在系统评估全模态感知、主动响应决策与多样化视频理解任务的协同能力,为流式视频理解模型提供统一的评测框架。
OMNIPRO is the first comprehensive benchmark dataset for fully active streaming video understanding, jointly developed by Renmin University of China and the Tencent WeChat Vision Team. It comprises 2700 manually validated samples, spanning 9 subtasks and 3 cognitive levels, and comprehensively covers 6 fundamental video understanding capabilities. Notably, 84% of these samples rely on audio signals, including both speech and non-speech content. The dataset is sourced from the test splits of two public datasets, LongVALE and COIN, with a total of 1771 source videos. Multimodal dense descriptions and structured question-answering pairs were generated using the Gemini 3 Flash model. This benchmark aims to systematically evaluate the collaborative capabilities of full-modal perception, active response decision-making, and diverse video understanding tasks, thereby providing a unified evaluation framework for streaming video understanding models.




