OmniMMI
收藏资源简介:
OmniMMI是一个专为Omni语言模型设计的全面的多模态交互基准,适用于流媒体视频环境。该数据集包含超过1,121个视频和2,290个问题,旨在解决现有视频基准中未被充分探索的两个关键挑战:流媒体视频理解和主动推理。数据集涵盖了六种不同的子任务,包含动作预测、动态状态定位、多轮依赖推理等。数据来源于YouTube和开源视频音频数据,经过人工审核和标注,适用于评估多模态语言模型在流媒体视频环境中的交互能力。
OmniMMI is a comprehensive multimodal interaction benchmark specifically designed for Omni language models, tailored for streaming video environments. The dataset consists of over 1,121 videos and 2,290 questions, aiming to address two critical challenges that remain under-explored in existing video benchmarks: streaming video understanding and proactive reasoning. It covers six distinct subtasks, including action prediction, dynamic state localization, multi-turn dependency reasoning, and more. The data is sourced from YouTube and open-source video-audio datasets, undergoes manual review and annotation, and is intended to evaluate the interactive capabilities of multimodal language models in streaming video scenarios.




