SOK-Bench
收藏资源简介:
SOK-Bench是一个由香港大学等机构创建的全新视频推理基准数据集,包含44,000个问题和10,000个视频片段,旨在评估模型在动态、开放世界和结构化知识背景下的推理能力。数据集通过结合大型语言模型和多模态大型语言模型自动生成问题-答案对、知识图谱和推理过程,确保了数据的高质量和多样性。该数据集特别适用于评估模型在理解和应用场景知识及通用知识解决问题的能力,为人工智能领域提供了一个重要的研究和测试平台。
SOK-Bench is a novel video reasoning benchmark dataset developed by institutions including the University of Hong Kong. It comprises 44,000 questions and 10,000 video clips, and is designed to evaluate the reasoning capabilities of models under dynamic, open-world, and structured knowledge contexts. The dataset automatically generates question-answer pairs, knowledge graphs, and reasoning processes by combining large language models (LLMs) and multimodal large language models (MLLMs), thus ensuring high data quality and diversity. This benchmark is particularly suitable for assessing models' abilities to comprehend and apply both scenario-specific knowledge and general knowledge to solve problems, serving as a critical research and testing platform for the field of artificial intelligence.



