KRIS-Bench
收藏资源简介:
KRIS-Bench是一个基于知识推理的图像编辑系统基准测试,旨在评估模型的知识推理能力。该数据集由东南大学、马克斯·普朗克计算机科学研究所等机构的研究人员共同创建,包含1,267个高质量的标注编辑实例,覆盖了22个编辑任务,跨越7个推理维度。数据集通过认知教育原理进行设计,将知识分为事实性、概念性和程序性三种类型,并提供了详细的分类体系,以支持更精细的评估。KRIS-Bench还引入了新的评估维度——知识合理性,以评估模型生成的编辑是否与现实世界的知识一致,并通过用户研究验证了评估协议的有效性。该数据集适用于图像编辑模型的研究和开发,旨在解决图像编辑中的知识推理问题。
KRIS-Bench is a knowledge reasoning-based benchmark for image editing systems, designed to evaluate the knowledge reasoning capabilities of models. This dataset was jointly created by researchers from institutions including Southeast University, Max Planck Institute for Computer Science, and other institutions. It comprises 1,267 high-quality annotated editing instances, covering 22 editing tasks and spanning 7 reasoning dimensions. The dataset is designed based on cognitive education principles, classifying knowledge into three categories: factual, conceptual, and procedural, and providing a detailed classification system to support more fine-grained evaluations. KRIS-Bench also introduces a new evaluation dimension—knowledge plausibility—to assess whether the edits generated by models align with real-world knowledge, and verified the effectiveness of its evaluation protocol through user studies. This dataset is suitable for the research and development of image editing models, aiming to address the knowledge reasoning challenges in image editing.
KRIS-Bench 数据集概述
数据集基本信息
- 名称: KRIS-Bench (Knowledge-based Reasoning in Image-editing Systems Benchmark)
- 开发者: Yongliang Wu等 (来自东南大学、马克斯·普朗克信息学研究所、上海交通大学等机构)
- 对应作者: Xinting Hu (†)
- 项目负责人: Xianfang Zeng (‡)
数据集简介
- 目的: 评估基于指令的图像编辑模型在知识推理任务上的表现
- 理论基础: 基于教育理论,将编辑任务分为三类知识类型:
- 事实性知识 (Factual)
- 概念性知识 (Conceptual)
- 程序性知识 (Procedural)
- 任务设计:
- 22个代表性任务
- 覆盖7个推理维度
- 包含1,267个高质量标注的编辑实例
评估方法
- 核心指标: 知识合理性 (Knowledge Plausibility)
- 评估增强:
- 使用知识提示
- 通过人类研究校准
性能排行榜
- 评估维度:
- 事实性知识 (包含属性感知、空间感知、时间感知)
- 概念性知识 (包含社会科学、自然科学)
- 程序性知识 (包含逻辑推理、指令分解)
| 排名 | 模型 | 事实性知识 | 概念性知识 | 程序性知识 | 综合得分 |
|---|---|---|---|---|---|
| 1 | GPT-4o OpenAI | 79.80 | 81.37 | 78.32 | 80.09 |
| 2 | Gemini 2.0 Google | 65.26 | 59.65 | 62.90 | 62.41 |
| 3 | Doubao ByteDance | 63.30 | 62.23 | 54.17 | 60.70 |
| 4 | BAGEL-Think ByteDance | 55.77 | 59.44 | 39.26 | 53.36 |
| 5 | BAGEL ByteDance | 47.71 | 52.17 | 40.23 | 47.76 |
| 6 | Step1X-Edit StepFun | 45.52 | 48.01 | 31.82 | 43.29 |
| 7 | Emu2 BAAI | 45.40 | 37.54 | 34.91 | 39.70 |
| 8 | AnyEdit ZJU | 39.26 | 41.88 | 31.74 | 38.55 |
| 9 | MagicBrush OSU | 41.84 | 39.24 | 26.54 | 37.15 |
| 10 | OmniGen BAAI | 33.11 | 28.02 | 23.89 | 28.85 |
| 11 | InstructPix2Pix UCB | 23.33 | 25.59 | 17.28 | 22.82 |
联系方式
- 结果提交: yongliang0223@gmail.com

- 1KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models东南大学, 马克斯·普朗克计算机科学研究所, 上海交通大学, StepFun, 加利福尼亚大学伯克利分校, 加利福尼亚大学默塞德分校 · 2025年



