遇见数据集

ImageAttributionBench

收藏
DataCite Commons2025-05-16 更新2025-05-17 收录
官方服务:

资源简介:

This is a benchmark dataset for image attribution,including 22 generation models and 460k images in total. This dataset is linked to NIPS 2025 benchmark submission.Submission Number: 1109 Abstract: The rapid advancement of generative AI has enabled the creation of highly realistic and diverse synthetic images, posing critical challenges for image provenance and misinformation detection. This underscores the urgent need for effective image attribution. However, existing attribution datasets are constrained by limited scale, outdated generation methods, and insufficient semantic diversity—hindering the development of robust and generalizable attribution models. To address these limitations, we introduce ImageAttributionBench, a large-scale and comprehensive dataset comprising images synthesized by a wide array of advanced generative models with state-of-the-art (SOTA) architectures. Covering multiple real-world semantic domains, the dataset offers rich diversity and scale to support and accelerate progress in image attribution research. To simulate real-world attribution scenarios, we evaluate several SOTA attribution methods on ImageAttributionBench under two challenging settings: (1) training on a standard balanced split and testing on degraded images, and (2) training and testing on semantically disjoint splits. In both cases, current methods exhibit consistently poor performance, revealing significant limitations in their robustness and generalization to unseen semantic content. Our work provides a rigorous benchmark to facilitate the development and evaluation of future image attribution methods.

本数据集为图像溯源(image attribution)基准数据集,共涵盖22种生成模型与46万张图像。 本数据集关联于NIPS 2025基准数据集投稿,投稿编号:1109。 摘要: 生成式AI(Generative AI)的快速发展使得能够生成高度逼真且多样化的合成图像,这为图像溯源与虚假信息检测带来了严峻挑战,凸显了对高效图像溯源技术的迫切需求。然而,现有溯源数据集存在规模有限、生成方法过时、语义多样性不足等局限,制约了鲁棒且可泛化的溯源模型的研发。为解决上述局限,我们提出ImageAttributionBench——一款大规模且全面的基准数据集,其包含由多种搭载当前最优(SOTA)架构的先进生成式AI模型合成的图像。该数据集覆盖多个真实世界语义领域,具备丰富的多样性与规模,可为图像溯源研究的发展与提速提供支撑。为模拟真实溯源场景,我们在ImageAttributionBench上针对两种极具挑战性的设置评估了多款SOTA溯源方法:(1)在标准平衡划分集上训练,在降质图像上测试;(2)在语义不相交的划分集上分别进行训练与测试。在两种设置下,当前方法均始终表现不佳,暴露出其在鲁棒性以及对未知语义内容的泛化能力上存在显著局限。本研究提供了一套严谨的基准数据集,以助力未来图像溯源方法的研发与评估。

提供机构:
Harvard Dataverse
创建时间:
2025-05-01
搜集汇总
数据集介绍
ImageAttributionBench 数据集图片
背景与挑战
背景概述
ImageAttributionBench是一个大规模图像归因基准数据集,包含22个先进生成模型生成的46万张图像,覆盖多个真实世界语义领域,旨在解决现有数据集的规模、方法过时和多样性不足等问题。该数据集通过平衡分割和语义不相交分割等挑战性设置评估方法性能,揭示了当前图像归因方法在鲁棒性和泛化性方面的显著局限,为未来研究提供了严谨的基准。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务