Structured Image Dataset
收藏资源简介:
本文介绍了一个用于结构化图像生成和编辑的数据集,包括一个全面的基准测试,一个具有思维链标注的大规模训练语料库和一个强大的统一模型。数据集由130万个高质量的结构化图像对组成,这些图像对来源于可执行的绘图程序,并辅以思维链推理标注。数据集的创建过程包括从可执行绘图程序中收集数百万个程序,将它们渲染成种子图像,然后在代码层面进行编辑以构建配对的代码-编辑示例,最后将这些示例渲染成图像-编辑对。该数据集旨在解决现代视觉生成模型在创建或编辑结构化视觉(如图表、图形和数学图形)方面的挑战,这些模型需要布局规划、文本渲染和多模态推理以确保事实准确性。
This paper presents a dataset for structured image generation and editing, which includes a comprehensive benchmark, a large-scale training corpus with chain-of-thought annotations, and a powerful unified model. The dataset consists of 1.3 million high-quality structured image pairs derived from executable drawing programs, supplemented with chain-of-thought reasoning annotations. The dataset creation process involves collecting millions of programs from executable drawing programs, rendering them into seed images, editing at the code level to construct paired code-editing examples, and finally rendering these examples into image-editing pairs. This dataset aims to address the challenges faced by modern visual generative models when creating or editing structured visual content such as charts, diagrams, and mathematical graphics, which require layout planning, text rendering, and multimodal reasoning to ensure factual accuracy.




