MegaStyle-1.4M
收藏资源简介:
MegaStyle-1.4M是由腾讯AIPD实验室联合多所高校构建的大规模风格数据集,通过文本到图像生成模型的风格映射一致性实现。该数据集包含170K风格提示词与400K内容提示词组合生成的140万高质量图像,数据源自LAION-Aesthetics、WikiArt和JourneyDB等开放数据集。其构建过程采用Qwen3-VL模型标注图像风格特征,并通过分层聚类算法平衡提示词分布。该数据集专注于解决风格迁移任务中风格一致性与多样性的平衡问题,为艺术创作、滤镜开发等应用提供支持。
MegaStyle-1.4M is a large-scale style dataset constructed by Tencent AIPD Lab in collaboration with multiple universities, grounded in the style mapping consistency of text-to-image generation models. This dataset contains 1.4 million high-quality images generated from the combination of 170K style prompts and 400K content prompts, with source data derived from open datasets including LAION-Aesthetics, WikiArt, and JourneyDB. During the construction process, the Qwen3-VL model was utilized to annotate image style features, and a hierarchical clustering algorithm was adopted to balance the distribution of prompts. This dataset focuses on addressing the trade-off between style consistency and diversity in style transfer tasks, providing support for applications such as artistic creation and filter development.
MegaStyle数据集概述
数据集名称
MegaStyle
数据集核心特性
- MegaStyle-1.4M: 包含约140万张风格图像的大规模风格数据集。
- 风格内一致性: 共享相同风格但内容不同的风格对。
- 风格间多样性: 包含大量多样的风格。
- 高质量: 数据集图像具有高质量。
数据构建方法
- 数据构建流程: 利用大型生成模型从给定风格描述生成相同风格图像的能力。
- 提示词库: 构建了包含17万条风格提示词和40万条内容提示词的多样化、平衡的提示词库。
- 生成方式: 通过内容-风格提示词组合生成大规模风格数据集。
数据集应用
- 风格编码器: 基于MegaStyle-1.4M,通过风格监督对比学习微调出MegaStyle-Encoder,用于提取富有表现力、风格特定的表示。
- 风格迁移模型: 训练了基于FLUX的风格迁移模型MegaStyle-FLUX。
- 模型效果: MegaStyle-FLUX在生成风格化图像时,能有效捕捉颜色、光线、纹理和笔触等细微差别,并与文本提示指定的内容以及参考图像的风格保持一致。
模型性能
- 对比基准: 与DEADiff、StyleShot、Attention-Distillation (Attn-Distill)、CSGO、StyleCrafter、InstantStyle和StyleAligned等最先进的风格迁移方法进行比较。
- 性能表现: MegaStyle-FLUX相比这些基线方法实现了更优的性能。
相关资源状态
- 论文: 即将发布
- 代码: 即将发布
- 数据集: 即将发布
- 模型: 即将发布

- 1MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping同济大学; 腾讯; 南洋理工大学; 香港科技大学; 福州大学; 香港大学; 新加坡国立大学 · 2026年



