Mogao
收藏资源简介:
Mogao数据集是专门为联合文本和图像生成而构建的大规模数据集,包含了一千万条数据。该数据集旨在促进统一模型在多模态理解和生成方面的研究和应用,尤其适用于Mogao模型。数据集的构建过程采用了高效的训练策略,可以同时优化教师强制文本标记和基于扩散的视觉标记。数据集的应用领域包括图像理解、文本到图像生成、图像编辑和合成,旨在解决多模态内容生成的问题。
The Mogao Dataset is a large-scale dataset specifically developed for joint text and image generation, comprising 10 million data instances. This dataset is intended to advance research and applications of unified models in multimodal understanding and generation, and is particularly tailored for the Mogao Model. The construction of this dataset employs an efficient training strategy that simultaneously optimizes teacher-forcing text tokens and diffusion-based visual tokens. Its applicable domains include image understanding, text-to-image generation, image editing and synthesis, aiming to address the challenges in multimodal content generation.
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
基本信息
- 标题: Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- 作者: Chao Liao, Liyang Liu, Xun Wang, Zhengxiong Luo, Xinyu Zhang, Wenliang Zhao, Jie Wu, Liang Li, Zhi Tian, Weilin Huang
- 提交日期: 2025年5月8日 (v1), 2025年5月11日 (v2)
- arXiv标识符: arXiv:2505.05472v1 [cs.CV]
- DOI: 10.48550/arXiv.2505.05472
- 分类: Computer Vision and Pattern Recognition (cs.CV)
摘要
Mogao是一个统一框架,通过因果方法实现交错多模态生成。其关键改进包括:
- 深度融合设计
- 双视觉编码器
- 交错旋转位置嵌入
- 多模态无分类器引导
这些改进使Mogao能够:
- 结合自回归模型和扩散模型的优势
- 处理任意交错的文本和图像序列
- 在大规模内部数据集上高效训练
实验表明,Mogao在以下方面表现优异:
- 多模态理解
- 文本到图像生成
- 生成高质量、连贯的交错输出
- 零样本图像编辑和组合生成
技术报告
- 类型: Mogao Technical Report
- 版本: v1 (2025年5月8日), v2 (2025年5月11日)
- PDF链接: View PDF

- 1Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation字节跳动种子实验室 · 2025年



