FiVA
收藏资源简介:
FiVA是一个细粒度的视觉属性数据集,专为文本到图像扩散模型设计。该数据集旨在从源图像中解耦不同的视觉属性,并在文本到图像生成过程中进行适应。
FiVA is a fine-grained visual attribute dataset specifically designed for text-to-image diffusion models. This dataset aims to decouple various visual attributes from source images and facilitate adaptation during the text-to-image generation process.
FiVA: Fine-grained Visual Attributes for T2I Models
简介
我们构建了一个细粒度的视觉属性数据集和一个框架,该框架能够从源图像中解耦不同的视觉属性,并在文本到图像生成过程中适应这些属性。
示例
我们的模型可以从多个参考图像中整合不同的属性 V(image, attr_name),并将它们集成到目标主体 T(subject) 中,同时还能根据不同的属性名称从同一参考图像中提取各种视觉属性。
发布
🚀 我们的代码和预训练模型将于2023年12月中旬发布。
引用
如果您发现我们的数据集或模型对您的研究和应用有用,请使用以下BibTeX引用: bibtex @inproceedings{wu2024fiva, title={Fi{VA}: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models}, author={Tong Wu and Yinghao Xu and Ryan Po and Mengchen Zhang and Guandao Yang and Jiaqi Wang and Ziwei Liu and Dahua Lin and Gordon Wetzstein}, booktitle={The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track}, year={2024}, url={https://openreview.net/forum?id=Vp6HAjrdIg} }




