GenIR
收藏资源简介:
GenIR数据集由中国科学院自动化研究所创建,旨在解决真实世界图像恢复中的数据不足问题。该数据集包含一百万张高质量图像,通过创新的图像-文本对构建、双提示微调及数据生成与过滤三个阶段生成。数据集的创建过程避免了传统的数据爬取方法,确保了版权合规和隐私安全。GenIR数据集主要应用于图像恢复领域,旨在提升模型对多样化和复杂化真实世界图像降质的处理能力。
The GenIR dataset was created by the Institute of Automation, Chinese Academy of Sciences, with the aim of addressing the shortage of training data for real-world image restoration. This dataset comprises one million high-quality images, which are generated through three sequential stages: innovative image-text pair construction, dual-prompt fine-tuning, and data generation and filtering. The development process of the dataset avoids traditional data crawling approaches, thus ensuring copyright compliance and privacy protection. The GenIR dataset is primarily utilized in the field of image restoration, with the objective of enhancing the model's ability to process diverse and complex real-world image degradations.
DreamClear 数据集概述
数据集名称
DreamClear
数据集描述
DreamClear 是一个用于真实世界图像恢复的高容量数据集,强调隐私安全的数据集构建。
数据集内容
- RealLQ250 基准: 包含 250 张真实世界的低质量(LQ)图像。
- 训练数据: 提供高质量(HQ)和低质量(LQ)图像对,用于图像恢复模型的训练。
数据集下载
- RealLQ250 基准: 可从 Google Drive 下载。
- 预训练模型: 可在 Huggingface 下载。
数据集使用
训练
- 准备训练数据: 生成高质量和低质量图像对。
- 提取文本特征: 使用 T5 模型提取文本特征。
- 训练 DreamClear: 使用提供的配置文件和预训练模型进行训练。
推理
- 图像恢复: 使用 DreamClear 模型将低质量图像恢复到高质量。
- 高级别基准测试: 提供分割和检测的测试指令。
数据集许可证
该数据集的代码和预训练权重基于 Apache 2.0 许可证。
引用
如果使用该数据集,请引用以下论文:
@article{ai2024dreamclear, title={DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset Curation}, author={Ai, Yuang and Zhou, Xiaoqiang and Huang, Huaibo and Han, Xiaotian and Chen, Zhengyu and You, Quanzeng and Yang, Hongxia}, journal={Advances in Neural Information Processing Systems}, year={2024} }




