遇见数据集

Skywork/Nano-banana-150k

收藏
Hugging Face2026-02-04 更新2026-02-07 收录
官方服务:

资源简介:

--- license: apache-2.0 task_categories: - image-to-image - text-to-image language: - en tags: - image-composition - multi-image - image-fusion - image-editing - unipic - 1-image-input pretty_name: Nano-banana-150k size_categories: - 10K<n<100K configs: - config_name: default data_files: - split: train path: "*.jsonl" --- # Nano-banana-150k: A Large-Scale Image Editing Instruction Dataset ## 📖 Overview Nano-banana-150k is an open-source image editing instruction dataset. It has been used in UniPic3 specifically for single-image editing training, serving as one of the core data sources for instruction-following image editing capabilities. This dataset is derived from the original Nano Banana dataset. We sincerely thank the original authors for making this valuable dataset publicly available to the community. 📎 Original dataset: https://huggingface.co/datasets/bitmind/Nano-banana-150k ## 🎨 Demo: Multi-Image Composition This dataset supports multi-image composition tasks, enabling models to combine and edit multiple input images based on natural language instructions. The following demo showcases the capabilities of **UniPic3** trained on this dataset: ![UniPic3 Multi-Image Composition Demo](unipic3_demo.png) *Example of multi-image composition and editing using UniPic3 trained on Nano-banana-150k dataset* ## 🎯 Key Features - **Large Scale**: 123,268 carefully curated image editing samples - **Diverse Tasks**: Covers multiple image editing scenarios including temporal transformation, background editing, action/gesture changes, hairstyle modification, and artistic portraits - **High Quality**: Detailed, natural language instructions with precise editing descriptions - **Production Ready**: Used in UniPic3 for real-world image editing applications - **Well Structured**: Clean JSONL format for easy integration with training pipelines ## 📊 Dataset Statistics | Task Type | Count | Description | |-----------|-------|-------------| | **Times-Change** | 18,178 | Temporal style transformation across different eras (1905, 1980, 2000, 2024, anime) | | **Background** | 32,765 | Background replacement and scene transformation | | **Action** | 22,605 | Pose and gesture modification (e.g., peace sign, thumbs up, waving) | | **Black Headshot** | 17,700 | Artistic black-and-white portrait generation | | **Hairstyle** | 16,012 | Hairstyle transformation and modification | | **Sweet Headshot** | 16,008 | Soft, warm-toned portrait generation | | **Total** | **123,268** | All image editing samples | ## 📁 Dataset Structure ### Files - `Nano-150k.jsonl`: Main dataset file containing all 123,268 samples in JSONL format (73.8 MB) - `Nano-150k.tar.gz`: Compressed archive containing all images referenced in the dataset (10.6 GB) ### Data Format Each line in `Nano-150k.jsonl` is a JSON object with the following structure: ```json { "task_type": "Times-Change" | "ic", "instruction": "Detailed natural language instruction for image editing...", "input_images": ["Nano-150k/Image/orignal/people/example.jpg"], "output_image": "Nano-150k/Image/output/time-change/1905/example.jpg", "category": "1905" // Optional, present for Times-Change tasks } ``` ### Field Descriptions - **`task_type`**: Type of editing task - `"Times-Change"`: Temporal style transformation tasks - `"ic"`: Image composition tasks (background, action, hairstyle, headshots) - **`instruction`**: Detailed natural language description of the desired image edit, including: - Specific visual changes to make - Style and aesthetic requirements - Background and scene modifications - Pose and gesture instructions - Lighting and composition details - **`input_images`**: List of input image paths (relative to dataset root) - **`output_image`**: Path to the expected output image (relative to dataset root) - **`category`**: (Optional) Category label, primarily used for Times-Change tasks to indicate the target era/style ### Directory Structure ``` Nano-150k/ ├── Nano-150k.jsonl # Main dataset file (73.8 MB) ├── Nano-150k.tar.gz # Compressed image archive (10.6 GB) └── Image/ ├── orignal/ │ └── people/ # Original input images └── output/ ├── time-change/ # Temporal transformation outputs │ ├── 1905/ │ ├── 1980/ │ ├── 2000/ │ ├── 2024/ │ └── anime/ ├── background/ # Background replacement outputs ├── action/ # Action/gesture modification outputs ├── hairstyle/ # Hairstyle transformation outputs ├── black_headshot/ # Black-and-white portrait outputs └── sweet_headshot/ # Soft portrait outputs ``` ## 🚀 Usage ### Loading the Dataset #### Using Hugging Face Datasets ```python from datasets import load_dataset # Load the dataset from Hugging Face dataset = load_dataset("Skywork/Nano-banana-150k", split="train") # Access a sample sample = dataset[0] print(sample["instruction"]) print(sample["input_images"]) print(sample["output_image"]) ``` #### Direct JSONL Loading ```python import json # Load from local file samples = [] with open("Nano-150k.jsonl", "r", encoding="utf-8") as f: for line in f: sample = json.loads(line.strip()) samples.append(sample) # Filter by task type times_change_samples = [s for s in samples if s["task_type"] == "Times-Change"] ic_samples = [s for s in samples if s["task_type"] == "ic"] ``` #### Using PyTorch DataLoader ```python from torch.utils.data import Dataset, DataLoader import json class Nano150kDataset(Dataset): def __init__(self, jsonl_path, image_root): self.samples = [] with open(jsonl_path, "r", encoding="utf-8") as f: for line in f: self.samples.append(json.loads(line.strip())) self.image_root = image_root def __len__(self): return len(self.samples) def __getitem__(self, idx): sample = self.samples[idx] # Load images, process instructions, etc. return sample dataset = Nano150kDataset("Nano-150k.jsonl", "Nano-150k/Image") dataloader = DataLoader(dataset, batch_size=32, shuffle=True) ``` ### Extracting Images **Note**: The `Nano-150k.tar.gz` file contains all the images referenced in the dataset. Whether you need to extract it depends on your use case: #### Option 1: Using Hugging Face Datasets (Recommended) When using the Hugging Face `datasets` library, you typically **don't need to manually extract** the tar file. The library can work with the JSONL file and image paths directly. Images will be loaded on-demand when accessed. ```python from datasets import load_dataset # Load dataset - images are accessed on-demand dataset = load_dataset("Skywork/Nano-banana-150k", split="train") ``` #### Option 2: Direct File Access If you need direct file system access to the images (e.g., for custom data loaders or batch processing), you'll need to extract the archive: ```bash # Extract the image archive (requires ~11 GB free space) tar -xzf Nano-150k.tar.gz # The images will be extracted to the Image/ directory structure: # Image/ # ├── orignal/people/... # └── output/... ``` **Storage Requirements**: - Compressed: 10.6 GB (tar.gz file) - Extracted: ~11 GB (estimated, actual size may vary) ### Filtering by Task Type ```python # Filter Times-Change samples by category def filter_by_category(samples, category): return [s for s in samples if s.get("category") == category] # Get all 1905 style transformations samples_1905 = filter_by_category(samples, "1905") # Get all 1980s style transformations samples_1980 = filter_by_category(samples, "1980") ``` ## 🔬 Task Categories ### 1. Times-Change (Temporal Style Transformation) Transforms subjects across different historical eras and styles: - **1905**: Edwardian/Victorian era styling - **1980**: 1980s retro aesthetic with neon colors and period fashion - **2000**: Early 2000s Y2K millennial style - **2024**: Modern contemporary fashion and settings - **Anime**: Anime-style transformations Example instruction: > "Transform the woman in the original portrait into a modern 1905–style sitter: keep her facial features and makeup tone but replace the gold veil and heavy South Asian jewelry with a high‑collared dark green velvet Edwardian dress trimmed with black lace..." ### 2. Background Replacement Replaces or modifies backgrounds while preserving the main subject. ### 3. Action/Gesture Modification Changes poses and gestures (e.g., peace sign, thumbs up, waving, crossing arms). ### 4. Hairstyle Transformation Modifies hairstyles while maintaining facial features and overall composition. ### 5. Black Headshot Generates artistic black-and-white portraits with various lighting and composition styles. ### 6. Sweet Headshot Creates soft, warm-toned portraits with gentle lighting and natural aesthetics. ## 🎓 Applications This dataset is designed for training and evaluating: - **Image-to-Image Translation Models**: Learn to transform images based on natural language instructions - **Controllable Image Editing**: Fine-grained control over specific image attributes - **Style Transfer Models**: Temporal and aesthetic style transformation - **Instruction-Following Vision Models**: Models that can follow detailed editing instructions ## 🔗 Related Work This dataset is part of the **Nano Banana** dataset series and has been used in: - **UniPic3**: A unified multi-image composition framework that leverages this dataset for training. For more details, see [Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling](https://arxiv.org/abs/2601.15664) ## 📝 Citation If you use this dataset in your research, please cite: ```bibtex @misc{wei2026skyworkunipic30unified, title={Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling}, author={Hongyang Wei and Hongbo Liu and Zidong Wang and Yi Peng and Baixin Xu and Size Wu and Xuying Zhang and Xianglong He and Zexiang Liu and Peiyu Wang and Xuchen Song and Yangguang Li and Yang Liu and Yahui Zhou}, year={2026}, eprint={2601.15664}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2601.15664}, } ``` ```bibtex @misc{wang2025skyworkunipicunifiedautoregressive, title={Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation}, author={Peiyu Wang and Yi Peng and Yimeng Gan and Liang Hu and Tianyidan Xie and Xiaokun Wang and Yichen Wei and Chuanxin Tang and Bo Zhu and Changshi Li and Hongyang Wei and Eric Li and Xuchen Song and Yang Liu and Yahui Zhou}, year={2025}, eprint={2508.03320}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2508.03320}, } ``` --- **Note**: This dataset contains 123,268 samples and requires appropriate storage space for the image archive. Ensure you have sufficient disk space before downloading the full dataset.

Nano-banana-150k is an open-source large-scale image editing instruction dataset specifically designed for single-image editing training, serving as one of the core data sources for the UniPic3 model. Derived from the original Nano Banana dataset, it contains 123,268 carefully curated image editing samples covering various scenarios such as temporal style transformation, background replacement, action/gesture modification, hairstyle transformation, and artistic portrait generation. The dataset is stored in JSONL format with detailed natural language instructions and precise editing descriptions, making it suitable for training and evaluating image-to-image translation models, controllable image editing, style transfer models, and instruction-following vision models.

提供机构:
Skywork
搜集汇总
数据集介绍
Skywork/Nano-banana-150k 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务