BM-6M
收藏资源简介:
ByteMorph-6M是一个专注于基于指令的非刚性运动图像编辑的大型基准数据集,包含了源图像、目标图像、编辑提示和对应的图像标题等信息,用于支持文本到图像和基于指令的图像编辑研究。
ByteMorph-6M is a large-scale benchmark dataset focused on instruction-based non-rigid motion image editing. It contains source images, target images, editing prompts, corresponding image captions and other relevant information, to support research on text-to-image and instruction-based image editing.
数据集概述:ByteMorph-6M
数据集基本信息
- 许可证: CC0 1.0
- 任务类别: 图像到图像
- 规模类别: 1M < n < 10M
- 下载大小: 44.63 GB
- 数据集大小: 45.1 GB
- 训练集样本数: 780,308
数据集结构
特征
image_id: 字符串类型,表示从生成的视频中采样的图像对名称。src_img: 图像类型,源图像。tgt_img: 图像类型,编辑后的目标图像。edit_prompt: 字符串类型,编辑的视觉语言模型(VLM)描述。edit_prompt_rewrite_instruction: 字符串类型,将VLM描述重写为编辑指令。src_img_caption: 字符串类型,源图像的描述。tgt_img_caption: 字符串类型,目标图像的描述。
数据示例
json { "image_id": "[video_name]frame[i]_[j]", "src_img": "...", "tgt_img": "...", "edit_prompt": "The camera angle shifts to a closer view...", "edit_prompt_rewrite_instruction": "Zoom in the camera angle...", "src_img_caption": "Several individuals are present...", "tgt_img_caption": "Several individuals are gathered..." }
数据集详情
- 原始视频来源: 由Seaweed生成,并采样为源-目标图像编辑对。
- 处理方式: 通过视觉语言模型(VLM)进一步过滤和标注。
用途
- 主要用途: 用于基于文本和指令的图像编辑研究。
- 目标用户: 计算机视觉、图像生成、图像处理和AIGC领域的研究人员和爱好者。
使用方法
bash git lfs clone https://huggingface.co/datasets/ByteDance-Seed/BM-6M
相关资源
- 项目页面: https://boese0601.github.io/bytemorph/
- 基准测试: https://huggingface.co/datasets/ByteDance-Seed/BM-Bench
- 数据集演示: https://huggingface.co/datasets/ByteDance-Seed/BM-6M-Demo
- Gradio演示: https://huggingface.co/spaces/Boese0601/ByteMorpher-Demo
- 模型检查点: https://huggingface.co/ByteDance-Seed/BM-Model
- 代码: https://github.com/ByteDance-Seed/BM-code




