Goku
收藏资源简介:
Goku是由中国科学技术大学与腾讯混元联合构建的大规模指令式视频编辑数据集,首次将任务边界从基础外观编辑扩展到多任务与结构化操作。该数据集包含200万高质量视频编辑对,涵盖10类核心编辑任务,视频分辨率为720p、时长3-10秒,数据源来自Koala-36M的精选视频片段,通过多模态大模型生成多样化编辑指令。构建过程采用自动化流水线设计,将复杂编辑分解为可控子问题,并引入渐进式过滤系统确保语义精度与时序连贯性。该数据集旨在解决现有方法在复杂结构编辑和多任务协同方面的局限性,为视频编辑模型的训练与评估提供全面基准。
Goku is a large-scale instructional video editing dataset jointly constructed by the University of Science and Technology of China and Tencent Hunyuan. It encompasses 2 million high-quality video editing pairs, and for the first time expands the task scope from basic appearance editing to multi-task and structural editing. This dataset covers 10 core editing tasks, with all videos at 720p resolution and each containing between 65 and 129 frames. The data is sourced from curated video clips from Koala-36M, and generated through an automated pipeline integrated with a progressive filtering system. Its construction employs a task decomposition strategy, utilizing Gemini 2.5-Pro to generate instructional prompts and ensure semantic fidelity and temporal consistency. This dataset is intended to provide training and evaluation benchmarks for complex video editing models, addressing the existing shortcomings of prior datasets in structural transformation and multi-task editing.
数据集概述
数据集名称:Goku
规模:200万(2 million)高质量、指令对齐的视频编辑对
任务范围:首个将任务边界从基础外观编辑扩展到多任务和结构操控(如精确控制主体运动)的大规模数据集
数据合成:设计了高效的数据合成流程,将复杂编辑分解为可控子问题,并引入渐进式过滤系统保证数据可靠性
数据集样本类别
| 类别 | 描述 |
|---|---|
| 多任务编辑 | 通过多轮编辑生成的编辑数据,示例包括2步和3步序列 |
| 结构编辑 | 包括摄像机运动和主体运动,如摄像机平移、倾斜、缩放,主体低头、闭眼、改变姿势等 |
| 参考添加 | 从参考图像中向视频添加指定对象 |
| 参考替换 | 用参考图像中的对象替换视频中的对象 |
| 添加编辑 | 向视频场景中添加新对象或元素 |
| 移除编辑 | 从视频场景中移除现有对象或元素 |
| 替换与修改 | 替换或修改视频中的特定对象和属性,如换装、改色、背景替换等 |
| 风格编辑 | 转移或修改整个视频的视觉风格,如城市速写、油画、波普艺术、皮克斯风格 |
基准与模型
基准:Goku-Bench
- 包含 1,000 个人工验证的测试用例
- 引入 7 项新颖的编辑专用评估指标
模型:Goku-Edit
- 采用多模态大语言模型(MLLM)作为文本编码器
- 解耦双分支设计:一个专用掩码分支处理结构控制,主分支负责外观渲染
- 在Goku-Bench上,指令跟随能力相较其他开源模型提升高达 +8%
与其他方法对比
在7个示例中,Goku-Edit与InsV2V、InsVIE、OmniVideo、Lucy等方法进行并排比较,覆盖多任务、结构编辑、外观编辑等场景。




