SAV Caption Dataset
收藏资源简介:
该页面包含SAV数据集中对象的伪标题和人工创建的标题,用于VoCap论文。具体包括:对于SAV验证集,由人工标注者为每个标注对象提供标题,每个对象由三个不同的标注者标注;对于SAV训练集,通过Gemini 1.5 Pro生成以对象为中心的标题
This page contains pseudo-titles and human-generated titles for objects in the SAV dataset, prepared for the VoCap paper. Specifically: For the SAV validation set, human annotators provided titles for each annotated object, with each object being annotated by three distinct annotators; For the SAV training set, object-centric titles were generated via Gemini 1.5 Pro.
vocap 数据集概述
数据集来源
该数据集为论文《VoCap: Video Object Captioning and Segmentation from Any Prompt》中使用的SAV数据集的伪标注和人工标注字幕。
数据集内容
- 验证集标注:SAV验证集中每个标注对象均由三名不同标注人员提供人工描述字幕
- 训练集标注:SAV训练集通过Gemini 1.5 Pro模型基于真实标注生成以对象为中心的自动字幕
数据格式
提供两个CSV文件:
sav_caption_val_human.csv:验证集人工标注sav_caption_train_automatic.csv(14MB):训练集自动生成标注
每行包含video_id、object_id、caption(逗号分隔)。验证集中大多数video_id、object_id对重复三次,对应三名标注人员的标注结果。
引用信息
@inproceedings{uijings25vocap, title={{VoCap}: Video Object Captioning and Segmentation from Any Prompt}, author={Jasper Uijlings and Xingyi Zhou and Xiuye Gu and Arsha Nagrani and Anurag Arnab and Alireza Fathi and David Ross and Cordelia Schmid}, booktitle={ArXiv}, year={2025}, }
许可信息
- 软件部分:Apache License 2.0(https://www.apache.org/licenses/LICENSE-2.0)
- 其他材料:Creative Commons Attribution 4.0 International License(CC-BY)(https://creativecommons.org/licenses/by/4.0/legalcode)
- 免责声明:非Google官方产品




