UGC-VideoCap
收藏资源简介:
UGC-VideoCap是一个大型基准数据集,专为详细描述包含音频和视觉信息的短视频而设计。它包括约1000个TikTok视频,每个视频都富含语音轨道和多样化的内容。数据集采用了严格的三阶段人工标注流程,包括仅音频、仅视觉和音频视觉联合语义的标注。此外,还包括超过4000个手工设计的开放式和多项选择题,以全面探索视觉和听觉方面的理解。UGC-VideoCap旨在促进全模态视频理解的研究,特别是在UGC和电影等现实世界中,丰富的多模态线索和细粒度语义是至关重要的。
UGC-VideoCap is a large-scale benchmark dataset specifically designed for detailed captioning of short videos containing both audio and visual information. It includes approximately 1,000 TikTok videos, each with rich audio tracks and diverse content. The dataset adopts a strict three-stage manual annotation workflow, covering annotations for audio-only, visual-only, and audio-visual joint semantics. Additionally, it contains over 4,000 manually designed open-ended and multiple-choice questions to comprehensively explore visual and auditory understanding. UGC-VideoCap aims to facilitate research on full-modal video understanding, especially in real-world scenarios such as UGC and film, where rich multimodal cues and fine-grained semantics are critically important.




