VALUE (Video-And-Language Understanding Evaluation)
收藏资源简介:
VALUE 是一个视频和语言理解评估基准,用于测试可推广到不同任务、领域和数据集的模型。它是 11 个 VidL(视频和语言)数据集的集合,涵盖 3 个流行任务:(i)文本到视频检索; (ii) 视频问答; (iii) 视频字幕。 VALUE 基准旨在涵盖广泛的视频类型、视频长度、数据量和任务难度级别。 VALUE 不只关注具有视觉信息的单通道视频,而是推广利用来自视频帧及其相关字幕的信息的模型,以及跨多个任务共享知识的模型。用于 VALUE 基准测试的数据集是:TVQA、TVR、TVC、How2R、How2QA、VIOLIN、VLEP、YouCook2 (YC2C、YC2R)、VATEX
VALUE is a video and language understanding evaluation benchmark designed to test models that generalize across diverse tasks, domains and datasets. It comprises 11 VidL (Video and Language) datasets, covering three popular tasks: (i) Text-to-Video Retrieval; (ii) Video Question Answering; (iii) Video Captioning. The VALUE benchmark aims to cover a wide range of video genres, video durations, data scales and task difficulty levels. Instead of focusing solely on single-channel videos with visual information, VALUE promotes models that leverage information from video frames and their associated captions, as well as models that share knowledge across multiple tasks. The datasets utilized for the VALUE benchmark are: TVQA, TVR, TVC, How2R, How2QA, VIOLIN, VLEP, YouCook2 (YC2C, YC2R), VATEX




