VALSE
收藏资源简介:
VALSE是由海德堡大学计算语言学系创建的一个新型基准数据集,旨在测试预训练的视觉和语言(V&L)模型对特定语言现象的视觉语言基础能力。VALSE包含六个测试,覆盖了各种语言结构,要求模型在视觉模态中定位语言现象,实现比以往更细致的评估。数据集构建过程中采用了支持有效干扰项构建的方法,并报告了对五种广泛使用的V&L模型进行评估的结果。VALSE数据集利用现有的高质量图像描述和视觉问答数据,设计用于利用预训练(或微调)V&L模型中的现有预测头,因此不包括任何重新训练,可视为零样本评估。数据集的应用领域是测量预训练V&L模型在语言视角下的未来进展,补充了以任务为中心的V&L评估。
VALSE is a novel benchmark dataset created by the Department of Computational Linguistics, Heidelberg University, which aims to test the visual-language grounding ability of pre-trained vision-and-language (V&L) models on specific linguistic phenomena. VALSE consists of six tests covering a wide range of linguistic structures, requiring models to locate linguistic phenomena in the visual modality and enabling more fine-grained evaluations than prior benchmarks. A method supporting the construction of effective distractors was adopted during the dataset construction, and evaluation results of five widely used V&L models are reported. The VALSE dataset leverages existing high-quality image captioning and visual question answering (VQA) data, and is designed to utilize the existing prediction heads in pre-trained (or fine-tuned) V&L models, thus eliminating any need for retraining and can be regarded as a zero-shot evaluation. The dataset is intended to measure future progress of pre-trained V&L models from a linguistic perspective, complementing task-centric V&L evaluations.

- 1VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena海德堡大学计算语言学系 · 2022年



