Open-vocabulary Video Question Answering (OVQA)
收藏资源简介:
Open-vocabulary Video Question Answering (OVQA) 是一个评估视频问答模型泛化能力的新基准。该数据集旨在通过考虑罕见和未见过的答案来衡量模型的泛化能力。OVQA 数据集通过引入一个基于图神经网络(GNN)的软词义化器,增强了模型对罕见和未见过答案的预测能力,从而提高了模型的泛化性能。此数据集适用于评估模型在长尾分布,包括未见过答案的情况下的表现,旨在解决现有模型偏向频繁答案而无法泛化到罕见和未见过答案的问题。
Open-vocabulary Video Question Answering (OVQA) is a novel benchmark for evaluating the generalization capabilities of video question answering models. This dataset is designed to assess a model's generalization ability by evaluating against rare and unseen answers. The OVQA dataset incorporates a graph neural network (GNN)-based soft lexicalizer to enhance the model's predictive performance for rare and unseen answers, thereby boosting its generalization capacity. This dataset is tailored for evaluating model performance under long-tailed distributions, including cases where answers are unseen, and aims to resolve the prevalent issue that existing models are biased towards frequent answers and cannot generalize to rare and unseen ones.




