Text-to-Video Quality Assessment DataBase (T2VQA-DB)
收藏资源简介:
T2VQA-DB是由上海交通大学创建的大规模数据集,包含10000个由9种不同文本到视频(T2V)模型生成的视频,每个视频都配有主观评分。数据集通过27名受试者的主观实验收集了每个视频的平均意见分数(MOS),旨在解决现有视频质量评估模型无法准确量化文本生成视频质量的问题。T2VQA-DB不仅用于训练和测试后续模型,还支持提出了一种基于Transformer的新模型T2VQA,该模型从文本-视频对齐和视频保真度两个角度提取特征,并利用大型语言模型进行质量预测,有效提升了文本生成视频质量评估的准确性。
T2VQA-DB is a large-scale dataset developed by Shanghai Jiao Tong University. It includes 10,000 videos generated by 9 different text-to-video (T2V) models, with each video accompanied by subjective ratings. The dataset collects the Mean Opinion Scores (MOS) for each video through subjective experiments involving 27 participants, aiming to address the limitation that existing video quality assessment models fail to accurately quantify the quality of text-generated videos. Besides being utilized for training and testing subsequent models, T2VQA-DB also supports the proposal of a novel Transformer-based model named T2VQA. This model extracts features from two perspectives: text-video alignment and video fidelity, and leverages large language models to conduct quality prediction, effectively improving the accuracy of text-generated video quality assessment.

- 1Subjective-Aligned Dataset and Metric for Text-to-Video Quality Assessment上海交通大学 · 2024年



