Multiple Multimodal Artificial Intelligence Preference Datasets in VQA (MMAIP-V)
收藏资源简介:
MMAIP-V是由快手科技和中国人民大学高瓴人工智能学院联合创建的高质量视频问答偏好数据集,旨在促进多模态大语言模型(MLLMs)的偏好学习。该数据集包含24,000条视频问答对,通过从多模态大语言模型的响应分布中采样,并利用外部评分函数进行响应质量评估构建而成。数据集的创建过程结合了多种视觉语言模型的反馈,确保了正负响应的高质量和多样性。MMAIP-V主要应用于视频问答领域的偏好学习,旨在解决现有数据集质量低、多样性不足的问题,从而提升MLLMs的指令遵循能力和减少幻觉现象。
MMAIP-V is a high-quality video question answering preference dataset jointly created by Kuaishou Technology and Gaoling School of Artificial Intelligence, Renmin University of China, aiming to facilitate preference learning for multimodal large language models (MLLMs). This dataset contains 24,000 video question-answering pairs, which are constructed by sampling from the response distribution of multimodal large language models and evaluating response quality using external scoring functions. The dataset creation process incorporates feedback from multiple visual-language models, ensuring the high quality and diversity of both positive and negative responses. MMAIP-V is mainly applied to preference learning in the field of video question answering, aiming to address the issues of low quality and insufficient diversity in existing datasets, thereby improving the instruction-following ability of MLLMs and reducing hallucinations.

- 1Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models快手科技,北京,中国 中国人民大学,高瓴人工智能学院,北京 · 2024年



