统一偏好数据集
收藏资源简介:
本文构建了一个大规模的统一偏好数据集,该数据集整合了现有的多个数据集,经过预处理后,包含了图像和视频的理解与生成任务的人类偏好数据。数据集分为 pairwise ranking 和 pointwise scoring 两种类型,覆盖了236000条数据,旨在训练能够跨视觉任务进行评估的统一奖励模型 UNIFIEDREWARD。该数据集的构建是为了解决现有奖励模型任务特定、适应性差的问题,并为多模态理解和生成任务提供了一种新的评估方法。
This paper constructs a large-scale unified preference dataset. Integrating multiple existing datasets and undergoing preprocessing, this dataset contains human preference data for image and video understanding and generation tasks. It is divided into two types: pairwise ranking and pointwise scoring, with a total of 236,000 entries, aiming to train the UnifiedReward unified reward model capable of conducting evaluations across visual tasks. This dataset is developed to address the issues of task-specificity and poor adaptability of existing reward models, and to provide a novel evaluation method for multimodal understanding and generation tasks.

- 1Unified Reward Model for Multimodal Understanding and Generation复旦大学, 上海创新研究院, 上海人工智能实验室, 上海科学院人工智能研究所 · 2025年



