VisualPRM400K
收藏资源简介:
VisualPRM400K是一个包含大约40万个多模态过程监督数据的数据集,每个样本包括一个图像、一个问题、一个分步解答以及每一步的正确性注释。该数据集由复旦大学、上海人工智能实验室等机构构建,旨在为多模态过程奖励模型提供训练数据。数据集中的图像和问题来自MMPR v1.1,而分步解答则是通过InternVL2.5系列模型采样得到。通过自动数据管道对每个步骤的正确性进行注释,用于训练VisualPRM模型,该模型能够预测每一步的正确性。
VisualPRM400K is a dataset consisting of approximately 400,000 multimodal process supervision samples. Each sample comprises an image, a question, a step-by-step solution, and correctness annotations for each individual step. Developed by institutions including Fudan University and the Shanghai AI Laboratory, this dataset aims to provide training data for multimodal process reward models. The images and questions within the dataset are sourced from MMPR v1.1, while the step-by-step solutions were sampled using the InternVL2.5 series of models. Correctness annotations for each step were generated via an automated data pipeline, and the dataset is employed to train the VisualPRM model, which can predict the correctness of each step.




