VerifyBench
收藏资源简介:
VerifyBench是一个专门设计用于评估基于参考的奖励系统的基准数据集,旨在填补现有奖励基准在评估推理模型训练中使用的验证系统方面的空白。数据集由来自现有开放数据集的指令和参考答案组成,并由多个开源和专有的大型语言模型生成响应。每个实例都经过至少两名人工标注者的验证,以确保标签的一致性和可靠性。VerifyBench-Hard是VerifyBench的一个更具挑战性的变体,专注于模型之间高度分歧的情况,为奖励系统的准确性提供了更严格的测试。
VerifyBench is a benchmark dataset specifically designed for evaluating reference-based reward systems, aiming to fill the gap in existing reward benchmarks regarding validation systems used during the training of reasoning models. The dataset comprises instructions and reference answers sourced from existing open datasets, with responses generated by multiple open-source and proprietary large language models. Each instance has been validated by at least two human annotators to ensure the consistency and reliability of the labels. VerifyBench-Hard is a more challenging variant of VerifyBench, focusing on scenarios with high levels of disagreement between models, which provides a stricter test for the accuracy of reward systems.




