VerifyBench
收藏资源简介:
在本文中,我们提出了VerifyBench,这是一个专门设计用于评估基于参考的奖励系统准确性的基准。为了创建VerifyBench,我们从现有的开放数据集中筛选了多样化的指令和参考答案配对。这些指令的响应由多个开源和专有的大型语言模型生成。每个响应的正确性通过自动模型判断和人工评估进行评估。VerifyBench中的每个实例都经过至少两名人类注释者的验证,以确保标签的一致性和可靠性,从而为奖励系统的评估提供了一个高质量的基准。
In this paper, we propose VerifyBench, a benchmark specifically designed to evaluate the accuracy of reference-based reward systems. To construct VerifyBench, we curated diverse instruction and reference answer pairs from existing open datasets. The responses to these instructions were generated by multiple open-source and proprietary large language models (LLMs). The correctness of each response was evaluated via both automated model judgments and human evaluation. Every instance in VerifyBench has been validated by at least two human annotators to ensure label consistency and reliability, thereby providing a high-quality benchmark for reward system evaluation.
VerifyBench 数据集概述
基本信息
- 数据集名称: VerifyBench
- 开发团队: 浙江大学、美团集团等机构联合开发
- 状态: 预印本,正在评审中
- 发布日期: 2025年5月
- 相关链接:
数据集描述
- 核心目标: 评估基于参考的奖励系统在大语言模型中的准确性
- 数据构成:
- 收集多样化指令与参考回答(源自现有开放数据集)
- 包含多个开源和专有LLM生成的响应
- 每个响应的正确性通过自动模型判断和人工评估双重验证
- 质量保证: 每个实例至少经过两名人类标注者验证
衍生数据集
- VerifyBench-Hard:
- 挑战性更强的变体
- 聚焦于领先模型产生高度冲突判断的争议性案例
- 样本基于高性能模型间的分歧模式精选
- 经过严格人工标注确保标签质量
主要贡献
- 构建VerifyBench基准测试,客观评估基于参考的奖励系统准确性
- 开发VerifyBench-Hard基准测试,突出当前模型的改进潜力
- 提供全面的实证分析,推动奖励系统准确性和RL训练的进步
引用格式
bibtex @misc{yan2025verifybench, title={VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models}, author={Yuchen Yan et al.}, year={2025}, eprint={2505.15801}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2505.15801}, }
联系方式
- 联系邮箱: yanyuchen@zju.edu.cn




