PPE-MBPP-Plus-Best-of-K
收藏资源简介:
# Overview This contains the MBPP-Plus correctness preference evaluation set for Preference Proxy Evaluations. The prompts are sampled from [MBPP-Plus](https://huggingface.co/datasets/evalplus/mbppplus). This dataset is meant for benchmarking and evaluation, not for training. [Paper](https://arxiv.org/abs/2410.14872) [Code](https://github.com/lmarena/PPE) # License User prompts are licensed under Apache-2.0, and model outputs are governed by the terms of use set by the respective model providers. # Citation ``` @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate Reward Models for RLHF}, author={Evan Frick and Tianle Li and Connor Chen and Wei-Lin Chiang and Anastasios N. Angelopoulos and Jiantao Jiao and Banghua Zhu and Joseph E. Gonzalez and Ion Stoica}, year={2024}, eprint={2410.14872}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2410.14872}, } ```
# 数据集概览 本数据集包含用于偏好代理评估(Preference Proxy Evaluations)的MBPP-Plus正确性偏好评估集。 该数据集的提示词采样自[MBPP-Plus](https://huggingface.co/datasets/evalplus/mbppplus)。 本数据集专为基准测试与评估设计,不得用于模型训练。 [论文](https://arxiv.org/abs/2410.14872) [代码](https://github.com/lmarena/PPE) # 授权协议 用户提示词采用Apache-2.0开源协议进行授权,模型输出需遵循对应模型提供商的使用条款。 # 引用格式 @misc{frick2024evaluaterewardmodelsrlhf, title={How to Evaluate Reward Models for RLHF}, author={Evan Frick and Tianle Li and Connor Chen and Wei-Lin Chiang and Anastasios N. Angelopoulos and Jiantao Jiao and Banghua Zhu and Joseph E. Gonzalez and Ion Stoica}, year={2024}, eprint={2410.14872}, archivePrefix={arXiv}, primaryClass={cs.LG}, url={https://arxiv.org/abs/2410.14872}, }




