HuggingFaceH4/orca_dpo_pairs
收藏资源简介:
OrcaDPO Pair数据集是OpenOrca数据集的一个子集,专门用于DPO偏好调优。数据集包含prompt、chosen和rejected三个部分,分别表示提示、被选择的回答和被拒绝的回答。数据集分为train_prefs和test_prefs两个部分,分别包含12359和500个样本。数据集的创建目的是为研究人员和开发者提供增强的文本数据,特别是用于增强FLAN Collection数据的推理能力。数据集的使用场景包括语言理解、自然语言处理、机器学习模型训练和模型性能评估。
The OrcaDPO Pair dataset is a subset of the OpenOrca dataset, specifically designed for DPO preference tuning. It contains three core components: prompt, chosen, and rejected, which respectively refer to the input prompt, the selected response, and the rejected response. The dataset is split into two subsets: train_prefs and test_prefs, which contain 12,359 and 500 samples respectively. The dataset is developed to provide enhanced textual data for researchers and developers, particularly to boost the reasoning capabilities of the FLAN Collection dataset. Typical application scenarios of the dataset include language understanding, natural language processing, machine learning model training, and model performance evaluation.
数据集概述
数据集名称
- OpenOrca数据集的预处理版本
数据集描述
- 该数据集是OpenOrca数据集的一个预处理版本,具体预处理步骤未详细说明。



