Reasoner-1o1-v0.3-HQ
收藏资源简介:
该数据集是通过使用带有验证器的思维链和不同难度级别创建的。问题生成使用了较高的温度设置,以鼓励在随机类别中的更好创造性。数据集未经过手工筛选,但欢迎审查并留下评论以讨论任何细节。数据集的创建使用了Meta-Llama-3.1-405B-Instruct和Meta-Llama-3.1-70B-Instruct模型。原计划创建一个包含2000个问题的数据集,但由于会话中途崩溃,数据集生成被中断,只能上传恢复的部分。作者提到可能会尝试再次生成数据集,并提到了在Llama-3.1-8B-Instruct模型上训练的小数据集的潜力。
This dataset was constructed using chain-of-thought with validators across varying difficulty levels. A higher temperature parameter was employed during question generation to encourage enhanced creativity across random categories. The dataset has not undergone manual filtering, and we welcome reviews and comments for discussions on any details. The dataset was created using the Meta-Llama-3.1-405B-Instruct and Meta-Llama-3.1-70B-Instruct models. The original plan was to build a dataset containing 2000 questions, but dataset generation was interrupted due to an in-session crash, and only the recovered portion could be uploaded. The authors noted that they may attempt to regenerate the full dataset, and also highlighted the potential of a small dataset trained on the Llama-3.1-8B-Instruct model.
Reasoner-1o1-v0.3-HQ 数据集详情
数据集描述
- 创建方法:使用链式思维(chain of thought)与验证器(verifier)生成,包含不同难度级别的问题。
- 生成设置:问题生成时使用了较高的温度设置(higher temperature setting),以促进跨随机类别的更好创造性。
- 数据来源:数据集由 Meta-Llama-3.1-405B-Instruct 和 Meta-Llama-3.1-70B-Instruct 生成。
- 数据集规模:原计划创建一个包含2000个问题的数据集,但由于会话中途崩溃,仅能恢复部分数据。
- 未来计划:可能会尝试重新生成数据集,但目前无法做出承诺。
- 潜在应用:在 Llama-3.1-8B-Instruct 模型上训练,展示了一定的潜力,尽管模型并非总是正确,但潜力巨大。
数据集状态
- 当前状态:仅恢复了部分数据,数据集不完整。
- 未来更新:可能会尝试重新生成数据集,但目前无法确定。




