HUNYUANPROVER数据集
收藏资源简介:
HUNYUANPROVER数据集由腾讯混元团队创建,旨在解决自动定理证明中的数据稀缺问题。该数据集包含30,000条数据实例,每条实例包括自然语言中的原始问题、自动形式化转换后的陈述以及由HunyuanProver生成的证明。数据集的生成过程涉及从130,000条高质量的自然语言到LEAN格式的陈述对开始,通过自动形式化模型将3000万条内部数学问题转换为形式化陈述,并经过多轮迭代生成证明数据。该数据集的应用领域主要集中在自动定理证明,旨在通过大规模数据生成和迭代优化提升定理证明模型的性能。
The HUNYUANPROVER dataset was developed by the Tencent Hunyuan Team to address the data scarcity challenge in automated theorem proving. This dataset contains 30,000 data instances, each comprising the original natural language problem, the automatically formalized statement, and the proof generated by HunyuanProver. The dataset generation workflow starts with 130,000 high-quality natural language-to-LEAN statement pairs, then leverages automated formalization models to convert 30 million internal mathematical problems into formal statements, and finally produces proof data through multiple iterative rounds. The primary application domain of this dataset is automated theorem proving, aiming to boost the performance of theorem proving models via large-scale data generation and iterative optimization.




