SYNTHLLM
收藏资源简介:
SYNTHLLM是一个可扩展的框架,能够将预训练语料库转化为多样化、高质量的合成数据集。该数据集通过自动提取和重组合多个文档中的高级概念,使用图算法在不同领域(如数学推理)生成合成数据。研究结果表明,SYNTHLLM生成的合成数据遵循修正的扩展规律,性能提升在达到3000亿个token后趋于平稳,且大型模型在较少的训练token数下即可达到最佳性能。
SYNTHLLM is a scalable framework that transforms pre-trained corpora into diverse, high-quality synthetic datasets. It automatically extracts and recombines high-level concepts from multiple documents, and leverages graph algorithms to generate synthetic data across various domains such as mathematical reasoning. Research findings demonstrate that the synthetic data generated by SYNTHLLM follows a modified scaling law, with performance improvements plateauing after reaching 300 billion tokens, and large models can achieve optimal performance with fewer training tokens.

- 1Scaling Laws of Synthetic Data for Language Models微软 · 2025年



