MaLA-LM/mala-code-reasoning-v2
收藏官方服务:
资源简介:
MaLA语料库(大规模语言适应语料库)是一个全面的多语种数据集,旨在支持大型语言模型的持续预训练。这个子集是第二个版本,包含了代码、推理数据和科学论文。
The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. This subset is the second version and contains code, reasoning data, and scientific papers.
提供机构:
MaLA-LM


