CofeAI/NanoData
收藏资源简介:
--- license: other license_name: other license_link: LICENSE task_categories: - text-generation language: - en size_categories: - 100B<n<1T --- ### Dataset Description To facilitate researchers to use [NanoLM](https://github.com/cofe-ai/nanoLM?tab=readme-ov-file) for comparative analysis across different model designs, we build a curated pre-training dataset from those of existing large-scale models (i.e., Llama, Falcon, GPT-3). It covers diverse domains to improve the generalization capabilities of the resultant models. #### Dataset Creation The data is mainly post-processed and filtered from [RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T) and [RedPajamaV2](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2). We develop a series of cleaning steps to remove redundant formatting, garbled characters, formula errors, duplicated paragraphs, low-quality text, and other unwanted content. After interleaved deduplication on document level of each independent subset, we finally obtain a high-quality dataset. #### Dataset Summary | Dataset | Num Tokens (B) | | -------------- | -------------- | | CommonCrawl | 67.00 | | C4 | 15.00 | | Wikipedia (En) | 5.14 | | Books | 4.48 | | ArXiv | 2.50 | | StackExchange | 2.00 | | Total | 97.12 | We release the data with approximate 100B tokens. Furthermore, we recommend users to add code dataset such as [Starcode](https://huggingface.co/datasets/bigcode/starcoderdata), [The Stack V2](https://huggingface.co/datasets/bigcode/the-stack-v2-dedup) to enrich model's performance on code and reasoning. ### Citation To cite NanoLM, please use: ``` @misc{yao2024nanolm, title={nanoLM: an Affordable LLM Pre-training Benchmark via Accurate Loss Prediction across Scales}, author={Yiqun Yao and Siqi fan and Xiusheng Huang and Xuezhi Fang and Xiang Li and Ziyi Ni and Xin Jiang and Xuying Meng and Peng Han and Shuo Shang and Kang Liu and Aixin Sun and Yequan Wang}, year={2024}, eprint={2304.06875}, archivePrefix={arXiv}, primaryClass={cs.CL} } ``` ### Acknowledgement The data is mainly curated and filtered from [RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T) and [RedPajamaV2](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2). We extend our gratitude to the original authors for their innovative work and for making it available to the community. ### License The code of NanoLM used to process the dataset and loss prediction is licensed under the Apache 2.0 license. For curated data, please refer to the licenses of the original ones. * [Common Crawl Foundation Terms of Use](https://commoncrawl.org/terms-of-use) * [C4 license](https://huggingface.co/datasets/allenai/c4#license) * Books: [the_pile_books3 license](https://huggingface.co/datasets/defunct-datasets/the_pile_books3#licensing-information) and [pg19 license](https://huggingface.co/datasets/deepmind/pg19#licensing-information) * [ArXiv Terms of Use](https://info.arxiv.org/help/api/tou.html) * [Wikipedia License](https://huggingface.co/datasets/legacy-datasets/wikipedia#licensing-information) * [StackExchange license on the Internet Archive](https://archive.org/details/stackexchange)
--- license: 其他 license_name: 其他 license_link: LICENSE task_categories: - 文本生成 language: - 英语 size_categories: - 1000亿 < 令牌数 < 1万亿 --- ### 数据集说明 为助力研究人员使用[NanoLM](https://github.com/cofe-ai/nanoLM?tab=readme-ov-file)开展不同模型设计的对比分析,我们从现有大规模模型(即Llama、Falcon、GPT-3)的预训练数据集中构建了经过精选的预训练数据集。本数据集覆盖多元领域,旨在提升所得模型的泛化能力。 #### 数据集构建流程 本数据集主要源自[RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T)与[RedPajamaV2](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2)数据集,经后处理与筛选得到。我们设计了一系列数据清洗流程,用于移除冗余格式、乱码字符、公式错误、重复段落、低质量文本等无效内容。在对每个独立子集进行文档级交叉去重后,最终得到高质量的预训练数据集。 #### 数据集概览 | 数据集名称 | 令牌数(十亿) | | ---------------- | -------------- | | CommonCrawl | 67.00 | | C4 | 15.00 | | 英文维基百科(Wikipedia (En)) | 5.14 | | 图书语料(Books) | 4.48 | | ArXiv | 2.50 | | StackExchange | 2.00 | | 总计 | 97.12 | 本数据集的发布规模约为1000亿令牌。此外,我们建议用户补充代码类数据集,例如[Starcode](https://huggingface.co/datasets/bigcode/starcoderdata)与[The Stack V2](https://huggingface.co/datasets/bigcode/the-stack-v2-dedup),以提升模型在代码与推理任务上的性能表现。 ### 引用方式 若需引用NanoLM,请使用以下格式: @misc{yao2024nanolm, title={nanoLM: an Affordable LLM Pre-training Benchmark via Accurate Loss Prediction across Scales}, author={Yiqun Yao and Siqi fan and Xiusheng Huang and Xuezhi Fang and Xiang Li and Ziyi Ni and Xin Jiang and Xuying Meng and Peng Han and Shuo Shang and Kang Liu and Aixin Sun and Yequan Wang}, year={2024}, eprint={2304.06875}, archivePrefix={arXiv}, primaryClass={cs.CL} } ### 致谢 本数据集主要源自[RedPajama](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T)与[RedPajamaV2](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2)数据集的精选与筛选工作。在此,我们向原始数据集的作者致敬,感谢其创新性的研究工作与开源共享精神。 ### 许可证声明 用于数据集处理与损失预测的NanoLM代码采用Apache 2.0许可证进行授权。 对于本精选数据集,请遵循其原始数据集的许可证要求: * [Common Crawl 基金会使用条款](https://commoncrawl.org/terms-of-use) * [C4 许可证](https://huggingface.co/datasets/allenai/c4#license) * 图书语料:遵循[the_pile_books3 许可证](https://huggingface.co/datasets/defunct-datasets/the_pile_books3#licensing-information)与[pg19 许可证](https://huggingface.co/datasets/deepmind/pg19#licensing-information) * [ArXiv 使用条款](https://info.arxiv.org/help/api/tou.html) * [维基百科 许可证](https://huggingface.co/datasets/legacy-datasets/wikipedia#licensing-information) * [互联网档案馆中的StackExchange许可证](https://archive.org/details/stackexchange)
数据集描述
数据集创建
本数据集是为了支持研究人员使用NanoLM进行不同模型设计的比较分析而构建的。数据主要从RedPajama和RedPajamaV2中经过一系列清洗步骤处理和过滤得到,包括去除冗余格式、乱码、公式错误、重复段落、低质量文本等。
数据集总结
| 数据集 | 令牌数量(B) |
|---|---|
| CommonCrawl | 67.00 |
| C4 | 15.00 |
| Wikipedia (En) | 5.14 |
| Books | 4.48 |
| ArXiv | 2.50 |
| StackExchange | 2.00 |
| 总计 | 97.12 |
数据集包含约100B令牌。建议用户添加如Starcode和The Stack V2等代码数据集以增强模型在代码和推理方面的性能。
许可证
数据集的原始数据遵循各自原始数据的许可证。具体包括:



