IndustryCorpus_ai
收藏资源简介:
该数据集是为了解决行业模型训练数据集存在的问题而构建的,包含3.4TB的高质量多行业分类的中英文预训练数据,其中1TB为中文数据,2.4TB为英文数据。数据集涵盖18个行业类别,包括医疗、教育、文学、金融等,并对中文数据进行了12种类型的标签标注。数据处理包括22个行业数据处理操作符的应用,基于规则和模型的过滤,以及文档级别的去重处理。
This dataset is constructed to address the issues existing in training datasets for industry-specific models. It contains 3.4 TB of high-quality multi-industry categorized pre-training data in both Chinese and English, with 1 TB being Chinese data and 2.4 TB being English data. The dataset covers 18 industry categories including healthcare, education, literature, finance and others, and 12 types of label annotations are conducted on the Chinese data. Data processing includes the application of 22 industry data processing operators, rule-based and model-based filtering, as well as document-level deduplication.
数据集概述
数据集基本信息
- 许可证:Apache-2.0
- 语言:中文、英文
- 数据量:超过1TB
- 任务类别:文本生成
数据集构建
- 原始数据来源:包括WuDaoCorpora、BAAI-CCI、redpajama、SkyPile-150B等超过100TB的开放源数据集。
- 处理方法:应用22个行业数据处理操作符进行清洗和过滤。
- 过滤后数据量:1TB中文数据,2.4TB英文数据。
数据标注
- 中文数据标注:包括字母数字比、平均行长度、语言置信度分数、最大行长度、困惑度等12种标签。
数据验证
- 验证方法:在医疗行业示范模型上进行持续预训练、SFT和DPO训练。
- 验证结果:客观性能提升20%,主观胜率82%。
数据集详细信息
- 行业分类:包括医疗、教育、文学、金融、旅游、法律、体育、汽车、新闻等18个类别。
- 基于规则的过滤:包括繁体中文转换、电子邮件移除、IP地址移除、链接移除、Unicode修复等。
- 基于模型的过滤:行业分类语言模型,准确率80%。
- 数据去重:MinHash文档级去重。
行业分类数据量
| 行业类别 | 数据量 (GB) | 行业类别 | 数据量 (GB) |
|---|---|---|---|
| 编程 | 4.1 | 政治 | 326.4 |
| 法律 | 274.6 | 数学 | 5.9 |
| 教育 | 458.1 | 体育 | 442 |
| 金融 | 197.8 | 文学 | 179.3 |
| 计算机科学 | 46.9 | 新闻 | 564.1 |
| 技术 | 333.6 | 影视 | 162.1 |
| 旅游 | 82.5 | 医学 | 189.4 |
| 农业 | 41.6 | 汽车 | 40.8 |
| 情感 | 31.7 | 人工智能 | 5.6 |
| 总计 (GB) | 3386.5 |




