遇见数据集

Ba2han/tokenized_nemotron_science

收藏
Hugging Face2025-12-28 更新2026-03-29 收录
官方服务:

资源简介:

--- license: other language: - en - tr tags: - nemotron - mcq - science - tokenized size_categories: - n<100K dataset_info: features: - name: text dtype: string - name: input_ids list: int32 - name: token_type_ids list: int8 - name: attention_mask list: int8 splits: - name: train num_bytes: 24535006235 num_examples: 1309675 download_size: 3872933304 dataset_size: 24535006235 configs: - config_name: default data_files: - split: train path: data/train-* --- # Tokenized Nemotron-Science MCQ (Augmented) This dataset is a derived version of [nvidia/Nemotron-Science-v1](https://huggingface.co/nvidia/Nemotron-Science-v1) (MCQ subset). ## Processing Steps 1. **Augmentation**: The original conversation format (User, Assistant, Reasoning) was expanded into **8 different text formats** (Markdown headers, ChatML style, JSON style, etc.). 2. **Language Filtering**: Examples containing more than **20%** characters from Chinese, Hindi, or Russian scripts were removed. 3. **Length Filtering**: - Dropped the **shortest 1%** of examples. - Dropped the **longest 5%** of examples. 4. **Tokenization**: Tokenized using [Ba2han/turkish_model-sft](https://huggingface.co/Ba2han/turkish_model-sft). ## Statistics - **Tokenizer**: `Ba2han/turkish_model-sft` - **Total Documents**: 1309675 - **Total Tokens**: 2,570,753,504 - **Average Tokens per Document**: 1962.89 ## Source - Original Dataset: `nvidia/Nemotron-Science-v1`

许可证:其他 语言: - 英语 - 土耳其语 标签: - Nemotron(Nemotron) - 多项选择题(Multiple Choice Question, MCQ) - 科学 - 分词化(Tokenized) 样本量类别: - 样本数少于100,000(n<100K) 数据集信息: 特征: - 字段名:text,数据类型:字符串 - 字段名:input_ids,数据类型:int32类型列表 - 字段名:token_type_ids,数据类型:int8类型列表 - 字段名:attention_mask,数据类型:int8类型列表 数据划分: - 划分名称:训练集(train),字节数:24535006235,样本数:1309675 下载大小:3872933304 字节 数据集总大小:24535006235 字节 配置: - 配置名称:默认(default),数据文件: - 数据划分:训练集,路径:data/train-* # 分词化Nemotron(Nemotron)-科学多项选择题(增强版) 本数据集为[nvidia/Nemotron-Science-v1](https://huggingface.co/nvidia/Nemotron-Science-v1)的衍生版本,取自其多项选择题(Multiple Choice Question, MCQ)子集。 ## 处理流程 1. **数据增强**:将原始对话格式(用户、助手、推理过程)扩展为**8种不同文本格式**(如Markdown标题、ChatML风格、JSON风格等)。 2. **语言过滤**:移除包含超过20%中文、印地语或俄语字符的样本。 3. **长度过滤**: - 剔除最短的1%样本。 - 剔除最长的5%样本。 4. **分词化**:使用[Ba2han/turkish_model-sft](https://huggingface.co/Ba2han/turkish_model-sft)完成分词操作。 ## 统计信息 - **分词器(Tokenizer)**:`Ba2han/turkish_model-sft` - **总文档数**:1309675 - **总词元(Token)数**:2570753504 - **单文档平均词元(Token)数**:1962.89 ## 来源 - 原始数据集:`nvidia/Nemotron-Science-v1`

提供机构:
Ba2han
二维码
社区交流群
二维码
科研交流群
商业服务