nemotron-nano2-safety-distill-gptoss
收藏资源简介:
Nemotron Nano 2 安全精简数据集,使用Nemotron Nano 2配方和GPT-OSS-20B及GPT-OSS-120B作为教师模型创建。由于资源限制,生成的推理步骤和响应没有经过守门模型的过滤。该数据集包含大约35,000个示例,可能会在未来版本中增加。数据来源于Nemotron内容安全数据集V2、gretel-v1、HarmfulTasks和RedTeam2K。
Nemotron Nano 2 Safety Trimmed Dataset is created using the Nemotron Nano 2 recipe, with GPT-OSS-20B and GPT-OSS-120B serving as the teacher models. Due to resource constraints, the generated inference steps and responses were not filtered by the moderation model. This dataset contains approximately 35,000 examples, and the sample size may be expanded in future versions. The data is sourced from the Nemotron Content Safety Dataset V2, gretel-v1, HarmfulTasks, and RedTeam2K.
Nemotron Nano 2 Safety Distill — GPT-OSS 数据集概述
数据集简介
- 基于Nemotron Nano 2安全配方创建的安全蒸馏数据集
- 使用GPT-OSS-20B和GPT-OSS-120B作为教师模型
- 包含约35,000个示例(截至2025年10月21日)
- 专为AI安全研究设计
数据来源
-
Aegis AI内容安全数据集v2.0
- 来源:https://huggingface.co/datasets/nvidia/Aegis-AI-Content-Safety-Dataset-2.0
-
Gretel安全对齐数据集v1
- 来源:https://huggingface.co/datasets/gretelai/gretel-safety-alignment-en-v1
-
恶意任务数据集
- 来源:https://github.com/CrystalEye42/eval-safety/blob/main/malicious_tasks_dataset.yaml
-
RedTeam-2K数据集
- 来源:https://huggingface.co/datasets/JailbreakV-28K/JailBreakV-28k/viewer/RedTeam_2K
数据集结构
配置子集
- aegis: 21,952个训练样本,1,244个验证样本,1,964个测试样本
- gretel-safety-alignment: 5,994个训练样本,1,181个验证样本,1,183个测试样本
- malicious-tasks: 225个训练样本
- redteam2k: 2,000个训练样本
数据特征
所有子集包含以下核心字段:
id: 数据点标识符prompt: 可能包含有害内容的输入提示reasoning_20b: GPT-OSS-20B的推理步骤response_20b: GPT-OSS-20B的响应reasoning_120b: GPT-OSS-120B的推理步骤response_120b: GPT-OSS-120B的响应metadata: 源数据集的附加元数据
元数据结构
aegis配置元数据字段:
- prompt_label
- prompt_label_source
- reconstruction_id_if_redacted
- response
- response_label
- response_label_source
- violated_categories
gretel-safety-alignment配置元数据字段:
- judge_response_reasoning
- judge_response_score
- judge_safe_response_reasoning
- judge_safe_response_score
- persona
- response
- response_probability_of_harm
- risk_category
- safe_response
- safe_response_probability_of_harm
- sub_category
- tactic
malicious-tasks配置元数据字段:
- category
- severity
- subcategory
redteam2k配置元数据字段:
- from
- policy
技术规格
- 任务类别: 文本生成、问答
- 语言: 英语
- 标签: gpt-oss、distillation、reasoning、ai-safety
- 规模类别: 10K-100K样本
警告说明
数据集包含可能有害的提示内容,仅限研究用途,使用时需负责任。




