Nemotron-RL-Instruction-Following-Adversarial-v1
收藏资源简介:
inverseIF数据集专注于设计对抗性提示,这些提示明确与AI模型的标准训练本能相冲突,涵盖了8种不同的“反惯例”模式。该数据集采用“模型破坏”方法,通过Nemotron-Nano-V2或Qwen3-235B-A22B-Thinking-2507生成四个候选响应,以测试负面约束是否足够困难以迫使模型表现出默认行为失败。这些响应由人类评委和GPT-5 LLM评委严格评估(要求两者之间至少有85%的一致率),只有当样本成功“破坏”模型时才会被接受(即四个响应中最多有一个通过严格标准,同时展示出差异性)。最终挑战性样本被格式化为包含对抗性提示、真实答案、候选响应和详细的双重评估指标的JSON文件。该数据集作为NVIDIA NeMo Gym的一部分发布,用于训练大型语言模型的强化学习环境。数据集包含100条记录,总存储量为72MB,采用CC-BY 4.0许可,适用于商业用途。
The inverseIF Dataset focuses on designing adversarial prompts that explicitly conflict with the standard training instincts of AI models, covering 8 distinct "anti-convention" patterns. This dataset adopts the "model-breaking" approach, generating four candidate responses via Nemotron-Nano-V2 or Qwen3-235B-A22B-Thinking-2507 to test whether the negative constraints are sufficiently difficult to force the model to fail at its default behavior. These responses are strictly evaluated by both human annotators and GPT-5 LLM annotators, with a required minimum agreement rate of 85% between the two groups. A sample is accepted only when it successfully "breaks" the model, i.e., at most one of the four candidate responses passes the strict criteria while exhibiting distinct differences. The final challenging samples are formatted into JSON files that contain adversarial prompts, ground-truth answers, candidate responses, and detailed dual evaluation metrics. This dataset is released as part of NVIDIA NeMo Gym, serving as a reinforcement learning environment for training large language models. The dataset contains 100 records with a total storage size of 72 MB, and is licensed under CC-BY 4.0 for commercial use.



