SafeSteerDataset
收藏资源简介:
SafeSteerDataset是由NASK国家研究机构等联合构建的对比数据集,包含2300组安全与不安全提示词对,覆盖性、暴力、仇恨等23个子类别。数据通过Gemini 2.5-Pro生成初稿后,经Qwen-8b嵌入模型筛选(余弦相似度>0.7),确保语义对齐。该数据集专为文本到图像模型的安全转向研究设计,用于精准隔离毒性激活流形,解决现有方法在良性提示上干扰图像质量的问题。
SafeSteerDataset is a contrastive dataset jointly constructed by NASK and other national research institutions. It contains 2300 pairs of safe and unsafe prompts, covering 23 subcategories such as sexual, violent, hateful content and others. The dataset was initially drafted by Gemini 2.5-Pro, then screened by the Qwen-8b embedding model with a cosine similarity threshold greater than 0.7 to ensure semantic alignment. This dataset is specifically designed for safety steering research on text-to-image models, aiming to accurately isolate the toxic activation manifold and resolve the issue where existing methods degrade image quality when handling benign prompts.

- 1Conditioned Activation Transport for T2I Safety SteeringNASK国家研究机构; 华沙理工大学; Tooplox; IDEAS研究所; CISPA亥姆霍兹信息安全中心 · 2026年



