SecureBreak
收藏资源简介:
SecureBreak是由帕维亚大学团队开发的面向AI安全的数据集,旨在检测大语言模型因安全对齐缺陷产生的有害输出。该数据集包含3059条经过人工标注的文本样本,数据源自对Llama、Qwen等主流开源模型在JailbreakBench对抗性提示下生成响应的系统收集,采用双人标注机制确保标注一致性(Cohen's Kappa=0.85)。其核心价值在于构建生成后过滤模块,既可作为阻断有害内容的最终防线,又能通过监督信号优化模型对齐流程,主要应用于AI安全、内容审核和伦理对齐研究领域。
SecureBreak is a dataset developed by the team from the University of Pavia for AI safety, aiming to detect harmful outputs generated by large language models (LLMs) due to safety alignment flaws. This dataset contains 3059 manually annotated text samples, which are systematically collected from the responses of mainstream open-source models such as Llama and Qwen under adversarial prompts from JailbreakBench. It adopts a double annotation mechanism to ensure annotation consistency (Cohen's Kappa = 0.85). Its core value lies in building post-generation filtering modules, which can not only serve as the final line of defense for blocking harmful content, but also optimize the model alignment process through supervision signals. It is mainly applied in the research fields of AI safety, content moderation and ethical alignment.
SecureBreak 数据集概述
数据集基本信息
- 数据集名称: SecureBreak
- 核心用途: 专门用于将响应分类为安全或不安全,旨在支持开发可靠的响应级别分类器。
- 版本: 1.0
- 数据格式: CSV
- 数据规模: 包含 3059 条记录。
- 特征数量: 9 个特征。

- 1SecureBreak -- A dataset towards safe and secure models帕维亚大学·电气、计算机与生物医学工程系 · 2026年



