S-Eval
收藏资源简介:
S-Eval是由浙江大学和阿里巴巴集团联合创建的大型语言模型安全评估数据集,包含220,000个评估提示,旨在系统地评估大型语言模型(LLMs)的安全性。数据集包括20,000个基础风险提示(10,000中文和10,000英文)和200,000个相应的攻击提示,这些攻击提示源自10种流行的对抗性指令攻击。S-Eval设计灵活,能够根据LLMs的快速演进和伴随的安全威胁,灵活配置和适应新的风险、攻击和模型,以持续更新基准。该数据集广泛应用于20个流行且具有代表性的LLMs评估中,结果证实S-Eval能更有效地反映和告知LLMs的安全风险,相比于现有基准。
S-Eval is a large language model (LLM) safety evaluation dataset jointly created by Zhejiang University and Alibaba Group. It contains 220,000 evaluation prompts, aiming to systematically assess the safety of large language models (LLMs). The dataset consists of 20,000 base risk prompts (10,000 in Chinese and 10,000 in English) and 200,000 corresponding adversarial prompts, which are derived from 10 prevalent adversarial instruction attack scenarios. S-Eval features a flexible design that enables flexible configuration and adaptation to new risks, attack methods, and models in line with the rapid evolution of LLMs and accompanying security threats, so as to support continuous benchmark updates. This dataset has been widely applied in the evaluation of 20 popular and representative LLMs, and the results confirm that S-Eval can more effectively reflect and inform the security risks of LLMs compared with existing benchmarks.




