AgentHarm
收藏资源简介:
AgentHarm数据集由英国人工智能安全研究所创建,旨在评估大型语言模型(LLM)代理在执行多步骤任务时的安全性和鲁棒性。该数据集包含110个基本恶意任务和330个增强任务,共计440个任务,涵盖11种危害类别,包括欺诈、网络犯罪和骚扰等。数据集通过合成工具和细粒度评分标准,确保任务的可靠性和安全性。AgentHarm数据集的应用领域主要集中在LLM代理的安全性研究,旨在解决代理在执行恶意任务时的潜在风险问题。
The AgentHarm dataset was developed by the UK AI Safety Institute, with the core goal of evaluating the safety and robustness of large language model (LLM) agents when executing multi-step tasks. This dataset includes 110 basic malicious tasks and 330 enhanced tasks, totaling 440 tasks covering 11 harm categories such as fraud, cybercrime, harassment and others. The reliability and security of the tasks are guaranteed through synthetic tools and fine-grained scoring criteria. The main application scenario of the AgentHarm dataset lies in the safety research of LLM agents, aiming to address the potential risks when agents execute malicious tasks.

- 1AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents英国人工智能安全研究所 · 2024年



