拒绝分类数据集
收藏资源简介:
拒绝分类数据集是由慕尼黑工业大学等机构的研究人员创建的,旨在分析大型语言模型(LLMs)在拒绝用户指令时的行为。该数据集包含8600个真实实例,由人工标注,并结合了合成数据,总计超过700万条拒绝实例。数据集的创建过程涉及对公开的IFT和RLHF数据集进行标注,并生成了多种拒绝类别的合成数据。该数据集主要用于评估和改进LLMs的安全性和可靠性,特别是在减少幻觉和提升模型拒绝不当指令的能力方面。
The Refusal Classification Dataset was developed by researchers from institutions including the Technical University of Munich (TUM) and other relevant organizations, with the core objective of analyzing the behaviors of large language models (LLMs) when they refuse user instructions. This dataset comprises 8,600 manually annotated real-world instances, supplemented by synthetic data, resulting in a total of over 7 million refusal samples. The dataset construction process involves annotating publicly available IFT and RLHF datasets, as well as generating synthetic data across multiple refusal categories. This dataset is primarily employed to evaluate and improve the safety and reliability of LLMs, particularly in mitigating hallucinations and enhancing the models' ability to reject inappropriate user instructions.

- 1Cannot or Should Not? Automatic Analysis of Refusal Composition in IFT/RLHF Datasets and Refusal Behavior of Black-Box LLMs慕尼黑工业大学 · 2024年



