AdvancedIF
收藏资源简介:
AdvancedIF是一个新的基准数据集,包含超过1600个提示和专家策划的量表,旨在评估大型语言模型在以下方面的能力:复杂指令跟随(每个提示包含6个以上的指令,包括格式、风格、结构、长度、负面约束和条件间指令的组合)、多轮指令跟随(能够遵循从前一个环节携带过来的指令),以及系统提示的可控性。
AdvancedIF is a novel benchmark dataset consisting of over 1,600 prompts and expert-curated evaluation scales, which is designed to assess the capabilities of large language models (LLMs) in three key aspects: complex instruction following, where each prompt includes more than 6 instructions combining format, style, structure, length requirements, negative constraints, and conditional inter-instructions; multi-turn instruction following, referring to the ability to follow instructions carried over from prior interaction rounds; and the controllability of system prompts.
数据集概述
基本信息
- 许可证:CC BY-NC 4.0
- 语言:英语
- 标签:指令遵循、多轮对话、大语言模型、基于评分标准
数据集简介
AdvancedIF是一个包含超过1,600个提示和专家设计的评分标准的新基准,旨在评估大语言模型在以下方面的能力:
- 复杂指令遵循:每个提示包含6个以上指令,结合了格式、风格、结构、长度、否定约束、拼写和条件间指令;
- 多轮指令遵循:遵循先前对话中传递的指令的能力;
- 系统提示可引导性:遵循系统提示中指令的能力。
详细信息
- 论文链接:https://arxiv.org/abs/2511.10507
- 评估脚本:https://github.com/facebookresearch/AdvancedIF
数据划分
- 测试集:1,645个样本




