CIMemories
收藏资源简介:
CIMemories是一个用于评估大型语言模型在处理持久记忆信息时上下文完整性的合成基准数据集。它包含具有超过100个属性的用户配置文件,并与多样化的任务上下文配对,用于检测模型在信息流控制方面的性能。
CIMemories is a synthetic benchmark dataset for evaluating the contextual completeness of large language models (LLMs) when handling persistent memory information. It comprises user profiles with over 100 attributes, paired with diverse task contexts to assess the model's performance in information flow control.
CIMemories数据集概述
数据集基本信息
- 名称:CIMemories
- 语言:英语
- 许可证:CC BY-NC 4.0
- 数据规模:10K-100K样本
数据集描述
CIMemories是一个用于评估大型语言模型中持久记忆上下文完整性的组合基准。该数据集专注于评估LLMs在基于任务上下文适当控制记忆信息流方面的能力。
核心特征
- 使用包含每个用户100多个属性的合成用户配置文件
- 包含多样化的任务上下文
- 评估属性级别的信息泄露违规情况
- 分析违规行为在任务和运行中的累积效应
评估发现
- 前沿模型的属性级别违规率高达69%
- 违规率随任务数量增加而上升(从1个任务的0.1%到40个任务的9.6%)
- 相同提示执行5次时违规率达到25.1%
- 隐私意识提示无法解决根本问题
研究意义
揭示LLMs在上下文感知推理能力方面的基本限制,表明需要不仅仅是更好的提示或规模扩展,而是真正的情境感知能力。
引用信息
bibtex @misc{mireshghallah2025cimemoriescompositionalbenchmarkcontextual, title={CIMemories: A Compositional Benchmark for Contextual Integrity of Persistent Memory in LLMs}, author={Niloofar Mireshghallah and Neal Mangaokar and Narine Kokhlikyan and Arman Zharmagambetov and Manzil Zaheer and Saeed Mahloujifar and Kamalika Chaudhuri}, year={2025}, eprint={2511.14937}, archivePrefix={arXiv}, primaryClass={cs.CR}, url={https://arxiv.org/abs/2511.14937}, }




