Episodic Memory Benchmark
收藏资源简介:
该数据集由华为技术有限公司的研究团队创建,旨在评估大型语言模型(LLMs)在情景记忆任务中的表现。数据集包含11个不同规模和多样性的子集,涵盖了丰富的时间和空间上下文信息,涉及特定实体和事件的详细描述。数据集的创建过程受到认知科学的启发,通过结构化方法生成合成的情景记忆任务,确保数据的连贯性和可控性。该数据集的应用领域主要集中在提升LLMs的情景记忆能力,解决其在处理复杂时空关系和多个相关事件时的不足,从而增强模型的推理能力和事实准确性。
This dataset was created by the research team of Huawei Technologies Co., Ltd., aiming to evaluate the performance of Large Language Models (LLMs) on episodic memory tasks. It comprises 11 subsets with varying scales and diversity, which contain rich temporal and spatial contextual information as well as detailed descriptions of specific entities and events. The development of this dataset is inspired by cognitive science, and synthetic episodic memory tasks are generated via structured methods to ensure the coherence and controllability of the data. The main application scenarios of this dataset focus on enhancing the episodic memory capabilities of LLMs, addressing their limitations in handling complex spatiotemporal relationships and multiple related events, thereby improving the model's reasoning ability and factual accuracy.

- 1Episodic Memories Generation and Evaluation Benchmark for Large Language Models华为技术有限公司 · 2025年



