MEMORYARENA
收藏资源简介:
MEMORYARENA是由斯坦福大学领衔的多机构团队构建的智能体记忆评估基准,包含766条人工设计的跨会话任务,平均任务步长57步,生成超过40K tokens的推理轨迹。数据集通过商品兼容性聚类和数学物理问题构造,模拟真实场景中智能体需长期记忆并复用早期信息的需求。其核心价值在于填补现有评估对记忆-行动耦合能力的空白,适用于验证智能体在渐进式搜索、形式化推理等复杂场景中的记忆效用。
MEMORYARENA is an agent memory evaluation benchmark developed by a multi-institution team led by Stanford University. It includes 766 manually designed cross-session tasks, with an average task length of 57 steps and producing reasoning traces exceeding 40K tokens. Constructed via product compatibility clustering and mathematical physics problems, the dataset simulates real-world scenarios where agents need to utilize long-term memory to reuse early-stage information. Its core value lies in filling the gap in existing evaluations regarding memory-action coupling capabilities, and it is applicable to verifying the memory utility of agents in complex scenarios such as progressive search and formal reasoning.



