Auto-SLURP
收藏资源简介:
Auto-SLURP是一个用于评估基于大型语言模型(LLMs)的多智能体框架的数据集,旨在测试智能个人助手的性能。该数据集基于原始的SLURP数据集,通过重新标记数据并整合模拟服务器和外部服务进行了扩展。它涵盖了语言理解、任务执行和响应生成等方面的评估,并且包含了一系列任务领域,如日历管理、媒体播放、交通调度和信息检索等。Auto-SLURP旨在解决当前缺乏专门用于评估多智能体框架性能的基准数据集的问题,为研究者提供了全面和灵活的评估平台。
Auto-SLURP is a dataset for evaluating large language model (LLM)-based multi-agent frameworks, specifically designed to test the performance of intelligent personal assistants. It is extended from the original SLURP dataset through data relabeling and the integration of simulated servers and external services. The dataset covers evaluations across language understanding, task execution and response generation, and includes a series of task domains such as calendar management, media playback, traffic scheduling and information retrieval. Auto-SLURP aims to fill the current gap in benchmark datasets dedicated to evaluating the performance of multi-agent frameworks, providing researchers with a comprehensive and flexible evaluation platform.




