AgentSense
收藏资源简介:
AgentSense是由复旦大学研究人员创建的一个用于评估语言模型社会智能的基准数据集。该数据集包含1225个多样化的社会场景,这些场景从大量剧本中提取,确保了场景和社交目标的多样性和现实性。数据集的创建过程采用了自下而上的方法,通过提取剧本中的场景模板并合成角色来多样化场景。AgentSense主要用于评估语言模型在复杂社会互动中的目标完成和隐含推理能力,旨在解决语言模型在复杂社交场景中的表现问题。
AgentSense is a benchmark dataset created by researchers at Fudan University for evaluating the social intelligence of language models. It contains 1,225 diverse social scenarios extracted from a large number of scripts, which ensures the diversity and realism of both the scenarios and their corresponding social goals. The dataset was constructed using a bottom-up methodology, by extracting scenario templates from scripts and synthesizing characters to diversify the scenarios. AgentSense is primarily used to assess the capability of language models in accomplishing goals and performing implicit reasoning during complex social interactions, aiming to address the performance limitations of language models in complex social scenarios.

- 1AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios复旦大学 · 2024年



