TMGBENCH
收藏资源简介:
TMGBENCH是由哈尔滨工业大学和香港大学联合创建的一个用于评估大型语言模型(LLMs)战略推理能力的系统性游戏基准。该数据集涵盖了144种基于Robinson-Goforth拓扑结构的2×2游戏类型,每种类型包含多个实例,并通过合成数据生成技术创建了多样化的故事背景游戏。数据集的创建过程包括主题控制和人工审查,确保数据的高质量和多样性。TMGBENCH旨在通过复杂的序列、并行和嵌套游戏结构,评估LLMs在多层次决策中的战略推理能力,解决现有基准在游戏类型覆盖、数据泄露和可扩展性方面的不足。
TMGBENCH is a systematic game-based benchmark co-developed by Harbin Institute of Technology and The University of Hong Kong for evaluating the strategic reasoning capabilities of Large Language Models (LLMs). This dataset encompasses 144 categories of 2×2 games based on the Robinson-Goforth topology, with multiple instances for each category, and generates diverse story-driven games using synthetic data generation techniques. The dataset construction process incorporates theme control and manual review to ensure high data quality and diversity. TMGBENCH aims to evaluate the strategic reasoning abilities of LLMs in multi-level decision-making through complex sequential, parallel, and nested game structures, addressing the limitations of existing benchmarks in terms of game type coverage, data leakage, and scalability.




