Quantitative Reasoning with Data (QRDATA)
收藏资源简介:
数据集QRDATA由北京大学王选计算机研究所和加州大学洛杉矶分校计算机科学系共同创建,包含411个问题,旨在评估大型语言模型在处理真实世界数据时的统计和因果推理能力。数据集内容涵盖从教科书、在线学习材料和学术论文中精心挑选的数据表,用于评估模型的自然语言推理、基于程序的推理和代理推理方法。数据集创建过程中,确保所有问题与数据匹配合理,通过手动构建确保数据集的质量。该数据集主要应用于评估和提升模型在数据基础上的高级定量推理能力,特别是在统计和因果推理方面。
The QRDATA dataset was co-developed by the Wangxuan Institute of Computer Technology at Peking University and the Department of Computer Science at the University of California, Los Angeles. It comprises 411 questions, aimed at evaluating the statistical and causal reasoning capabilities of large language models (LLMs) when handling real-world data. The dataset covers data tables carefully selected from textbooks, online learning materials, and academic papers, and is utilized to assess models' natural language reasoning, program-based reasoning, and AI agent-based reasoning methods. During the dataset's development, strict alignment between all questions and their corresponding data was ensured, and the dataset's quality was validated through manual curation. This dataset is primarily used to evaluate and improve models' advanced quantitative reasoning abilities based on real-world data, with a particular focus on statistical and causal reasoning scenarios.

- 1Are LLMs Capable of Data-based Statistical and Causal Reasoning? Benchmarking Advanced Quantitative Reasoning with Data北京大学王选计算机研究所 · 2024年



