CritPt (Complex Research using Integrated Thinking -Physics Test)
收藏资源简介:
CritPt 数据集旨在评估大型语言模型(LLM)在现代物理研究中的推理能力。该数据集由 71 个复合研究挑战组成,模拟了初级研究项目的全规模,并分解为 190 个更简单的检查点任务,以提供更细致的洞察。所有问题都是由 50 多位活跃的物理研究人员根据他们自己的研究新创建的,每个问题都经过精心策划,以接受猜测抵抗和机器可验证的答案,并通过高度定制的自动化评分管道进行评估。CritPt 为评估 LLM 在现实物理研究工作流程中的价值提供了一个强大的框架,这是定义 AI 在科学发现中未来角色的一个基本但尚未充分探索的组成部分。
The CritPt dataset is designed to evaluate the reasoning capabilities of Large Language Models (LLMs) in modern physics research. The dataset consists of 71 complex research challenges, each simulating the full scale of an early-career research project, and is broken down into 190 simpler checkpoint tasks to provide more granular insights. All questions are newly created by over 50 active physics researchers based on their own original research. Each problem is meticulously curated to resist guesswork and paired with machine-verifiable answers, and is evaluated via highly customized automated scoring pipelines. CritPt provides a robust framework for evaluating the value of LLMs in real-world physics research workflows, a fundamental yet under-explored component in defining the future role of AI in scientific discovery.

- 1Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research BenchmarkArgonne National Laboratory, University of Illinois Urbana-Champaign, Virginia Tech, Ohio State University, Northeastern University, Columbia University, University of Florida, Perimeter Institute for Theoretical Physics, University of Waterloo, University of Connecticut, University of Cologne, The Chinese University of Hong Kong, Harvard University, ETH Zurich, Paul Scherrer Institute, University of Washington Seattle, University of Chicago, University of Colorado Boulder, University of California Los Angeles, University of California San Diego, University of Tennessee Knoxville, National Institute of Theory and Mathematics in Biology, Princeton University · 2025年



