OmniMatBench
收藏资源简介:
OmniMatBench是由中国科学技术大学、上海人工智能实验室等机构联合构建的多模态材料科学推理基准数据集,旨在系统评估人工智能模型在跨学科材料科学领域的知识应用与工程推理能力。该数据集包含3,171条专家精心策划的问答与计算问题,覆盖19个材料学子领域,数据来源于经典材料学知识资源,并经过多阶段专家验证流程确保科学严谨性。数据集构建过程采用知识-结构-加工-应用(KSPA)框架,通过专家提取、人工标注与自动验证相结合的方式,形成包含问题、证明记录和评分要点的结构化数据。该数据集主要应用于评估多模态大语言模型在材料科学领域的推理能力,解决现有基准在工程应用、公式计算和跨领域知识整合方面的不足,为开发可靠的材料科学研究助手提供关键基准。
OmniMatBench is a multimodal materials science reasoning benchmark dataset jointly constructed by institutions including the University of Science and Technology of China and the Shanghai AI Laboratory, aiming to systematically evaluate the knowledge application and engineering reasoning capabilities of AI models across interdisciplinary materials science fields. This dataset contains 3,171 expert-curated question-answering and computational questions covering 19 sub-fields of materials science. Its data is sourced from classic materials science knowledge resources, and has undergone a multi-stage expert validation process to ensure scientific rigor. The dataset construction follows the Knowledge-Structure-Processing-Application (KSPA) framework, combining expert extraction, manual annotation and automatic verification to form structured data including questions, proof records and scoring rubrics. This dataset is primarily used to evaluate the reasoning capabilities of multimodal large language models in the materials science domain, addressing the shortcomings of existing benchmarks in engineering applications, formula-based calculations and cross-domain knowledge integration, providing a critical benchmark for developing reliable materials science research assistants.
数据集名称
OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields
核心概述
OmniMatBench 是一个针对材料科学领域设计的多模态推理基准测试集,旨在评估多模态大语言模型(MLLMs)在材料科学中的综合推理能力,填补了现有基准在从材料知识到应用推理方面的空白。
数据集组成
- 规模与来源:包含 3,171 个由专家精选的问答与计算问题。
- 覆盖范围:横跨 19 个材料科学子领域,涵盖四大类别:
- 基础材料知识
- 结构材料与工程材料
- 材料加工与制造
- 功能材料与应用材料
评估结果与发现
- 模型表现:对 13 个开源和闭源 MLLMs 进行了评估,最佳模型的总体得分仅为 0.372,揭示了当前模型在材料科学推理方面存在显著差距。
- 具体问题:
- 不同子领域间的表现差异很大。
- 模型表现出固定的推理启发式策略,材料知识分布不均。
- 在辅以公式、检索和代码的情况下,对高层知识的应用能力有限。
意义与目的
该基准为评估当前多模态大语言模型在材料科学研究中的能力与局限性提供了关键洞察,并为开发可靠的 AI 助手奠定了基础。




