TUTORBENCH
收藏资源简介:
TUTORBENCH是一个专为评估大型语言模型(LLM)辅导能力而设计的数据集和评估基准。该数据集由人类专家精心策划的1490个样本组成,涵盖高中和AP课程。样本来自三个常见的辅导任务:生成针对学生困惑的适应性解释、对学生作品的反馈和评估,以及通过有效的提示生成促进主动学习。为了应对辅导的内在复杂性,样本附有特定的评分标准,用于评估模型响应。TUTORBENCH使用可靠的细粒度自动评估方法,该方法使用LLM-judge和样本特定的评分标准。我们评估了16个前沿的LLM在TUTORBENCH上的表现,并详细分析了它们的性能和行为。结果表明,没有一种前沿的LLM能够达到超过56%的得分,显示出巨大的改进空间。我们发现LLM在展现全面辅导技能方面存在不足,所有前沿模型在与这些技能相关的评分标准上都不到60%的通过率。我们还发现,不同的模型家族展现出不同的优势和局限性:Claude模型在支持主动学习方面表现优于其他模型,而在其他两种使用案例中则落后。通过发布TUTORBENCH,我们提供了一个全面且未饱和的基准,以指导下一代AI的发展。
TUTORBENCH is a dataset and evaluation benchmark specifically designed for assessing the tutoring capabilities of Large Language Models (LLMs). The dataset consists of 1,490 expert-curated samples covering high school and AP courses. The samples originate from three common tutoring tasks: generating adaptive explanations tailored to students’ confusion, providing feedback and assessment on student work, and facilitating active learning through effective prompting. To address the inherent complexity of tutoring, each sample is paired with specific grading rubrics for evaluating model responses. TUTORBENCH employs a reliable fine-grained automatic evaluation method that leverages LLM-judge and sample-specific grading rubrics. We evaluated 16 state-of-the-art LLMs on TUTORBENCH and conducted a detailed analysis of their performance and behaviors. The results indicate that no state-of-the-art LLM achieves a score exceeding 56%, revealing significant room for improvement. We find that LLMs lack proficiency in comprehensive tutoring skills: all state-of-the-art models achieve a pass rate of less than 60% on the grading rubrics related to these skills. We also discover that different model families exhibit distinct advantages and limitations: Claude models outperform others in supporting active learning, yet lag behind in the other two use cases. By releasing TUTORBENCH, we provide a comprehensive and under-explored benchmark to guide the development of next-generation AI.




