OpsEval
收藏资源简介:
OpsEval是一个专为大型语言模型设计的全面IT运营基准套件。该数据集由中国科学院和清华大学等机构合作创建,包含7184个多选题和1736个问答格式的问题,涵盖英语和中文。OpsEval旨在评估LLMs在各种关键场景中的能力水平,通过专家评审确保评估的可信度。数据集的应用领域包括根因分析、运维脚本生成和警报信息汇总,旨在优化专为IT运营定制的LLMs。
OpsEval is a comprehensive IT operations benchmark suite specifically designed for large language models (LLMs). Developed through collaboration between institutions including the Chinese Academy of Sciences and Tsinghua University, this dataset comprises 7,184 multiple-choice questions and 1,736 question-answering formatted questions, covering both English and Chinese languages. OpsEval aims to evaluate the competency levels of LLMs across various critical IT operational scenarios, with expert reviews conducted to ensure the credibility of the assessment framework. Its application areas include root cause analysis, operation and maintenance script generation, and alert information summarization, with the ultimate goal of optimizing LLMs tailored for IT operational scenarios.




