遇见数据集

open-llm-leaderboard/details_TheBloke__robin-13B-v2-fp16

收藏
Hugging Face2023-08-27 更新2024-03-04 收录
官方服务:

资源简介:

该数据集是在模型TheBloke/robin-13B-v2-fp16的评估运行期间自动创建的,用于在Open LLM Leaderboard上进行评估。数据集由61个配置组成,每个配置对应一个评估任务。数据集是从1次运行中创建的,每次运行可以在每个配置中找到特定的分割,分割名称使用运行的时间戳。train分割始终指向最新的结果。此外,一个名为results的配置存储了所有运行的聚合结果,并用于计算和显示Open LLM Leaderboard上的聚合指标。

This dataset was automatically created during the evaluation run of the model TheBloke/robin-13B-v2-fp16 for evaluation on the Open LLM Leaderboard. It consists of 61 configurations, each corresponding to one evaluation task. This dataset was generated from a single run, where each configuration contains a specific data split, and the split name uses the timestamp of the run. The "train" split always points to the most recent results. Additionally, a configuration named "results" stores the aggregated results across all runs, and is used to calculate and display the aggregated metrics on the Open LLM Leaderboard.

提供机构:
open-llm-leaderboard
原始信息汇总

数据集概述

数据集简介

该数据集是在评估模型 TheBloke/robin-13B-v2-fp16 在 Open LLM Leaderboard 上的自动创建的。数据集包含 61 个配置,每个配置对应一个评估任务。

数据集结构

数据集由 1 次运行创建,每个运行可以在每个配置中找到特定的分割,分割名称使用运行的时间戳。"train" 分割始终指向最新的结果。

额外配置

一个额外的配置 "results" 存储所有运行的聚合结果,用于计算和显示 Open LLM Leaderboard 上的聚合指标。

数据加载示例

python from datasets import load_dataset data = load_dataset("open-llm-leaderboard/details_TheBloke__robin-13B-v2-fp16", "harness_truthfulqa_mc_0", split="train")

最新结果

这些是最新结果(来自 2023-07-31T15:48:06.598529 运行)的示例: python { "all": { "acc": 0.49056004249413854, "acc_stderr": 0.034895228964178376, "acc_norm": 0.49452555601900244, "acc_norm_stderr": 0.03487806793899599, "mc1": 0.34149326805385555, "mc1_stderr": 0.016600688619950826, "mc2": 0.5063100731922137, "mc2_stderr": 0.014760623429029368 }, "harness|arc:challenge|25": { "acc": 0.5401023890784983, "acc_stderr": 0.01456431885692485, "acc_norm": 0.5648464163822525, "acc_norm_stderr": 0.014487986197186045 }, "harness|hellaswag|10": { "acc": 0.5945030870344553, "acc_stderr": 0.004899845087183104, "acc_norm": 0.8037243576976698, "acc_norm_stderr": 0.003963677261161229 }, "harness|hendrycksTest-abstract_algebra|5": { "acc": 0.33, "acc_stderr": 0.04725815626252606, "acc_norm": 0.33, "acc_norm_stderr": 0.04725815626252606 }, "harness|hendrycksTest-anatomy|5": { "acc": 0.4666666666666667, "acc_stderr": 0.043097329010363554, "acc_norm": 0.4666666666666667, "acc_norm_stderr": 0.043097329010363554 }, "harness|hendrycksTest-astronomy|5": { "acc": 0.4868421052631579, "acc_stderr": 0.04067533136309173, "acc_norm": 0.4868421052631579, "acc_norm_stderr": 0.04067533136309173 }, "harness|hendrycksTest-business_ethics|5": { "acc": 0.45, "acc_stderr": 0.05, "acc_norm": 0.45, "acc_norm_stderr": 0.05 }, "harness|hendrycksTest-clinical_knowledge|5": { "acc": 0.4679245283018868, "acc_stderr": 0.03070948699255655, "acc_norm": 0.4679245283018868, "acc_norm_stderr": 0.03070948699255655 }, "harness|hendrycksTest-college_biology|5": { "acc": 0.4722222222222222, "acc_stderr": 0.04174752578923185, "acc_norm": 0.4722222222222222, "acc_norm_stderr": 0.04174752578923185 }, "harness|hendrycksTest-college_chemistry|5": { "acc": 0.24, "acc_stderr": 0.04292346959909284, "acc_norm": 0.24, "acc_norm_stderr": 0.04292346959909284 }, "harness|hendrycksTest-college_computer_science|5": { "acc": 0.38, "acc_stderr": 0.04878317312145632, "acc_norm": 0.38, "acc_norm_stderr": 0.04878317312145632 }, "harness|hendrycksTest-college_mathematics|5": { "acc": 0.31, "acc_stderr": 0.04648231987117317, "acc_norm": 0.31, "acc_norm_stderr": 0.04648231987117317 }, "harness|hendrycksTest-college_medicine|5": { "acc": 0.44508670520231214, "acc_stderr": 0.03789401760283646, "acc_norm": 0.44508670520231214, "acc_norm_stderr": 0.03789401760283646 }, "harness|hendrycksTest-college_physics|5": { "acc": 0.17647058823529413, "acc_stderr": 0.0379328118530781, "acc_norm": 0.17647058823529413, "acc_norm_stderr": 0.0379328118530781 }, "harness|hendrycksTest-computer_security|5": { "acc": 0.62, "acc_stderr": 0.048783173121456316, "acc_norm": 0.62, "acc_norm_stderr": 0.048783173121456316 }, "harness|hendrycksTest-conceptual_physics|5": { "acc": 0.4, "acc_stderr": 0.03202563076101735, "acc_norm": 0.4, "acc_norm_stderr": 0.03202563076101735 }, "harness|hendrycksTest-econometrics|5": { "acc": 0.30701754385964913, "acc_stderr": 0.04339138322579861, "acc_norm": 0.30701754385964913, "acc_norm_stderr": 0.04339138322579861 }, "harness|hendrycksTest-electrical_engineering|5": { "acc": 0.4068965517241379, "acc_stderr": 0.04093793981266237, "acc_norm": 0.4068965517241379, "acc_norm_stderr": 0.04093793981266237 }, "harness|hendrycksTest-elementary_mathematics|5": { "acc": 0.25925925925925924, "acc_stderr": 0.02256989707491841, "acc_norm": 0.25925925925925924, "acc_norm_stderr": 0.02256989707491841 }, "harness|hendrycksTest-formal_logic|5": { "acc": 0.31746031746031744, "acc_stderr": 0.04163453031302859, "acc_norm": 0.31746031746031744, "acc_norm_stderr": 0.04163453031302859 }, "harness|hendrycksTest-global_facts|5": { "acc": 0.38, "acc_stderr": 0.04878317312145632, "acc_norm": 0.38, "acc_norm_stderr": 0.04878317312145632 }, "harness|hendrycksTest-high_school_biology|5": { "acc": 0.49032258064516127, "acc_stderr": 0.028438677998909558, "acc_norm": 0.49032258064516127, "acc_norm_stderr": 0.028438677998909558 }, "harness|hendrycksTest-high_school_chemistry|5": { "acc": 0.32019704433497537, "acc_stderr": 0.032826493853041504, "acc_norm": 0.32019704433497537, "acc_norm_stderr": 0.032826493853041504 }, "harness|hendrycksTest-high_school_computer_science|5": { "acc": 0.47, "acc_stderr": 0.05016135580465919, "acc_norm": 0.47, "acc_norm_stderr": 0.05016135580465919 }, "harness|hendrycksTest-high_school_european_history|5": { "acc": 0.6303030303030303, "acc_stderr": 0.037694303145125674, "acc_norm": 0.6303030303030303, "acc_norm_stderr": 0.037694303145125674 }, "harness|hendrycksTest-high_school_geography|5": { "acc": 0.5606060606060606, "acc_stderr": 0.03536085947529479, "acc_norm": 0.5606060606060606, "acc_norm_stderr": 0.03536085947529479 }, "harness|hendrycksTest-high_school_government_and_politics|5": { "acc": 0.6683937823834197, "acc_stderr": 0.03397636541089118, "acc_norm": 0.6683937823834197, "acc_norm_stderr": 0.03397636541089118 }, "harness|hendrycksTest-high_school_macroeconomics|5": { "acc": 0.44871794871794873, "acc_stderr": 0.025217315184846482, "acc_norm": 0.44871794871794873, "acc_norm_stderr": 0.025217315184846482 }, "harness|hendrycksTest-high_school_mathematics|5": { "acc": 0.23333333333333334, "acc_stderr": 0.02578787422095932, "acc_norm": 0.23333333333333334, "acc_norm_stderr": 0.02578787422095932 }, "harness|hendrycksTest-high_school_microeconomics|5": { "acc": 0.4411764705882353, "acc_stderr": 0.0322529423239964, "acc_norm": 0

搜集汇总
数据集介绍
构建方式
在开放大语言模型排行榜(Open LLM Leaderboard)的评估体系中,该数据集专为记录模型TheBloke/robin-13B-v2-fp16的评测结果而自动生成。其构建基于一次完整的评估运行,涵盖了61个不同的评测任务配置,每个配置对应一个特定的基准测试。数据集的每个配置内均以运行时间戳命名分割,其中“train”分割始终指向最新一次的评测结果,而额外的“results”配置则聚合了所有运行的总体指标,用于在排行榜上计算和展示综合性能。
使用方法
研究者可通过Hugging Face的datasets库便捷地调用该数据集。具体使用时,需指定目标任务的配置名称(如“harness_truthfulqa_mc_0”)及所需的分割(如“train”)。例如,执行`load_dataset("open-llm-leaderboard/details_TheBloke__robin-13B-v2-fp16", "harness_truthfulqa_mc_0", split="train")`即可加载最新一次的TruthfulQA评测详情。此外,通过访问“results”配置或直接查阅存储库中的JSON文件,可获取模型在所有任务上的综合性能汇总,便于进行横向对比与深入研究。
背景与挑战
背景概述
大语言模型(LLM)的迅猛发展催生了对其性能进行系统化评估的迫切需求。在此背景下,HuggingFace团队于2023年发起了Open LLM Leaderboard项目,旨在通过标准化基准测试集,为开源社区提供一个透明、可复现的模型性能比较平台。该数据集正是围绕TheBloke/robin-13B-v2-fp16这一模型在Leaderboard上的评估过程生成的,记录了其在61个任务配置上的详细表现,涵盖ARC挑战、HellaSwag、MMLU(涵盖数学、医学、法律等57个学科)及TruthfulQA等核心评测。该数据集由Clementine等人创建,其核心研究问题在于如何通过细粒度的任务结果追踪,揭示模型在不同认知维度上的优势与短板。该数据集不仅为开发者提供了评估模型改进效果的量化依据,也推动了LLM评估方法论向更严谨、更全面的方向发展,对开源社区构建可信赖的模型评价体系具有里程碑意义。
当前挑战
该数据集所应对的核心挑战在于如何高效、公正地评估大语言模型在多样化任务上的真实能力。具体而言,1)领域问题层面,LLM的评估长期面临单一指标无法反映模型泛化性的困境,例如模型可能在常识推理(HellaSwag)上表现优异,却在对抗性问答(ARC Challenge)或事实一致性(TruthfulQA)上显著退化,而该数据集通过覆盖数理、人文、医学等57个学科的多选题测试,系统性地暴露了这一不均衡性;2)构建过程层面,挑战包括如何确保不同任务间的评分标准统一(如准确率与归一化准确率的权衡)、如何处理多次运行(run)结果的版本管理与时间戳对齐,以及如何将61个独立配置的庞杂结果以可复用的结构化格式(Parquet文件)存储,同时保证数据加载接口的简洁性与向后兼容性。
常用场景
经典使用场景
该数据集专为评估大语言模型在多样化自然语言理解任务中的表现而设计,广泛应用于模型性能的标准化基准测试。其存储了TheBloke/robin-13B-v2-fp16模型在61个配置上的详细评估结果,涵盖ARC挑战集、HellaSwag常识推理、TruthfulQA事实性问答以及涵盖57个学科的MMLU基准等经典任务。研究者可通过加载特定配置与分割,复现模型在每项任务上的精确得分与误差范围,从而系统性地量化和比较模型在推理、知识掌握和真实性等方面的能力。
解决学术问题
该数据集解决了大语言模型评估中缺乏细粒度、可复现的标准化记录这一关键问题。通过结构化存储每次运行的完整结果,它使得研究者能够精确追溯模型在不同子任务上的表现差异,从而深入分析其优势与短板。例如,MMLU基准的细粒度学科得分揭示了模型在专业领域知识上的局限性,而TruthfulQA的MC1与MC2指标则量化了模型生成真实内容的倾向。这些数据为诊断模型偏差、评估知识边界和指导模型优化提供了坚实依据。
实际应用
在工业界与学术界的实际应用中,该数据集为模型选型与迭代提供了关键参考。开发者可依据其在ARC挑战上的推理准确率或HellaSwag上的常识理解得分,筛选出最适合特定场景的模型版本。此外,数据集的标准化格式便于集成到自动化评估流水线中,支持持续监控模型在部署前后的性能退化。其公开的细粒度结果也构成了社区驱动的模型排行榜基础,帮助从业者快速定位具有竞争力的预训练模型。
数据集最近研究
最新研究方向
在大规模语言模型评估领域,开放大语言模型排行榜(Open LLM Leaderboard)已成为衡量模型性能的标杆性平台。针对TheBloke/robin-13B-v2-fp16模型的评估数据集,当前研究前沿聚焦于通过多维度、细粒度的任务配置(涵盖ARC挑战、HellaSwag、TruthfulQA及涵盖57个学科的MMLU基准)来系统性地剖析模型的推理能力、知识广度与诚实性。这一方向与业界对模型透明度和可复现性的追求紧密相连,尤其是通过标准化评估框架(如lm-evaluation-harness)自动生成详尽的性能指标,为社区提供了横向对比的权威依据。该数据集的意义在于,它不仅揭示了13B参数级别模型在复杂认知任务上的潜力与局限,更推动了评估范式的演进——从单一准确率转向对错误率、标准化得分及多选一致性等深层特质的挖掘,从而为下一代语言模型的优化指明了关键瓶颈。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务