BIG-Bench Extra Hard (BBEH)
收藏资源简介:
BIG-Bench Extra Hard (BBEH)是由谷歌DeepMind创建的数据集,旨在通过替代BIG-Bench Hard (BBH)中的每个任务来测试模型的一般推理能力。BBEH中的每个新任务都是在BBH的相应任务的基础上构建的,它们在相似的推理领域中测试类似的或更多的技能,但难度更大。该数据集保留了BBH的高多样性,并包含了200个问题/任务,除了Disambiguation QA任务有120个问题。BBEH旨在提供一个更准确的衡量模型一般推理能力的指标,挑战当前最先进的模型。
BIG-Bench Extra Hard (BBEH) is a dataset developed by Google DeepMind, which aims to test models' general reasoning abilities by replacing each task in BIG-Bench Hard (BBH) with newly constructed tasks. Each new task in BBEH is built upon its corresponding task in BBH, testing similar or enhanced skills within the same reasoning domains but with considerably higher difficulty. This dataset retains the high diversity of BBH, and contains 200 questions/tasks, with the exception of the Disambiguation QA task which includes 120 questions. BBEH is intended to provide a more accurate metric for evaluating models' general reasoning capabilities, posing challenges to current state-of-the-art models.

- 1BIG-Bench Extra Hard谷歌DeepMind · 2025年



