BioMysteryBench-preview
收藏资源简介:
BioMysteryBench是一个由Anthropic创建的基准测试公开样本,包含5个问题。数据集主要包含两部分:问题描述文件(problems.csv或problems.parquet)和每个问题对应的数据文件(data.zip或data/<id>.zip)。问题描述文件中每行代表一个问题,包含问题标识符(id)、任务提示(question)、评分标准(answer_rubric,含预期答案)、允许访问的网络域(allowed_domains)以及人类是否可解的标记(human_solvable)。数据文件需解压到工作目录中使用。该数据集适用于评估模型在问题解决任务上的性能。完整基准测试需申请访问权限。
BioMysteryBench is a public sample of a benchmark created by Anthropic, containing 5 problems. The dataset mainly consists of two parts: the problem description file (problems.csv or problems.parquet) and the corresponding data files for each problem (data.zip or data/<id>.zip). Each line in the problem description file represents a problem and includes the following fields: problem identifier (id), task prompt (question), grading rubric (answer_rubric, including expected answers), allowed web domains (allowed_domains), and a flag indicating whether the problem is human-solvable (human_solvable). The data files need to be extracted to the working directory for use. This dataset is suitable for evaluating model performance on problem-solving tasks. Full benchmark access requires permission.
BioMysteryBench (公开样本) 数据集详情
基本信息
- 创建者: Anthropic
- 数据集类型: 基准测试(Benchmark)公开样本
- 样本规模: 包含5个问题
数据内容
数据集包含以下文件:
1. problems.csv / problems.parquet
每个问题对应一行数据,包含以下字段:
- id: 问题标识符
- question: 向模型展示的任务提示
- answer_rubric: 评分标准(包含预期答案)
- allowed_domains: 求解环境可以访问的网络域名
- human_solvable: 标记是否至少有一位人类基准测试者解决了该问题(
yes表示可解决,no表示无人解决)
2. data.zip(样本)或 data/<id>.zip(完整集)
每个问题的数据文件,需要在求解前解压到工作目录中。
完整数据集获取
如需访问完整基准测试数据集,请通过以下链接申请权限: https://huggingface.co/datasets/Anthropic/BioMysteryBench-full




