terminal_bench_2_a1_inferredbugs_20260627_013005-traces
收藏资源简介:
该数据集记录了智能体或AI模型在特定任务中的交互对话及其执行结果,适用于对话系统评估和任务导向的交互分析。每个样本包含多轮对话(其中每轮有角色和内容)、智能体标识、模型信息(包括模型名称和提供商)、任务描述、运行日期、运行标识(如run_id、trial_name、episode)、任务结果、验证器输出以及数据来源追踪。数据集规模为832个训练样本,总大小约69MB,以结构化格式存储,支持对模型或智能体在对话任务中的表现进行详细分析。
This dataset records interactive dialogues and execution results of agents or AI models in specific tasks, suitable for dialogue system evaluation and task-oriented interaction analysis. Each sample includes multiple rounds of conversations (with roles and content per round), agent identifiers, model information (including model name and provider), task description, run date, run identifiers (such as run_id, trial_name, episode), task results, verifier output, and trace source for data provenance. The dataset comprises 832 training samples, with a total size of approximately 69MB, stored in a structured format to support detailed analysis of model or agent performance in dialogue tasks.
- 数据集名称:
laion/terminal_bench_2_a1_inferredbugs_20260627_013005-traces - 数据集大小:下载大小约18.95 MB,数据集总大小约69.46 MB
- 数据分割:仅包含训练集(
train),共832个样本 - 数据特征:
conversations:对话列表,包含content(字符串类型)和role(字符串类型)字段agent:代理名称(字符串类型)model:模型名称(字符串类型)model_provider:模型提供商(字符串类型)date:日期(字符串类型)task:任务描述(字符串类型)episode:回合编号(字符串类型)run_id:运行ID(字符串类型)trial_name:试验名称(字符串类型)result:结果(字符串类型)verifier_output:验证器输出(字符串类型)trace_source:追踪来源(字符串类型)
- 配置:默认配置(
default),数据文件路径为data/train-*




