ToolMind
收藏资源简介:
ToolMind是一个大型的开源工具使用数据集,包含推理轨迹,旨在提高代理型大型语言模型的推理和工具调用能力。该数据集通过图采样的方式合成超过160k的回合,涉及超过20k的工具。数据集通过将功能组织为图结构中的节点,采样图上的路径来构建复杂高质量的用户意图,并通过多智能体方式合成轨迹。每个轨迹回合都通过思考模型进行推理回答和正确性过滤,只保留正确且有价值的回合。在Tau-bench、Tau2-bench和BFCL-v4代理基准测试中,使用ToolMind进行微调的模型表现出显著的改进。
ToolMind is a large-scale open-source tool-use dataset containing reasoning trajectories, designed to enhance the reasoning and tool invocation capabilities of agent-based large language models. This dataset synthesizes over 160k interaction turns via graph sampling, involving more than 20k tools. It constructs complex and high-quality user intents by organizing functionalities as nodes in a graph structure and sampling paths on the graph, then synthesizes reasoning trajectories through a multi-agent framework. Each trajectory turn is subjected to reasoning response generation and correctness filtering via a dedicated reasoning model, retaining only valid and valuable turns. Models fine-tuned with ToolMind have exhibited substantial performance improvements across the Tau-bench, Tau2-bench, and BFCL-v4 agent benchmark tests.
ToolMind数据集概述
数据集基本信息
- 许可证: Apache-2.0
- 任务类别: 文本生成
- 语言: 英语
- 标签: 函数调用、工具调用、合成数据
- 官方名称: ToolMind
数据集构成
数据文件配置
- 配置名称: test
- 数据文件分割:
- graph-based-synthetic-data: data/graphsyn.jsonl
- xlam-function-calling-60k: data/xlam-function-calling-60k-query.jsonl
- When2Call: data/When2Call-query.jsonl
- glaive-function-calling-v2: data/glaive-function-calling-v2-query.jsonl
- ToolACE: data/ToolACE-query.jsonl
- BUTTONInstruct: data/BUTTONInstruct-query.jsonl
- APIGen-MT-5k: data/APIGen-MT-5k-query.jsonl
- Tau-bench training set: data/tau-train-query.jsonl
数据集规模与特点
- 包含超过160,000轮对话
- 基于20,000多个工具合成
- 通过图采样和多智能体模拟构建复杂用户意图
- 包含推理轨迹和质量过滤
合成流程
数据收集与增强
- 从开源数据集收集函数:xlam-function-calling-60k、glaive-function-calling-v2、ToolACE
- 使用语言模型完善函数描述和参数类型
- 使用Conan-embedding-v1进行向量化
图构建
- 将函数表示为图中的节点
- 基于输入输出参数语义相似度构建边
- 引入随机边构建增加拓扑多样性
随机游走采样
- 使用长度5-20的随机游走采样函数链
- 限制节点访问次数避免过采样
多智能体轨迹合成
- 三个模型分别模拟用户、助手和函数
- 使用思维模型进行质量过滤和错误校正
- 仅保留有效轮次
混合训练数据
整合的开放源代码数据包括:
- xlam-function-calling-60k
- When2Call
- glaive-function-calling-v2
- ToolACE
- BUTTONInstruct
- APIGen-MT-5k
- Tau-bench训练集
性能表现
整体性能
在Tau2-bench和BFCL-v4评估中,使用ToolMind微调的模型相比基线有显著提升:
- Qwen3-8B模型在各项指标上均有明显改进
- Qwen3-14B模型在多数指标上表现更优
消融研究
- 增强的开放源代码数据(20万)在多个指标上带来提升
- 合成数据(16万)在特定领域表现突出
- 完整ToolMind数据集(36万)综合性能最佳
训练注意事项
- 数据基于助手消息进行分割
- 训练时仅对每个样本的最后一条助手消息计算损失
局限性
- 模型输出可能存在意外内容
- 可能生成包含偏见或歧视的有害内容
- 不承担传播不当信息造成的后果责任
引用与联系
- 使用请引用此Huggingface项目
- 联系方式:nanbeige@126.com




