工具调用(Tool Use / Function Calling)
一句话定义:模型输出一个结构化的”我要调用这个函数、参数是这些”的请求,由外部系统真正执行,再把结果喂回模型。这是模型从”说话”走向”做事”的接口。
1. 它包含四个可学习的决策
Toolformer(arXiv:2302.04761)最早把工具调用形式化:
“a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction.”
Toolformer 把「用工具」拆成了四个可学习的决策——调哪个 API、什么时候调、传什么参数、怎么把返回值融进后面的预测。这四个决策基本就是工具调用能力的全部内容。
| 决策 | 失败时会怎样 |
|---|---|
| 选哪个工具 | 用了不相关的工具 |
| 什么时候调 | 该调不调、不该调乱调 |
| 传什么参数 | 参数类型/取值错误 |
| 如何把结果并入后续生成 | 拿到结果却不会用 |
2. 它是训练出来的,不是参数涌出来的(最重要的一条)
ToolLLM论文(arXiv:2307.16789)的原话:
“they remain significantly limited in tool-use capabilities … The reason is that current instruction tuning largely focuses on basic language tasks but ignores the tool-use domain.”
开源模型在工具使用能力上明显受限,也就是「用外部工具(API)去完成人类指令」这件事。原因在于:当时的指令微调主要聚焦在基础语言任务上,把工具使用这个领域整个忽略了。
2026 年的实证支持这个判断——在 BFCL V4 榜单上(Last Updated 2026-04-12):
| 模型 | Overall Acc |
|---|---|
| Nanbeige4-3B-Thinking(3B) | 51.4 |
| xLAM-2-3b-fc-r(3B,专用工具模型) | 41.22 |
| Phi-4(14B) | 28.79 |
| Gemma-3-27b-it(27B) | 29.47 |
| Llama-3.3-70B(70B) | 31.9 |
| Llama-4-Scout-17B-16E | 28.13 |
3B 的模型把 70B 的模型甩在后面——参数量与工具调用能力不单调。这是”参数大 ≠ 工具调用强”最直接的证据。
3. 现在怎么测:BFCL 的重心已经转移
BFCL V4(2025-07 发布)的总分构成(官方逐字):
Overall Score = Agentic(40%) + Multi-Turn(30%) + Live(10%) + Non-Live(10%) + Hallucination(10%)
官方博客原话的含义:单轮函数调用已经饱和、无法区分前沿模型,所以榜单把 70% 的权重挪到了跨轮次保持状态、搜索、记忆、以及知道什么时候不该调用。
“The single hardest capability is abstention: the Hallucination category rewards a model for correctly refusing to call a function when no available tool fits, and tool-tuned models are systematically biased toward calling something.”
最后半句非常重要:专门调过工具调用的模型,会倾向于”总得调点什么”——这是个系统性的偏置。
4. 稳定性:单次跑分会骗人
τ-bench论文(arXiv:2406.12045)提出的 pass^k 揭穿了这个问题:
gpt-4o 在 τ-retail 上 pass^1 = 61.2%,但 pass^8 < 25%。
详见 一致性指标pass的k次方:“能做成一次”和”每次都能做成”是两种能力。
5. 工具调用能力的一半在模型外
这是本概念最容易被忽略的部分。Anthropic 官方的说法:
“When we evaluate ‘an agent,’ we’re evaluating the harness and the model working together.”
Terminal-Bench 2.0 官方榜给出的数字:
- Claude Opus 4.6:Meta-Harness 76.4% vs Claude Code 58.0%(差 18.4 点)
- GPT-5.3-Codex:LemonHarness 84.5% vs Terminus 2 64.7%(差 19.8 点)
详见 模型与框架分责 与 Terminal-Bench与Harness分责。
而 Lilian Weng 引 Lin et al. 2026 的研究,把这件事拆成了两个独立的轴:
“harness-updating refers to the capability of producing useful harness edits and harness-benefit denotes the capability of utilizing the updated harness… a range of model of different sizes and core intelligence, from Qwen3.5-9B to Claude Opus 4.6, were observed … to show similar harness updating capability; the 9B harness proposer is able to write a skill procedurally isomorphic to Opus. To best utilize a harness, a model needs to invoke skills/tools correctly and timely and be good at long-horizon instruction following.”
结论:写框架的能力跨模型尺寸基本持平;真正分化的是”用好框架的能力”——即”及时正确地调用工具 + 长程指令遵循”。
6. 下一步:智能体强化学习(agentic RL)
2025–2026 年的方向不再是模仿数据,而是在真实可执行环境里做强化学习:
- ToolVerse(arXiv:2607.15660)从近 400 个真实 MCP(MCP 模型上下文协议)中构建约 4,500 个工具的可执行训练环境
- ATLAS(arXiv:2603.06713)针对小模型在大工具空间里的三个失效模式:“eager tool loading saturates context, execution errors compound over time, and sparse rewards limit learning”
相关页面
- 能力维度 —— 工具调用是独立一维
- 一致性指标pass的k次方 —— 怎么正确测量它
- 模型与框架分责 —— 一半能力在模型外
- MCP 模型上下文协议 —— wiki 已有概念页