BFCL(Berkeley Function Calling Leaderboard)
类型:函数调用 / 工具调用评测榜单 维护方:UC Berkeley Gorilla 团队 官网:https://gorilla.cs.berkeley.edu/leaderboard.html 当前版本:V4(2025-07-17 发布,榜单 Last Updated 2026-04-12)
它是”工具调用能力”这个单一维度上最权威的评测。
它怎么判分(这点很关键)
“Throughout, BFCL relies on AST (Abstract Syntax Tree) based, or state-transition based verification ensuring determinism and minimal fluctuations”
即:用抽象语法树/状态转移做程序化判定,不用 LLM 当裁判——所以分数波动小、可复现。这一点与 评测污染 里那些靠裁判打分的榜单形成对比。
区分两种模式:
“FC = native support for function/tool calling. Prompt = walk-around for function calling, using model’s normal text generation capability.”
榜单区分两种模式——FC 指模型原生支持函数调用,Prompt 指模型不支持、只能靠普通文本生成来模拟。看榜时一定要分清这两类,它们的分数不可直接比较。
V4 的权重构成(官方逐字)
Overall Score = Agentic(40%) + Multi-Turn(30%) + Live(10%) + Non-Live(10%) + Hallucination(10%)
| 子项 | 权重 | 条目数 | 测什么 |
|---|---|---|---|
| Agentic | 40% | 665 | Web Search(200)+ Memory(465) |
| Multi-Turn | 30% | 800 | Base / Missing Function / Missing Parameter / Long Context |
| Live | 10% | 1351 | 真实用户贡献的函数 |
| Non-Live | 10% | 1150 | 合成函数 |
| Hallucination | 10% | 1122 | 该不该拒绝调用 |
| Format Sensitivity | 不计分 | 5200 | 格式鲁棒性 |
“Overall Accuracy is the unweighted average of all the sub-categories.”
总分是所有子类别的未加权平均分——每个子类权重相同,不按题目数量加权。
权重变化告诉我们什么
单轮函数调用(Live + Non-Live)只占 20%,而 Agentic + Multi-Turn 占了 70%。
二手解读说得很直白(与官方数据一致):
“That weighting is an admission by the benchmark’s own authors: single-turn accuracy is saturated and no longer separates frontier models, so the leaderboard moved the goalposts to where models still fail — holding state across turns, searching, remembering, and knowing when NOT to call anything.”
这个权重分配等于榜单作者自己承认——单轮调用的准确率已经饱和,再考也区分不出前沿模型的高下了,所以把重心挪到了模型还做不好的地方:跨轮次保持状态、搜索、记忆,以及知道什么时候不该调用。
“The single hardest capability is abstention … tool-tuned models are systematically biased toward calling something.”
最重要的一条:参数量与它不单调
V4 榜单(Last Updated 2026-04-12):
| 模型 | 参数量 | Overall Acc |
|---|---|---|
| Claude-Opus-4-5 (FC) | 未公开 | 77.47 |
| Nanbeige4-3B-Thinking (FC) | 3B | 51.4 |
| xLAM-2-32b-fc-r (FC) | 32B | 54.66 |
| Qwen3-32B (FC) | 32B | 48.71 |
| ToolACE-2-8B (FC) | 8B | 42.44 |
| xLAM-2-3b-fc-r (FC) | 3B | 41.22 |
| Phi-4 (FC) | 14B | 28.79 |
| Gemma-3-27b-it (FC) | 27B | 29.47 |
| Llama-3.3-70B (FC) | 70B | 31.9 |
| Llama-4-Scout-17B-16E | 17B×16E | 28.13 |
| Qwen3-0.6B (FC) | 0.6B | 23.93 |
3B 的 Nanbeige4 高于 70B 的 Llama-3.3。这是 ToolLLM论文 那句”工具调用是训练出来的,不是参数涌出来的”在 2026 年的直接印证。
引用注意事项
- ⚠️ 榜单 Agentic 子项的数值在纯文本抓取中发生了列粘连(出现 298.47 / 355.17 等”准确率”,实为跑分成本 USD)。引用子项分数前请以官网表格为准。
- 只可引用 Overall Acc 与明确的子项总分。
相关页面
- 工具调用 —— 概念页
- 能力维度 —— 它是独立一维的证据来源
- 一致性指标pass的k次方 —— 它测的是 pass