BFCL(Berkeley Function Calling Leaderboard)

类型:函数调用 / 工具调用评测榜单 维护方:UC Berkeley Gorilla 团队 官网:https://gorilla.cs.berkeley.edu/leaderboard.html 当前版本:V4(2025-07-17 发布,榜单 Last Updated 2026-04-12)

它是”工具调用能力”这个单一维度上最权威的评测。


它怎么判分(这点很关键)

“Throughout, BFCL relies on AST (Abstract Syntax Tree) based, or state-transition based verification ensuring determinism and minimal fluctuations”

即:用抽象语法树/状态转移做程序化判定,不用 LLM 当裁判——所以分数波动小、可复现。这一点与 评测污染 里那些靠裁判打分的榜单形成对比。

区分两种模式:

“FC = native support for function/tool calling. Prompt = walk-around for function calling, using model’s normal text generation capability.”

榜单区分两种模式——FC 指模型原生支持函数调用,Prompt 指模型不支持、只能靠普通文本生成来模拟。看榜时一定要分清这两类,它们的分数不可直接比较。

V4 的权重构成(官方逐字)

Overall Score = Agentic(40%) + Multi-Turn(30%) + Live(10%) + Non-Live(10%) + Hallucination(10%)
子项权重条目数测什么
Agentic40%665Web Search(200)+ Memory(465)
Multi-Turn30%800Base / Missing Function / Missing Parameter / Long Context
Live10%1351真实用户贡献的函数
Non-Live10%1150合成函数
Hallucination10%1122该不该拒绝调用
Format Sensitivity不计分5200格式鲁棒性

“Overall Accuracy is the unweighted average of all the sub-categories.”

总分是所有子类别的未加权平均分——每个子类权重相同,不按题目数量加权。

权重变化告诉我们什么

单轮函数调用(Live + Non-Live)只占 20%,而 Agentic + Multi-Turn 占了 70%。

二手解读说得很直白(与官方数据一致):

“That weighting is an admission by the benchmark’s own authors: single-turn accuracy is saturated and no longer separates frontier models, so the leaderboard moved the goalposts to where models still fail — holding state across turns, searching, remembering, and knowing when NOT to call anything.”

这个权重分配等于榜单作者自己承认——单轮调用的准确率已经饱和,再考也区分不出前沿模型的高下了,所以把重心挪到了模型还做不好的地方:跨轮次保持状态、搜索、记忆,以及知道什么时候不该调用。

“The single hardest capability is abstention … tool-tuned models are systematically biased toward calling something.”

最重要的一条:参数量与它不单调

V4 榜单(Last Updated 2026-04-12):

模型参数量Overall Acc
Claude-Opus-4-5 (FC)未公开77.47
Nanbeige4-3B-Thinking (FC)3B51.4
xLAM-2-32b-fc-r (FC)32B54.66
Qwen3-32B (FC)32B48.71
ToolACE-2-8B (FC)8B42.44
xLAM-2-3b-fc-r (FC)3B41.22
Phi-4 (FC)14B28.79
Gemma-3-27b-it (FC)27B29.47
Llama-3.3-70B (FC)70B31.9
Llama-4-Scout-17B-16E17B×16E28.13
Qwen3-0.6B (FC)0.6B23.93

3B 的 Nanbeige4 高于 70B 的 Llama-3.3。这是 ToolLLM论文 那句”工具调用是训练出来的,不是参数涌出来的”在 2026 年的直接印证。

引用注意事项

  • ⚠️ 榜单 Agentic 子项的数值在纯文本抓取中发生了列粘连(出现 298.47 / 355.17 等”准确率”,实为跑分成本 USD)。引用子项分数前请以官网表格为准。
  • 只可引用 Overall Acc 与明确的子项总分。

相关页面