一句话摘要:τ-bench 测的是客服场景下的”工具-智能体-用户”三方交互。最强的 gpt-4o 单次成功率 61.2%,但要求连续 8 次都成功的 pass^8 掉到 25% 以下——单次跑分掩盖了可靠性问题。

原始标题

τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains(Sierra / Princeton,arXiv:2406.12045,2024-06-17)

来源:https://arxiv.org/abs/2406.12045


它解决什么问题

多数工具调用评测只看”这一轮函数调没调对”。但真实的客服场景要求:遵守领域规则、跨多轮保持状态、跟真人来回沟通。τ-bench 要把这些都纳入。

关键要点

1. 任务设计

  • τ-retail:115 个零售客服任务;τ-airline:50 个航司客服任务。
  • 判分不看对话文本,看对话结束时的数据库状态是否与标注的目标状态一致,外加回复是否包含必要信息:

“The reward of a task episode r = r_action × r_output ∈ {0,1} is based on (1) whether the final database is identical to the unique ground truth outcome database (r_action), and (2) whether the agent’s responses to the user contain all necessary information (r_output).”

判分不看对话说了什么,只看两件事的乘积——(1) 对话结束时数据库的状态是不是和标注的目标状态完全一致;(2) 模型给用户的回复里有没有包含所有必要信息。

2. pass^k 的提出(本论文最重要的贡献)

“For real-world agent tasks requiring reliability and consistency like customer service, we propose a new metric – pass^k (pass hat k), defined as the chance that all k i.i.d. task trials are successful, averaged across tasks.”

对于客服这类要求「可靠且一致」的真实任务,我们提出一个新指标 pass^k(读作 pass hat k),定义为「k 次独立运行全部成功」的概率,再对所有任务取平均。

指标含义适用
pass@kk 次里至少一次成功探索类任务(写代码、找解法)
pass^kk 次全部成功面向用户的服务(必须每次都对)

3. 单次成功率 vs 稳定性(核心数字)

模型τ-retail pass^1τ-airline pass^1平均
gpt-4o61.235.248.2
gpt-4-turbo57.732.445.1
claude-3-opus44.234.739.5
mistral-large30.722.426.6
gemini-1.5-pro21.714.017.9
gpt-3.5-turbo20.010.815.4
meta-llama-3-70B(text-ReAct)14.814.414.6

“Our experiments show that even state-of-the-art function calling agents (like gpt-4o) succeed on <50% of the tasks, and are quite inconsistent (pass^8 <25% in retail).”

我们的实验显示,即使是最先进的函数调用智能体(比如 gpt-4o),成功率也不到 50%,而且非常不稳定(在零售场景中 pass^8 低于 25%)。

§5.1 原文:“Even for the best-performing gpt-4o function calling agent which has a >60% average task success, pass^8 drops to <25%.“

4. 领域规则(policy)本身就是能力的一部分

移除系统提示词里的领域规则后,gpt-4o 在 τ-airline 上从 33.2 掉到 10.8(降 22.4 个百分点)。

⚠️ 论文自身存在表格不一致:Table 2 中 gpt-4o 的 airline 基线为 35.2,Table 3 为 33.2。引用时须注明。

5. 后续 τ²-bench(双控环境,arXiv 2506.07982)

“Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider.”

从”用户只提供信息”切到”用户也能动手改世界”的双控环境:gpt-4.1 掉 18 个百分点,o4-mini 掉 25 个百分点。

“around 20% pass^1 when agents must shift from autonomous operation to guiding a user.”

当任务从「模型自己干」切换到「模型引导用户一起干」时,单次成功率会掉大约 20 个百分点。

重要引用(英文原文)

“pass^k … defined as the chance that all k i.i.d. task trials are successful, averaged across tasks.”

pass^k ……定义为「k 次独立运行全部成功」的概率,再对所有任务取平均。

边界与争议

  • Anthropic 官方给出的选择指南(2026-01-09):

    “pass@k for tools where one success matters, pass^k for agents where consistency is essential.”

  • Anthropic 同时警告不要僵化地检查路径:

    “There is a common instinct to check that agents followed very specific steps like a sequence of tool calls in the right order. We’ve found this approach too rigid and results in overly brittle tests, as agents regularly find valid approaches that eval designers didn’t anticipate.”

    人们常有一种本能——去检查智能体有没有严格按指定步骤走(比如工具调用的顺序对不对)。我们发现这种做法太死板,会让测试变得极其脆弱,因为智能体经常能找到评测设计者没预料到、但同样有效的解法。

  • 数据是 2024 年的,模型的绝对分数已大幅改善,但”pass^1 与 pass^k 之间存在巨大落差”这个结构性现象仍然成立。

与本文其他页面的关系