Terminal-Bench
类型:终端环境下的智能体评测 论文:arXiv:2601.11868(2.0 版本) 规模:89 个任务(2.0),k=5 次试验 覆盖:软件工程、系统管理、数据处理、模型训练、安全 官网:https://www.tbench.ai/leaderboard/terminal-bench/2.0
它在本知识库中的独特价值:它是”同一模型换框架差多少”这个问题最大的自然实验场。
它测什么
给模型一个终端(命令行),让它完成真实任务,用程序化测试套件判定成败(pass@1,3 次重复)。
“We show that frontier models and agents score less than 65% on the benchmark”(论文发表时的口径)
为什么它成了 harness 研究的天然实验场
因为多个团队会拿同一个模型配上各自的 harness 去打榜,官方榜把这些条目都保留了下来。
Claude Opus 4.6(同一份权重)
| 榜位 | harness | 准确率 |
|---|---|---|
| 11 | Meta-Harness(Stanford IRIS) | 76.4% ± 2.4 |
| 14 | Capy | 75.3% ± 2.4 |
| 20 | TongAgents | 71.9% ± 2.7 |
| 36 | Terminus 2(官方最小 agent) | 62.9% ± 2.7 |
| 50 | Claude Code(Anthropic 自家) | 58.0% ± 2.9 |
最高与最低相差 18.4 个百分点,而厂商自家的 Claude Code 排在第 50 位。
GPT-5.3-Codex(同一份权重)
| 榜位 | harness | 准确率 |
|---|---|---|
| 2 | LemonHarness | 84.5% ± 2.6 |
| 15 | Simple Codex(OpenAI 自家简化版) | 75.1% ± 2.4 |
| 32 | Terminus 2 | 64.7% ± 2.7 |
相差 19.8 个百分点。
交叉验证:LangChain 只改 harness
“We used a simple recipe to iteratively improve deepagents-cli … 13.7 points from 52.8 to 66.5 on Terminal Bench 2.0. We only tweaked the harness and kept the model fixed, gpt-5.2-codex.”
我们只改框架、模型完全固定(gpt-5.2-codex),分数就从 52.8 涨到 66.5,涨了 13.7 个点。
“Our coding agent went from Top 30 to Top 5 … We only changed the harness.”
我们的编程智能体从第 30 名开外升到了前 5 名,而唯一的改动就是换了框架。
2026-08-31 的榜首(Artificial Analysis 口径)
| 名次 | 模型 | 分数 |
|---|---|---|
| 1 | GPT-5.6 Sol (xhigh) | 89.5% |
| 2 | Claude Opus 5 (Adaptive Reasoning, Max Effort) | 89.1% |
| 3 | Grok 4.6 (high) | 88.4% |
注意:这个口径是”每个模型的最佳 harness 成绩”,与上面那组”同一模型不同 harness”的数据不是同一张表。
使用注意事项
- 误差棒为 ±2.2 到 ±2.9:18 个点的差距远大于误差棒、结论稳健;但相邻名次之间 1 个点的差别不可区分。
- 各 harness 条目的提交日期从 2026-02 到 2026-05 不等,不完全是同一时间截面。
- 官方榜自述:“A Terminal-Bench team member ran the evaluation and verified the results.”