偏好对战榜(Chatbot Arena / LMArena)
一句话定义:让两个匿名模型回答同一个问题,由人类投票”哪个更好”,再用统计模型把海量投票折算成一个分数。它躲开了静态榜单的污染问题,但引入了全新的偏见。
1. 机制
Chatbot Arena 论文(arXiv:2403.04132)把评测按两个坐标分类:
| 有标准答案(ground truth) | 人类偏好(human preference) | |
|---|---|---|
| 静态题目 | MMLU、HellaSwag、GSM-8K | — |
| 实时题目(live) | — | Chatbot Arena |
论文的自我定位与对静态榜的批评:
“the test sets in these benchmarks are static, meaning they can become contaminated over time … for many complex tasks, establishing a definitive ground truth is not only challenging but sometimes unattainable.”
这些基准的测试题是静态的,意味着它们会随时间推移被污染;而且对许多复杂任务来说,要定出一个标准答案不仅困难,有时根本做不到。
评分算法(LMSYS 官方博客):
“The standard method for computing the Arena Score (i.e., the Bradley-Terry coefficients, which we formerly called the Elo score) is to run a logistic regression of Y_i onto X_i.”
Arena 分数的标准算法是 Bradley-Terry 系数(以前对外叫 Elo 分),做法是对胜负结果跑一次逻辑回归。
2. 官方自曝的偏见:长度压倒一切
Arena风格控制实验(LMSYS 官方,2024-08-29):
“Why is GPT-4o-mini so good? Why does Claude rank so low, when anecdotal experience suggests otherwise? We have answers for you. We controlled for the effect of length and markdown, and indeed, the ranking changed.”
官方直接回应了社区的两个疑问,并给出答案——把「回答长度」和「markdown 排版」这两个因素控制住之后,排名确实变了。
| style 特征 | 系数(Control Both) |
|---|---|
| Length(回答长度) | 0.249 |
| Markdown List | 0.031 |
| Markdown Header | 0.024 |
| Markdown Bold | 0.019 |
“length was the dominant style factor. All other markdown effects are second order.”
长度是压倒性的风格因素,所有其他 markdown 排版效果都是次要的。
控制风格后的排名变动:
| 模型 | 原排名 → 控制后 |
|---|---|
| claude-3-opus | 16 → 10 |
| gpt-4o-mini | 6 → 11 |
| grok-2-mini | 6 → 18 |
⚠️ 官方自己声明这是观察性研究:“There are possible unobserved confounders such as positive correlation between length and substantive quality … a chain-of-thought explanation for a reasoning question”——写得长有时恰恰是因为想得深,全消掉会误伤。
3. 数据特权问题
排行榜幻象论文(arXiv:2504.20879)的发现:
- Meta 在 Llama-4 发布前私测了 27 个变体,只公开最好的那个
- Google 拿走了 19.2%、OpenAI 拿走了 20.4% 的对战数据
- 83 个开源权重模型加起来只占 29.7%
- 用对战榜数据训练能让 ArenaHard 胜率从 23.5% 翻倍到 49.9%,但 “This improvement does not translate to out-of-distribution performance on benchmarks such as MMLU”
卷首题词(Goodhart 定律):“Any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.”
4. 误差棒已经压过名次差(2026 年现状)
2026-08 的快照(CASRAI 口径):
“Claude Fable 5 — 1506±5; Claude Opus 4.6, High effort — 1505±4; Claude Opus 4.7, High effort — 1502±4; Muse Spark 1.2, xHigh — 1498±10”
这是 2026 年 8 月的前四名快照。重点不是具体名次,而是每个分数后面那个 ±4 到 ±5 的误差范围。
“Note the error bars: the top three scores overlap once you account for ±4-5 points, meaning the leaderboard cannot cleanly distinguish ‘best’ from ‘second best’ among them on any given day”
同一来源的另一句很关键:“a leaderboard measuring which AI output humans prefer to read is a different instrument from one measuring reasoning accuracy or factual reliability”
5. 排名算法本身也有问题
《Ranking Unraveled》(arXiv:2411.14483, ACL 2025)的结论很直接:
“We do not recommend ELO as an algorithm to rank LLM performance.”
我们不推荐用 Elo 作为给大语言模型排名的算法。
传递性保持率对比:
| 算法 | Arena 数据集 |
|---|---|
| Elo | 68.24% |
| Markov | 51.38% |
| Glicko | 56.54% |
| Bradley-Terry | 77.29% |
⚠️ 该论文的引文来自二次转录,未逐字核对 arXiv 正文。
更激进的一篇(arXiv:2605.06656)认为全局榜单本身可能无意义:
“the best-fit global Bradley-Terry (BT) ranking is misleading. Nearly 2/3 of the decisive votes cancel out, and even the top 50 models according to the global BT ranking are statistically indistinguishable (pairwise win probabilities are at most 0.53 within the top 50 models)… What appears as global noise is in fact a mixture of coherent but conflicting subpopulations.”
⚠️ 该论文为 2605.* 编号,本次未直接抓取 arXiv 原文,引文来自审稿综述转录。
6. 怎么正确使用对战榜
- ✅ 它测的是”人类更喜欢读哪个回答”,这个信号在真实对话产品里是有价值的。
- ⚠️ 它不等于”更能干”——长度、格式、语气会显著影响分数。
- ⚠️ 误差棒内的名次差没有意义:25–30 分以内的差距属于噪声。
- ⚠️ 别用它来选做工具调用/编程任务的模型——那是完全不同的维度,见 工具调用 与 Terminal-Bench与Harness分责。
相关页面
- 能力维度 —— 为什么”人类偏好”只是其中一维
- 评测污染 —— 静态榜的另一类问题
- LMArena —— 实体页