一句话摘要:Schaeffer 等人(NeurIPS 2023 杰出论文)证明,涌现能力是研究者选的指标造成的假象:用”精确匹配”这类非线性/不连续指标看,能力会突然出现;换成连续指标,就只剩下平滑、可预测的改善。

原始标题

Are Emergent Abilities of Large Language Models a Mirage?(Schaeffer, Miranda, Koyejo, Stanford,arXiv:2304.15004,NeurIPS 2023 Outstanding Paper)

来源:https://openreview.net/attachment?id=ITw9edRDlD&name=pdf 作者页(含机制总结):https://rylanschaeffer.github.io/publications/2023-neurips-emergent-abilities-mirage


它解决什么问题

涌现能力论文 声称某些能力”突然出现、无法预测”。反方问了一个很尖锐的问题:会不会不是模型变了,而是你拿的尺子变了?

关键要点

1. 核心论点

“Here, we present an alternative explanation for emergent abilities: for a particular task and model family, when analyzing fixed model outputs, emergent abilities appear due to the researcher’s choice of metric rather than due to fundamental changes in models with scale. Specifically, nonlinear or discontinuous metrics produce seemingly emergent abilities, whereas linear or continuous metrics produce smooth, continuous, predictable changes in model performance.”

我们提出另一种解释——涌现能力之所以出现,是因为研究者选了什么指标,而不是模型随规模发生了根本变化。具体来说,非线性或不连续的指标会制造出「涌现」的假象。

2. 涌现的两个定义属性(引言原文)

“These quotations collectively identify the two defining properties of emergent abilities in LLMs: 1. Sharpness, transitioning seemingly instantaneously from not present to present; 2. Unpredictability, transitioning at seemingly unforeseeable model scales.”

这些引文共同确立了涌现能力的两个定义属性:一是陡峭性(从无到有看似瞬间完成),二是不可预测性(转折发生在看似无法预见的规模上)。反方要推翻的正是这两条。

3. 关键机制

“Our doubt stems from the observation that emergent abilities seem to appear only under metrics that nonlinearly or discontinuously scale any model’s per-token error rate.”

一个直觉例子:一道多步数学题要连对 5 步才算对。单步正确率从 60% 平滑升到 80%,“全对率”却从 0.6⁵≈8% 跳到 0.8⁵≈33%——看起来像”突然会了”,实际底下是连续改善。

4. 作者本人的一句话总结

“Key Insight: When using nonlinear or discontinuous metrics (like exact-match accuracy), smooth, continuous improvements in model performance can appear as sharp, discontinuous ‘emergent’ abilities. By changing to continuous metrics (like token-level accuracy or Brier score), the apparent emergence disappears and is replaced by smooth, predictable improvement.”

核心洞察——用非线性或不连续指标(如精确匹配准确率)时,平滑连续的性能提升会看起来像陡峭、不连续的「涌现」;换成连续指标(如 token 级准确率或 Brier 分数),涌现就消失了,取而代之的是平滑可预测的提升。

5. 三重验证

“we (1) make, test and confirm three predictions on the effect of metric choice using the InstructGPT/GPT-3 family on tasks with claimed emergent abilities; (2) make, test and confirm two predictions about metric choices in a meta-analysis of emergent abilities on the Beyond the Imitation Game Benchmark (BIG-Bench); and (3) show how to choose metrics to produce never-before-seen seemingly emergent abilities in multiple vision tasks across diverse deep network architectures.”

第 (3) 点是杀手锏:他们能凭换指标,在任何架构上制造出看起来像涌现的曲线。

重要引用(英文原文)

“Via all three analyses, we provide evidence that emergent abilities disappear with different metrics or with better statistics, and may not be a fundamental property of scaling AI models.”

综合这三项分析,我们给出的证据是:换用不同的指标、或改用更好的统计方法,涌现能力就会消失——因此涌现可能并不是「缩放 AI 模型」的一个根本属性。

边界与争议

  • 反方并未否定”能力会变强”,否定的是”突变且不可预测”这两个属性。这一点常被误读。
  • 正方(Wei et al.)的回应:即便是指标假象,“在某个规模以下这项能力确实没法用”这个工程事实仍然成立。
  • 第三方旁证(早于争论):BIG-bench(arXiv:2206.04615)2022 年就发现”tasks that exhibit ‘breakthrough’ behavior at a critical scale often involve multiple steps or components, or brittle metrics”——与反方结论方向一致。

与本文其他页面的关系

  • 涌现能力 —— 概念页,正反合观
  • 能力维度 —— 为什么不同维度”解锁”的观感不同
  • 评测污染 —— 榜单可信度的另一重问题