涌现能力(Emergent Abilities)
一句话定义:小模型完全没有、大模型突然具备的能力。这个概念在 2022 年提出,2023 年被一篇 NeurIPS 杰出论文质疑为”指标造成的幻觉”,至今没有定论。
1. 正方的定义
涌现能力论文(Wei et al., TMLR 2022):
“We consider an ability to be emergent if it is not present in smaller models but is present in larger models. Thus, emergent abilities cannot be predicted simply by extrapolating the performance of smaller models.”
我们把一种能力称为「涌现的」,条件是它在小模型上不存在、却在大模型上存在。因此涌现能力没法靠外推小模型的表现来预测。
两个定义属性(反方论文归纳):
“1. Sharpness, transitioning seemingly instantaneously from not present to present; 2. Unpredictability, transitioning at seemingly unforeseeable model scales”
涌现的两个定义属性:一是陡峭性(从无到有的转变看起来是瞬间完成的),二是不可预测性(转变发生在看起来无法预见的模型规模上)。
思想源头是物理学家 P.W. Anderson 1972 年的《More Is Different》:
“Emergence is when quantitative changes in a system result in qualitative changes in behavior.”
涌现就是「系统里量的变化,导致了行为上质的变化」。
政策含义(也是这句话最有影响力的地方):“additional scaling could further expand the range of capabilities of language models”
2. 反方的反驳
涌现能力是幻觉吗论文(Schaeffer et al., NeurIPS 2023 Outstanding Paper):
“emergent abilities appear due to the researcher’s choice of metric rather than due to fundamental changes in models with scale. Specifically, nonlinear or discontinuous metrics produce seemingly emergent abilities, whereas linear or continuous metrics produce smooth, continuous, predictable changes.”
涌现能力的出现,是研究者选了什么指标造成的,而不是模型随规模发生了根本变化。具体来说:非线性或不连续的指标会制造出「看似涌现」的能力,而线性或连续的指标只会给出平滑、连续、可预测的变化。
一个直觉例子
一道数学题要连对 5 步才算对。假设单步正确率从 60% 平滑升到 80%:
- 全对率 = 0.6⁵ ≈ 8% → 0.8⁵ ≈ 33%
底下是连续改善,画出来的曲线却像”突然开窍”。涌现可能长在尺子上,不长在模型里。
反方最有力的证据是他们能凭换指标,在视觉任务上凭空制造出涌现曲线:
“we … show how to choose metrics to produce never-before-seen seemingly emergent abilities in multiple vision tasks across diverse deep network architectures.”
3. 更早的伏笔
BIG-bench(arXiv:2206.04615,2022)在争论爆发前一年就观察到了同一件事:
“tasks that improve gradually and predictably commonly involve a large knowledge or memorization component, whereas tasks that exhibit ‘breakthrough’ behavior at a critical scale often involve multiple steps or components, or brittle metrics”
“突破性行为”往往出现在”多步骤”或”指标脆弱”的任务上——这与 Schaeffer 的结论方向一致。
4. 目前的中立立场
| 问题 | 答案 |
|---|---|
| ”能力会随规模突变” | 大概率是指标假象。换成连续指标,涌现消失。 |
| “某些能力在某个规模以下确实没法用” | 工程事实成立。无论曲线是真是假,小模型在那道题上就是 0 分。 |
| “继续放大还能解锁新能力吗” | 没有定论。这取决于涌现是真是假。 |
一句话记住:“涌现”作为数学性质可疑,作为工程观察有用。
5. 为什么这个概念对”参数与智能”重要
- 如果涌现为真 → 参数量存在”能力门槛”,跨过去就是质变,那么堆规模是唯一路径。
- 如果涌现为假 → 能力是连续积累的,那么数据、后训练、推理期算力(推理期缩放)都可以在不改参数量的前提下把能力推上去。
- 2026 年的实践更支持后者:同一份权重,换推理策略就能拉开 27 个百分点(见 推理期缩放 的 s1 案例),这更像是”连续的能力被不同的使用方式释放出来”,而不是”参数跨过了某道门”。
相关页面
- 能力维度 —— 不同维度的”解锁”节奏不同
- 缩放定律 —— 平滑改善的那一半
- 评测污染 —— 指标问题的另一面