一句话摘要:OpenAI 团队 2020 年的实证研究发现,语言模型的损失随模型规模、数据规模、训练算力呈**幂律(power-law)**下降,并且大模型样本效率更高——因此固定算力下最优策略是”训极大的模型、喂相对少的数据、远未收敛就停”。

原始标题

Scaling Laws for Neural Language Models(Kaplan, McCandlish, Henighan, Brown, Chess, Child, Gray, Radford, Wu, Amodei,arXiv:2001.08361,2020-01-23)

来源:https://arxiv.org/abs/2001.08361 ;HTML 全文 https://arxiv.org/html/2001.08361v1


它解决什么问题

在 2020 年之前,“把模型做大能变好”只是经验直觉,没有人给出可外推的定量关系。这篇论文要回答的是:如果给你一笔固定的训练算力,应该把它花在”更大的模型”上,还是”更多的数据”上?

关键要点(按原文结构)

1. 三条幂律

记号: = 非嵌入参数量(non-embedding parameters), = 训练 token 数, = 训练算力。

约束条件公式指数尺度常数
参数受限
数据受限 tokens
算力受限 PF-days

算力估算式:

“Accounting for the backwards pass (approximately twice the compute as the forwards pass), we then define the estimated non-embedding compute as C≈6N floating point operators per training token.”

即:训练算力 ≈ 6 × 参数量 × token 数。这是所有”参数/数据/算力”换算的起点。

2. 大模型样本效率(sample efficiency)更高

“Large models are more sample-efficient than small models, reaching the same level of performance with fewer optimization steps (Figure 2) and using fewer data points (Figure 4).”

大模型比小模型更「省样本」——用更少的优化步数、更少的数据点,就能达到同样的水平。

3. 固定算力下的最优解:大模型 + 早停

“When working within a fixed compute budget C but without any other restrictions on the model size N or available data D, we attain optimal performance by training very large models and stopping significantly short of convergence (see Figure 3).”

在算力预算固定、模型大小和数据量都不受限时,最优策略是训练非常大的模型,并且在远未收敛时就提前停下。

“As the computational budget C increases, it should be spent primarily on larger models, without dramatic increases in training time or dataset size. This also implies that as models grow larger, they become increasingly sample efficient.”

算力预算增加时,应该主要用来把模型做大,而不是大幅增加训练时间或数据集规模。这也意味着模型越大,样本效率越高。

4. 经验最优分配(这句被后来的 Chinchilla 直接推翻)

“which closely matches the empirically optimal results N∝C_min^{0.73}, B∝C_min^{0.24}, and S∝C_min

拆解:参数量应随算力的 0.73 次方增长,数据量只随 0.27 次方增长——即”算力涨了主要用来加参数,数据只要慢慢加”。这正是 Chinchilla 论文要修正的结论。

5. 数据应随模型次线性增长

“Equation (1.1) and (1.2) together suggest that as we increase the model size, we should increase the dataset size sublinearly according to D∝N

把两个经验公式合起来看,模型变大的时候,数据量只需要按 D∝N^0.74 次线性增长就够了。注意:这条被后来的 Chinchilla 直接推翻了。

6. 论文的自我外推

“C*∼10^4 PF-Days, N*∼10^12 parameters, D*∼10^12 tokens, L*∼1.7 nats/token”

这是论文按公式往外推的「拐点」量级——算力约 10^4 PF-Days、参数约 10^12、token 约 10^12、损失约 1.7 nats/token。

重要引用(英文原文)

“We study empirical scaling laws for language model performance on the cross-entropy loss. The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude.”

我们研究的是语言模型在交叉熵损失上的经验缩放规律——损失随模型规模、数据集规模、训练算力呈幂律下降,有些趋势跨越了七个数量级以上。

“Our results strongly suggest that larger models will continue to perform better, and will also be much more sample efficient than has been previously appreciated. Big models may be more important than big data.”

边界与争议(必读)

  • 实验规模有限:模型参数 768 到 1.5B(不含嵌入),数据 22M 到 23B tokens。所有结论都是从小规模实验外推出来的。
  • 结论被推翻了一半:2022 年的 Chinchilla 用 400 多个模型重做实验,指出 Kaplan 的 高估了参数的重要性,正确比例应接近 0.5 / 0.5。详见 Chinchilla计算最优论文。
  • “大模型样本效率更高”这部分没有被推翻,反而被后续实践(计算最优与过训练)以另一种方式验证了。

与本文其他页面的关系

  • 缩放定律 —— 概念页,讲幂律关系本身
  • Chinchilla计算最优论文 —— 直接的修正者
  • 模型参数 —— 参数量这个变量到底指什么