一句话摘要:Allen-Zhu 与 Li 用受控合成数据测出,语言模型能且只能每参数存储 2 bits 知识,即使量化到 int8 也成立——据此一个 7B 模型能装下 14B bits,超过英文维基百科加教科书的总和。
原始标题
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws(Zeyuan Allen-Zhu & Yuanzhi Li,arXiv:2404.05405,2024-04-08)
来源:https://arxiv.org/abs/2404.05405
⚠️ 编号勘误:网络上常把这条结论误标为 arXiv 2309.02427。2309.02427 实为《Cognitive Architectures for Language Agents (CoALA)》,与知识容量无关。本系列三部分正确编号为:
- Part 3.1 知识存储与提取 = 2309.14316
- Part 3.2 知识操纵 = 2309.14402
- Part 3.3 知识容量缩放定律 = 2404.05405(本页)
它解决什么问题
以往衡量模型能力都用”损失(loss)“或”榜单分数”这类间接指标。这篇论文换了个问法:一个模型里到底装了多少比特的事实知识? 作者用受控的合成数据集,把知识表示成 (实体, 属性, 值) 三元组(如 (USA, capital, Washington D.C.)),直接数比特。
关键要点
1. 核心结论:2 bits/parameter
“Through multiple controlled datasets, we establish that language models can and only can store 2 bits of knowledge per parameter, even when quantized to int8, and such knowledge can be flexibly extracted for downstream applications. Consequently, a 7B model can store 14B bits of knowledge, surpassing the English Wikipedia and textbooks combined based on our estimation.”
这句话的”only can(只能)“是关键:作者强调 2 bits 是上限,不是平均值。训练不足时达不到。
2. 影响容量的五个因素
“More broadly, we present 12 results on how (1) training duration, (2) model architecture, (3) quantization, (4) sparsity constraints such as MoE, and (5) data signal-to-noise ratio affect a model’s knowledge storage capacity.”
更广泛地说,我们给出了 12 条结果,覆盖五个因素如何影响模型的知识存储容量:训练时长、模型架构、量化、稀疏约束(如 MoE)、以及数据的信噪比。
3. 架构的反直觉发现
“The GPT-2 architecture, with rotary embedding, matches or even surpasses LLaMA/Mistral architectures in knowledge storage, particularly over shorter training durations. This arises because LLaMA/Mistral uses GatedMLP, which is less stable and harder to train.”
带旋转位置编码的 GPT-2 架构,在知识存储上匹配甚至超过 LLaMA/Mistral 架构,尤其在训练时长较短时。原因是 LLaMA/Mistral 用的 GatedMLP 更不稳定、更难训练。这条很反直觉——「更先进的架构」反而更不擅长存知识。
4. 数据质量(信噪比)的杠杆极大
“Prepending training data with domain names (e.g., wikipedia.org) significantly increases a model’s knowledge capacity. Language models can autonomously identify and prioritize domains rich in knowledge, optimizing their storage capacity.”
这条对”数据工程比参数量更重要”是最直接的证据:同样的数据,加个域名前缀,容量就变了。
5. 论文提出的终极问题
“Despite rumors of GPT-4 having over 1T parameters, is it necessary to store all human knowledge? Could a 10B model, if trained sufficiently with high-quality data, match GPT-4’s knowledge capacity? Our paper seeks to address these questions.”
尽管有传言说 GPT-4 有超过 1T 参数,但真的有必要存下人类全部知识吗?一个 100 亿参数的模型,如果用高质量数据充分训练,能不能达到 GPT-4 的知识容量?本文就是在回答这些问题。
重要引用(英文原文)
“Scaling laws describe the relationship between the size of language models and their capabilities. Unlike prior studies that evaluate a model’s capability via loss or benchmarks, we estimate the number of knowledge bits a model stores.”
缩放定律描述的是语言模型规模与其能力之间的关系。不同于以往用损失值或榜单分数来评估能力的研究,我们直接去估算一个模型存储了多少比特的知识。
边界与争议
- 该论文摘要外的细分数字(“知识需曝光 1000 次""垃圾数据混训使容量降 20 倍”等)在本次抓取中只找到二手中文解读,未获 arXiv 正文逐字核实。写作时只可引用摘要中的一手结论,细分数字须另核原文 PDF。
- 2 bits/param 是”存储”上限,不等于”能用”。能存进去不等于能取出来——这正是 知识操纵与思维链论文 的核心发现。
- 该结论对 MoE 的含义:稀疏约束会影响容量,因此”总参数 × 2 bits”不能直接套用到 混合专家 模型上。