一句话摘要:模型能把每个人的出生月份答对 100%、训了 25,000 条样本,却回答不了”这个人出生在偶数月吗”——除非让它先说出月份再判断。推理和知识是两种独立的能力,后者不随前者自动获得。
原始标题
Physics of Language Models: Part 3.2, Knowledge Manipulation(Allen-Zhu & Li,arXiv:2309.14402,v1 2023-09-25,v2 2024-07-16,结论未变)
来源:https://arxiv.org/html/2309.14402v2
它解决什么问题
知识容量缩放定律论文 回答了”模型里装了多少知识”。这一篇接着问:装进去的知识,模型能灵活使用吗? 作者定义了四类知识操纵任务:
| 任务 | 例子 |
|---|---|
| 检索(retrieval) | “A 的 X 属性是什么?“ |
| 分类(classification) | “A 的 X 属性是奇数还是偶数?“ |
| 比较(comparison) | “A 和 B 在 X 属性上谁更大?“ |
| 反向检索(inverse search) | “谁的 X 属性等于 T?“ |
关键要点
1. 一步都推不动(Result 3)——最锋利的一条
“Specifically, for the binary classification ‘Was Anya born in an even month’, language models fail without CoT — i.e., without first generating the month ‘October’ and then assessing its parity. This remains true even if the model is sufficiently trained • to answer everyone’s birth month with 100% accuracy, • on 25,000 QA samples, more than needed to classify 12 months to 2 classes. This reveals that language models cannot efficiently be trained+finetuned to perform even a single step of knowledge manipulation during inference time without CoT.”
翻译成人话:模型能把”Anya 出生在十月”背得滚瓜烂熟,却没法直接回答”十月的奇偶性”。它必须先把”十月”这个词吐出来,再对着这个词判断奇偶。
2. 改善”抽取”不改善”操纵”(Result 4 / Result 5)
”• Including sufficient CoT samples in training does not enhance non-CoT inference (Result 4); • Improving model’s knowledge extraction don’t improve its manipulation ability (Result 5).”
两条结论——训练时加入足够的思维链样本,改善不了推理时不做思维链的表现(结果 4);改善模型的知识抽取能力,也改善不了它的知识操纵能力(结果 5)。这是两条独立的通道。
3. 反向检索彻底失败:模型不是数据库
“We discover that language models cannot perform this task, regardless of training methods, data, or model size, unless the knowledge is already presented inversely in the data (Result 8).² This suggests that language models cannot be used as databases.”
我们发现语言模型做不到反向检索(「谁的某属性等于 T?」),无论换什么训练方法、数据或模型规模,除非数据里本来就有倒过来写的内容。作者的结论很重:语言模型不能被当作数据库使用。
4. 比较任务即使海量训练也近乎随机
“the accuracy of comparing knowledge among 100 options is barely random guess, even with 2,500,000 training samples, more than enough to learn to rank 100 objects”
让模型在 100 个选项里比较知识,准确率几乎等于瞎猜——即使训了 250 万条样本,而这个数据量用来学会给 100 个对象排序绰绰有余。
5. 规模也救不了(Result 6、8)
“We also demonstrate that modern large models like GPT-4 or Llama-3 struggle with these tasks, suggesting these limitations may be inherent to generative language models and not easily overcome by scaling up.”
我们还证明,像 GPT-4 和 Llama-3 这样的现代大模型在这些任务上同样吃力——这说明上述局限可能是生成式语言模型固有的,不容易靠加大规模来克服。
作者自证时效性:“When we prepared this paper we used GPT-4 of 2023. As of May 10, 2024, such counter-examples still apply to GPT-4 and Llama-3.”
重要引用(英文原文)
“Language models can store vast factual knowledge, yet their ability to flexibly use this knowledge for downstream tasks (e.g., via instruction finetuning) remains questionable. … We show that language models excel in knowledge retrieval but struggle even in the simplest classification or comparison tasks unless Chain of Thoughts (CoTs) are employed during both training and inference.”
语言模型能存储海量事实知识,但它们灵活使用这些知识的能力仍然存疑。我们证明:模型擅长知识检索,但连最简单的分类或比较任务都搞不定,除非在训练和推理时都用思维链。
“Our primary contribution is a controlled, synthetic experiment that confirms these weaknesses are inherent to language models: they cannot efficiently manipulate knowledge from pre-training data, even when such knowledge is perfectly stored in the models, despite adequate training and sufficient model size.”
边界与争议
- 这是受控合成实验,用的是人工构造的传记数据集,批评者会质疑与真实语料的差距。但作者专门用 GPT-4 / Llama-3 做了验证(Result 6、8)。
- “模型不能当数据库”这句话要正确理解:它指的是不能靠参数存储去支撑反向检索类操作,不是说参数里没有知识。工程上的解法是外挂检索(RAG)和数据库,而不是继续堆参数。
- 这条结论直接支撑了文章的核心论点之一:知识与推理是两条独立的轴,见 能力维度。