长上下文(Long Context)
一句话定义:模型一次能”看到”多少 token。但要区分两个截然不同的数字:标称窗口(能塞进去多少) 和 有效长度(塞进去之后还能正常思考多少)——后者往往只有前者的几分之一。
1. 两个数字,别搞混
| 含义 | 谁能决定 | |
|---|---|---|
| 标称上下文窗口 | 位置编码与注意力机制支持的最大输入长度 | 架构设计 + 长文本训练 |
| 有效上下文长度 | 在该长度下性能仍可接受的实际长度 | 训练数据分布、注意力实现、任务类型 |
2026 年主流旗舰的标称窗口已经普遍到 1M token(DeepSeek-V4-Pro、GLM-5.2、Kimi K2.5、混元 Hy4 等),但这不代表它们能在 1M 长度上正常工作。
2. “标称 ≠ 有效”的硬证据
NoLiMa长上下文论文(arXiv:2502.05167,ICML 2025)——它堵上了大海捞针测试的字面匹配漏洞:
“We evaluate 13 popular LLMs that claim to support contexts of at least 128K tokens. … At 32K, for instance, 11 models drop below 50% of their strong short-length baselines. Even GPT-4o, one of the top-performing exceptions, experiences a reduction from an almost-perfect baseline of 99.3% to 69.7%.”
标称 128K 的模型,在 32K(四分之一处)就有 11 个掉了一半。
RULER(arXiv:2404.06654,NVIDIA)的同期结论:
“Despite achieving nearly perfect accuracy in the vanilla NIAH test, almost all models exhibit large performance drops as the context length increases. While these models all claim context sizes of 32K tokens or greater, only half of them can maintain satisfactory performance at the length of 32K.”
尽管在经典的大海捞针测试里几乎满分,几乎所有模型在上下文变长时都会大幅掉分。它们都自称支持 32K 以上,但只有一半能在 32K 这个长度上保持令人满意的表现。
3. 大海捞针(NIAH)为什么不够用
“NIAH is fundamentally a simple retrieval task … this benchmark typically assesses direct lexical matching, which may not be representative of flexible, semantically oriented tasks.”
NoLiMa 的改进:让”问题和针之间几乎没有字面重叠”,模型必须推断潜在关联才能定位。
RULER 的改进:在 vanilla NIAH 之外增加 multi-hop tracing(多跳追踪) 和 aggregation(聚合) 任务。
4. 进阶现象:上下文腐坏(context rot)
ContextRot技术报告(Chroma,2025-07-14)用 18 个模型 × 8 种长度 × 11 个位置系统测试:
“models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows.”
模型不是均匀地使用它全部的上下文窗口;相反,随着输入越来越长,它的表现会越来越不可靠。
几个具体发现:
- 干扰项会雪上加霜:“Even a single distractor reduces performance … adding four distractors compounds this degradation further”
- 语义相似度越低,衰减越快
- 反直觉:“models perform worse when the haystack preserves a logical flow of ideas. Shuffling the haystack and removing local coherence consistently improves performance”
- 家族差异:“Claude models … tend to abstain when uncertain. In contrast, GPT models show the highest rates of hallucination”
详见 上下文腐坏。
5. 长上下文的技术成本:KV cache
长上下文的第二座大山不是注意力计算,而是 KV cache(键值缓存)——为了不重复计算,模型会把已处理的 token 的 Key/Value 存下来。
KV cache 字节数 = batch × 序列长度 × 2 × 层数 × KV 头数 × 每头维度 × 每参数字节数
具体算例(Llama 3.1 70B,FP16):
| 上下文 | KV cache |
|---|---|
| 128K | 约 42 GB |
| 1M | 约 328 GB(还不含 140 GB 的模型权重) |
另一个对照(Llama-2 7B,32 层 / 32 KV heads @ 32K):16 GB KV cache,而 FP16 权重只要约 14 GB——KV cache 比模型本身还大。
这也是为什么各大厂在 2026 年都在做注意力稀疏化:DeepSeek 的 DSA、GLM-5.2 的 IndexShare(1M 上下文下单 token FLOPs 降 2.9×)等。2026 年的新战场不是激活参数,而是长上下文。
6. 工程含义
-
❌ 别把”1M 窗口”当成”可以把整个代码库丢进去”。
-
✅ 检索优于堆料:Chroma 的 LongMemEval 实验显示,同一批信息,聚焦提示(约 300 tokens)下的表现显著高于完整提示(平均约 113K tokens)。
“adding irrelevant context, and thereby adding an additional step of retrieval, significantly impacts a model’s ability to maintain reliable performance”
-
✅ 这正是 上下文工程 的价值:重要的不是信息是否在上下文里,而是它怎么被呈现。