上下文腐坏(Context Rot)

一句话定义:模型的表现随输入长度增加而持续、非均匀地退化——不是到某个长度突然崩掉,而是”越长越不可靠”,且退化模式与位置、干扰项、文本结构都有关。

术语由 Chroma 的 2025 年技术报告确立(ContextRot技术报告)。


1. 核心现象

“models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows.”

报告的测试规模:18 个模型 × 8 种输入长度 × 11 个答案位置,涵盖 Claude Opus 4 / Sonnet 4、o3 / GPT-4.1、Gemini 2.5 Pro、Qwen3-235B-A22B 等。

2. 五条具体规律(逐条来自原文)

  1. 普遍衰减:“Across all experiments, model performance consistently degrades with increasing input length.”
  2. 相似度越低衰减越快:“Lower similarity needle-question pairs increases the rate of performance degradation.”
  3. 干扰项非均匀影响:“Distractors have non-uniform impact … Even a single distractor reduces performance relative to the baseline, and adding four distractors compounds this degradation further.”
  4. 相似度的影响不是统一的
  5. 结构本身有影响:“The structural pattern of the haystack consistently shows an impact on how models process long inputs.”

3. 最反直觉的一条:连贯文本比打乱文本更难

“Surprisingly, we find that structural coherence consistently hurts model performance.”

文本的连贯结构会持续损害模型的表现。这和直觉完全相反,是这份报告最有名的发现之一。

“models perform worse when the haystack preserves a logical flow of ideas. Shuffling the haystack and removing local coherence consistently improves performance.”

当草堆文本保持逻辑连贯时,模型表现更差;把句子打乱、去掉局部连贯性,表现反而稳定提升。可能的解释是连贯文本更容易让模型被上下文带偏。

“Across all 18 models … models perform better on shuffled haystacks than on logically structured ones.”

一个可能的解释:连贯文本里”看起来像答案”的干扰更多;打乱后局部线索与全局语义脱钩,反而减少了干扰。这也提醒我们:把文档切碎再喂,未必是坏事。

4. 真实场景的对照实验(LongMemEval)

“We use LongMemEval_s and filter for tasks … to end up with 306 total prompts. These prompts average out to ~113k tokens.” 而聚焦提示 “average out to ~300 tokens”。

他们从 LongMemEval_s 里筛出 306 条提示,平均长度约 11.3 万 token;而对应的「聚焦版」提示平均只有约 300 token。

“We verify that the models are highly capable of succeeding on the focused inputs, then observe consistent performance degradation with the full inputs.”

先确认模型在短提示(约 300 token)上确实能做对,然后观察到换成完整长提示(约 11.3 万 token)时,表现会出现一致性下降。

“adding irrelevant context, and thereby adding an additional step of retrieval, significantly impacts a model’s ability to maintain reliable performance.”

这是”先检索再喂”优于”整包丢进去”的最直接证据。

5. 模型家族之间的行为分化

“Claude models consistently exhibit the lowest hallucination rates. … they tend to abstain when uncertain, explicitly stating that no answer can be found. In contrast, GPT models show the highest rates of hallucination, often generating confident but incorrect responses when distractors are present.”

这里的启示:面对长上下文时的”拒答倾向”是一个独立的模型特质,与知识量无关。选模型时要按你的容错需求来定。

6. 独立交叉验证

研究关键数字
NoLiMa(arXiv:2502.05167)13 个标称 ≥128K 的模型中,11 个在 32K 就掉到短上下文基线的 50% 以下;GPT-4o 从 99.3% → 69.7%
RULER(arXiv:2404.06654)17 个模型中,标称 ≥32K 的只有一半能在 32K 保持令人满意的表现;vanilla NIAH 上却几乎满分

7. 工程对策(一句话)

“Whether relevant information is present in a model’s context is not all that matters; what matters more is how that information is presented. We demonstrate that even the most capable models are sensitive to this, making effective context engineering essential for reliable performance.”

具体做法见 上下文工程:先检索、去重、压缩、按相关度重排、把关键信息放在开头或结尾(模型对序列开头最敏感——“Accuracy is highest when the unique word is placed near the beginning”)。

相关页面