A Frustrating Phenomenon
Anyone who's used AI coding tools for complex features has probably experienced this: at the start of a conversation, the model performs like a genius — clear thinking, clean code. But as the conversation grows longer, it starts making the same mistakes repeatedly, even reproducing bugs you've already corrected.
You explicitly tell it "don't use this API," it politely apologizes, then uses it again in the next generation. You copy the entire context into a new conversation, and it suddenly gives you the correct solution. The model hasn't degraded — the context itself is rotting.
This topic has been generating a lot of buzz in AI engineering circles lately. Some call it Context Rot. After digging into it, I found it actually encompasses two completely different mechanisms, and the industry's responses to each are strikingly different.
First Type: Attention Dilution
The first mechanism is relatively easy to understand. Although large model context windows keep expanding — from 4K to 128K to millions of tokens — "fitting" and "using well" are two different things.
Transformer's attention mechanism dilutes signals when processing long sequences. It's like being in a two-hour meeting: the key decision made in the first 30 minutes might already be buried under subsequent discussions by minute 90. Models are the same: early important information gradually loses its "presence" in attention weights under the flood of massive tokens.
This isn't a bug — it's determined by the mathematical nature of the attention mechanism. When processing each token, the model needs to "look back" at all previous tokens. When "previous" becomes very long, each token gets less attention. The performance degradation curve often starts well before reaching the context limit.
The good news: this problem has clear engineering solutions. Anthropic introduced progressive disclosure in their tech stack, using context compression, summary generation, and Context Folding to shorten effective context length. The results are remarkable — reportedly reducing token consumption by 84% while improving agentic search performance by 39%. Claude's /compact command is a product of this same approach.
Second Type: Context Poisoning
The second mechanism is more insidious and far more difficult to deal with.
It happens like this: the context isn't long — maybe only 10-20K tokens, far from attention dilution levels. But the problem is that an incorrect judgment appeared early in the conversation — the model made a wrong assumption about some API behavior, or misunderstood the requirements — and this incorrect judgment gets repeatedly cited and reinforced in subsequent turns.
The model continues reasoning on a flawed foundation, like building additional floors on a crooked foundation. Each layer looks "reasonable," but the entire structure has veered off course. Even worse, when you try to correct a specific error, the model finds "evidence" from those poisoned early contexts to defend the mistake.
This is why asking AI to fix a bug makes things worse. Not because the model isn't smart enough, but because it's using wrong premises to understand your correction instructions. It's not "fixing" — it's "interpreting your fix in the wrong way."
Currently, there's essentially no good technical solution for context poisoning. The most effective method is frustratingly simple: start fresh. Open a new conversation and re-describe the clean requirements. Without those "toxic" contexts, the model can actually give you the correct solution.
The Fundamental Difference
On the surface, both types of rot cause the model to "get dumber," but their root causes are completely different:
- Attention dilution is a mathematical limitation — an inherent flaw of the attention mechanism on long sequences, addressable through engineering (compression, summarization, segmented processing). It's a "quantity" problem, fundamentally optimizable
- Context poisoning is a cognitive deficiency — it requires the model to have "metacognition" ability: recognizing that its foundational assumptions are wrong and proactively dismantling established reasoning chains. Current Transformer architectures essentially lack this capability
Analogy: attention dilution is like someone who can't hear you clearly in a noisy environment — you can solve it with noise-canceling headphones (compression techniques). Context poisoning is like someone stubbornly believing a false fact and building extensive reasoning on it — it's nearly impossible to correct that initial error without making them "forget" all the downstream deductions.
Humans have this problem too, but we have a key advantage: we can realize "wait, my earlier assumption might be wrong." This metacognitive ability is what current large models lack most.
The Sub-Agent Debate: Two Industry Camps
To combat context poisoning, major companies have widely adopted an architectural strategy: sub-agents. The core idea is to slice a large task into pieces, each handled by a brand-new, clean sub-agent. The sub-agent returns only results, discarding all intermediate context.
This is essentially engineered "start fresh" — since a single conversation's context will rot, don't let any conversation live long enough for it to happen. Each sub-agent handles a task small enough and context short enough that neither type of rot has time to develop.
But this strategy has sparked considerable controversy in the industry.
Anthropic's position is clear: multi-agent systems are the future. They claim their multi-agent systems achieve 90% performance improvement over single agents, with clear advantages on complex tasks. From an engineering perspective, this makes sense — using clean contexts for sliced tasks avoids both attention dilution and context poisoning.
But Cognition (the Devin team) pushed back. They publicly argued "don't build multi-agent systems," claiming sub-agents don't solve problems but create new ones — how do you determine task slicing granularity? How do you synchronize information between sub-agents? How do you guarantee global consistency? Every new agent boundary is a new potential failure point.
I think both sides have valid points. Sub-agents are indeed an effective weapon against context poisoning, but they're not a silver bullet. They transform a "context management" problem into a "distributed coordination" problem — and the latter has never been an easy problem in computer science.
Final Thoughts
The Context Rot concept deserves attention because it reveals a fact many haven't realized: context window size doesn't equal effective utilization length. A 128K window doesn't mean you can dump 128K of content in and expect the model to understand it all.
As a developer who uses AI tools daily, my coping strategy is simple:
- Compact regularly — proactively compress once conversations exceed a certain length, don't wait for the model to start degrading
- Start fresh on wrong assumptions — once you detect the model reasoning on incorrect premises, don't try to correct it in the current conversation, just open a new one
- Front-load critical context — place the most important constraints and requirements at the beginning or end of your prompt, leveraging the attention mechanism's U-shaped distribution
Context rot isn't model degradation — it's another cognitive calibration of model capability boundaries. Rather than complaining "AI is getting dumber," understand why it happens and use it correctly.
After all, the tool doesn't get dumber — we just haven't learned how to talk to it properly yet.