Context rot has a name now: here's what four months of living with it taught me
A new piece lays out the mechanisms behind context rot and gives concrete token thresholds. I have been hitting those thresholds by hand since February. This post covers where the analysis lands and where my experience pushes back.
Someone sent me a piece on context rot this week. Reading it was like finding a clean schematic for a machine I had been operating by feel for four months.
The article makes a tight argument: context rot (performance degradation as context lengthens) is a real, structural phenomenon, not something a bigger window fixes. It names three mechanisms behind it and puts actual numbers on where it starts to matter. I have been writing about exactly this since February, but always from the practitioner's side: what broke, and the habit that stopped it breaking. The article supplies the part I never had: why it breaks, stated precisely. So this post covers where the analysis confirms what I found the hard way, where it sharpens it, and where my months of living with the problem push back on it.
Context rot mechanisms
I used the phrase "context rot" in my first post on this back in February, almost casually, to describe the drift I kept seeing: by session three or four, the AI suggests a library I ruled out on day one, generates a migration that conflicts with a schema decision from two days ago, and stops following the naming conventions I established. My explanation at the time was hand-wavy: "lost in the middle," attention shifting to recent messages. That was true as far as it went, but it was a description dressed up as an explanation.
The article separates the mechanisms, and the distinction is more useful than my single bucket:
Attention dilution. Transformer attention spreads across all tokens, so every token gets proportionally less weight as the context grows. Early critical information does not vanish; it gets outvoted by the sheer volume of later content. This is the one I had gestured at.
Conflicting instructions. Older instructions do not disappear when newer ones supersede them. They sit there as latent contradictions that surface unpredictably. I had felt this one but never named it, and naming it explains a failure mode that always confused me. The model did not forget my constraint. The constraint is still in there, now competing with three later messages that softened or contradicted it, and the model is adjudicating a conflict I did not know I had created.
Coherence degradation. Synthesizing across dozens of prior exchanges gets harder as they pile up. Responses become "locally sensible but globally inconsistent": each individual answer is fine, and the trajectory drifts. That phrase is the cleanest description I have seen of the slow-motion failure I described in my legal-work example: an AI five months into a case producing a motion that is technically competent and strategically wrong, because it is reasoning soundly over a degraded picture of the whole.
The key point: context rot is not forgetting. The information is right there in the window. The model is confused by accumulation rather than starved of data. I knew this empirically, which is why I kept insisting the fix is not "remind it harder," but I did not have the three-way decomposition to explain why reminding harder sometimes makes things worse (you have just added another instruction to the conflicting pile).
Token thresholds
The article offers practical, explicitly non-binding thresholds from production agent tasks: under ~20K tokens is generally fine; 20K–50K shows noticeable degradation on complex multi-constraint tasks; past ~50K, rot becomes a real reliability problem.
In my second post I wrote about the "maximum effective context window," the gap between a model's advertised window and the point where its performance actually holds up, and said I wished I had had that concept six months earlier. What I did not have was numbers. I could tell you a fresh session with a 500-word handover note beat a four-hour session where everything was technically "in context," but I could not tell you where the cliff was.
These thresholds line up with my experience closely enough that I trust them as a starting heuristic, with one caveat the article is careful to flag and I want to amplify: they are task-shaped, not universal. My CSS-debugging sessions tolerate far more accumulation than my multi-constraint architecture sessions, because the former is a sequence of locally-scoped problems and the latter demands global coherence, the dimension that degrades first. The number that actually matters is not tokens but how many live constraints the current task must hold simultaneously. A 60K-token session reading mostly inert reference material can be fine; a 30K-token session juggling eight interacting decisions is already in trouble.
So I would reframe the thresholds as a proxy for constraint load: they hold until your task gets constraint-dense, at which point the cliff arrives earlier than the token count predicts.
Mitigations and my .md files
The article's engineering recommendations (compression, pruning, session splitting, structured curation) are, almost line for line, the system I described building out of markdown files in February. I find this more reassuring than redundant: two people reasoning from different directions (it from mechanism, me from scar tissue) converging on the same four moves suggests the moves are right.
- Context compression → my
HANDOVER.md. A two-hour session with 150 exchanges distilled into a structured state document at the session boundary. The article's "summarize earlier exchanges" is the automated cousin of the thing I ask Claude to do by hand at the end of every session. - Context pruning → "know when to reset." The most counterintuitive habit I have, and the one that helps most: when a session starts making subtle errors, the best move is to end it, generate a handover, and start clean. That is pruning taken to its logical extreme: prune everything, keep only the distilled state.
- Session splitting → "modularize aggressively." Break the work into discrete tasks with clear interface contracts, finish one, save the result externally, start the next with only the relevant context. The next module does not need the debugging history of the last one; it needs the current contract.
- Structured context management → the entire premise of the starter prompt. Actively curate what enters context instead of letting it grow organically.
The one thing I would add from experience: compression quality decides everything, and it is lossy in a way that is easy to get wrong. A lazy handover ("worked on CSS and deployment") is nearly useless. A good one ("changed header border to gradient; deployed via Portainer; discovered Docker volume overlays content images, fix: copy to public/images") lets the next session start instantly. The article treats compression as a strategy; I would treat it as a skill with a sharp quality gradient. The decisions and constraints have to survive, and the debugging tangents have to be dropped. Automated compression that summarizes by recency or token-proportion rather than by durability of the information will compress the wrong things: it will keep the recent CSS noise and drop the load-bearing schema decision from day one. That is the same prioritization failure that causes context rot, relocated into the tool that is supposed to fix it.
Extensions to the argument
The article is strong on the single-session problem. The last four months convinced me the more interesting frontier is what happens when you stop fighting accumulation and start architecting around it.
Context rot is the reason agent teams work. Once you accept that a single context degrades as it accumulates constraints, the case for decomposition stops being about parallelism and starts being about rot avoidance. I argued in March that agent teams need explicit context boundaries the way human teams do. The mechanism the article describes is why: a discovery agent that only scans and inventories never accumulates the constraint load that produces coherence degradation, because its context stays scoped to one concern. The monolithic "do everything" agent is a context-rot generator by construction, which is a worse problem than its speed. Splitting concerns across agents is session-splitting applied to the org chart. The article's fourth mitigation, scaled up, is the multi-agent architecture.
The platforms are now absorbing every one of these mitigations. When I wrote about Claude Fable 5 last week, what struck me was that the new tier did not grow the context window; it shipped the management layer. Task budgets let the model pace itself against a known limit. Server-side compaction automates my HANDOVER.md pattern. File-based memory produced a 3× jump on long-horizon tasks. The four mitigations in this article are the capabilities the frontier labs are building into the infrastructure. The article frames context governance as something you must do; I would add that it is becoming something the platform increasingly does with you, and the practitioners who internalized the discipline by hand are the ones who will know how to drive those features, because they understand what the features are for.
Observability is still the missing piece. The one place neither the article nor the platforms have caught up to where I wish they were: I still cannot see what is actually in the model's context at any given moment. When a session goes sideways, I cannot inspect which mechanism fired, whether a constraint was diluted, conflicted with a later instruction, or coherence simply decayed. I diagnose by symptom and reach for the same reset either way. The three-mechanism decomposition is useful here, because it gives me a vocabulary for what to look for, but until tools surface context composition the way they surface a stack trace, context engineering stays a craft practiced partly blind.
The unchanged conclusion
I have ended several posts in this series on the same sentence, and this new piece does not move me off it; it reinforces it from the mechanism side: context is a resource that needs to be managed, not a bucket that needs to be bigger.
What the article adds is the why, decomposed cleanly enough to be actionable. What four months of doing it by hand adds is the texture: the thresholds bend with constraint density, the compression is a skill with a steep quality gradient, the single-session mitigations scale up into agent architecture, and the platforms are racing to absorb all of it. The two accounts reach the same conclusion from opposite ends. When the theory and the practice agree this closely, it is usually because the conclusion holds.