Context hygiene: managing what an AI keeps across sessions
The bottleneck in long-running AI sessions is not model intelligence. It is context management. Here is what I have learned about keeping an AI on track across complex, multi-session projects.
The more I work with LLMs on real projects, the more the core skill looks like context management. Prompting matters less than keeping the model's context correct.
We talk a lot about which model is smarter or which tool writes better code. In practice, what makes or breaks a long-running AI collaboration is whether the model still remembers what you told it three hours ago. Usually it does not.
Silent context loss
Every LLM has a context window, a fixed amount of text it can hold at once. Gemini 1.5 Pro claims 2 million tokens. Claude has 200K. Having a large context window and using it well are different problems.
Researchers call the effect "lost in the middle." Models recall information at the beginning and the end of a conversation better than the middle, and the middle is where most of the work lives. The model does not forget in the human sense. It stops attending to that part of the input.
When the conversation exceeds the window, the system compresses it. It summarizes the history and decides what is important. If your architecture constraints get compressed into "user is building a web app," you are working on a degraded version of your own instructions, and you probably will not notice until something breaks.
This is silent context loss. It does the most damage of any failure mode in AI-assisted work, because the model looks like it is still following your instructions while it is acting on a worn-down copy of them.
Software projects
I have been building projects with AI coding assistants for a while. The pattern repeats.
The first session goes well. You lay out the architecture, define the stack, set constraints. The assistant follows your patterns and generates code that fits. You feel productive.
By the third or fourth session it drifts. You ask for a feature and the assistant suggests a library you ruled out in session one. It writes a migration that conflicts with a schema decision from two days ago. It drops the naming conventions you set.
The model is not getting worse. The early decisions, the ones that matter most, are buried under hundreds of later messages about bug fixes and small tweaks, and the model's attention has moved to the last few exchanges. This is context rot.
I hit it building this blog. By the time I was debugging CSS, the assistant had lost the original design-system decisions. They were still in the history, but they were competing with dozens of newer messages about Docker volumes and git lock files.
Better prompts do not fix this. External state the model can read does. For code that means a living document, an architecture decision record, a progress file, a CLAUDE.md, that holds the decisions that matter. Each session starts by reading that file rather than trusting conversational memory.
Legal work
The stakes rise when you move from code to consequential decisions.
Consider using AI across several months on a complex legal case. You start by uploading contracts, case law, depositions, hundreds of pages of source material. The AI helps identify precedents, draft motions, analyze opposing arguments. It is genuinely useful.
Legal cases evolve. New evidence surfaces. Rulings narrow what is admissible. Strategy pivots. Each development pushes the original foundation, the contract clauses, the jurisdictional nuances, the initial theory of the case, further into the middle of the context where attention fades.
Five months in, you ask for a response to a new motion. The AI produces something competent but strategically wrong, because it has forgotten a preliminary ruling from month two that limited the exact line of argument it is now proposing. Or it cites a contract clause without the interpretation you fixed early on.
In legal work the AI hallucinates by omission. It does not invent facts, usually. It forgets constraints. A legally sound argument that ignores a prior ruling is worse than wrong; it is potentially malpractice.
The pattern matches the coding case, and so does the fix: hierarchical context management. The core facts, parties, jurisdiction, theory, key rulings, stay in a reference document injected at the start of every session. Evidence is retrieved when needed. The AI is not trusted to hold the strategy on its own.
Practical context hygiene
After enough sessions going sideways, I have settled on a few principles that help.
Keep a state file. For any project spanning multiple sessions, maintain a document with the current state: decisions made, constraints set, what is done, what is next. Give it to the AI at the start of every session. Do not trust conversational history to carry this forward.
Re-anchor before big asks. Before requesting anything significant, a new feature, a strategic decision, a complex analysis, restate the constraints that matter. "Using our established PostgreSQL schema with the audit-logging pattern from week one, design the new endpoint." A few tokens here saves hours of fixing drift.
Checkpoint regularly. Every ten to fifteen exchanges, ask the AI to summarize its current understanding of the project state. This forces it to consolidate, and it gives you a chance to catch misalignment. A wrong summary shows you the context rot before it does damage.
Know when to reset. This one is the least intuitive. When a conversation gets long and the AI starts making subtle errors, end the session. Have it write a handover note with the current state, pending tasks, and established constraints, and start fresh. A clean context window with a good handover note beats a bloated conversation where the important parts are buried.
Modularize aggressively. Break complex work into discrete tasks with clear inputs and outputs. When a module is done, save the result externally, and start the next one with only the relevant context. The AI does not need every debugging session. It needs the current interface contract.
Context engineering
We are past the point where prompt engineering meant a clever one-shot instruction. The skill now is context engineering: managing the information across a whole project, deciding what the AI needs now, what can be retrieved later, and what can be safely dropped.
Models will keep improving at this on their own. Agentic retrieval, where the AI searches its own files and history, already helps. For now the human maintains context hygiene. The model is a capable collaborator with a specific kind of amnesia, and the results come from working with that limit.
People who treat context management as a first-class concern get steady results from these tools. People who leave it to the model repeat the same conversation and see the same forgetting.