Your .md files are a memory architecture (you just don't call it that)
The manual context management patterns that work today map closely onto the agent memory architectures being built for the near future.
I've written two posts now about managing context in AI sessions: one about the problem and one about a practical fix using .md files as external memory. While writing those, I kept running into industry research that pointed at the same thing: the crude file-based system I use daily is basically a hand-rolled version of what agent memory researchers are trying to automate.
It is either validating or embarrassing, depending on how you look at it.
Files as memory architecture
My starter prompt creates five markdown files for every project: CLAUDE.md (project constitution), PROGRESS.md (build log), DECISIONS.md (architecture decision record), TECHNICAL.md (implementation details), and HANDOVER.md (session transition state). Each gets read and written at specific times, with specific rules about what goes where.
I built this by trial and error over a few months. But it maps almost exactly to a pattern described in the agent memory literature as "hierarchical memory with separated stores":
- Long-term knowledge →
CLAUDE.md: rarely changes, loaded every session, defines the project's identity and constraints - Episodic memory →
DECISIONS.mdandPROGRESS.md: records of what happened and why, searchable when you need to understand past choices - Working memory →
HANDOVER.md: the current task state, overwritten each session, what the agent needs right now
The research calls these "different storage and retrieval strategies for different memory types." I call them "files I read in a specific order." It is the same idea in different packaging.
Manual semantic compression
One of the bigger trends in context engineering right now is semantic compression: instead of putting raw conversation logs into the context window, systems generate layered summaries (session-level and topic-level) and keep only those plus the most recent raw exchanges.
That is exactly what HANDOVER.md does at session boundaries. A two-hour session with 150 exchanges gets compressed into a structured document: current task state, decisions made, constraints that carry forward, and next steps. The raw conversation is gone, but the important parts survive.
The difference is that I do this compression manually (or rather, I ask Claude to do it at the end of each session). The automated systems build it into the infrastructure: running compression continuously, clustering by topic, and allocating token budgets across context layers. The core insight is the same. Raw history is a poor format for context, and structured summaries work better. The compression also has to be lossy in the right way, preserving decisions and constraints while dropping the debugging tangents.
The quality of the handover document matters a great deal. A lazy summary ("worked on CSS and deployment") is nearly useless. A structured one ("changed header border to gradient, deployed via Portainer, discovered Docker volume overlays content images, fix: copy to public/images") lets the next session pick up immediately. The automated systems will need to learn this same distinction, and I suspect many will struggle with it at first.
Maximum effective context window
There is a finding from recent research that I wish I had had six months ago: the gap between a model's advertised context window and its "maximum effective context window," the point up to which performance actually holds up.
Models with 200K token windows sometimes start degrading at a few thousand tokens on certain tasks. The exact threshold depends on the task, but the pattern is consistent: there is an inflection point where adding more context starts to hurt instead of help. Attention dilutes, so the model has more text to search through and less ability to find the right piece at the right time.
This explains something I observed but could not articulate: why a fresh session with a good 500-word handover note consistently outperforms a 4-hour session where all the information is technically "in context." The handover note sits well below any model's effective window. The 4-hour session is well past it, even though it is within the theoretical limit.
The practical implication: context is not a bucket you fill until it is full, it is more like a workbench with limited surface area. Keep the active working set small and well-organized, and put everything else in files you can pull in when needed.
Agent-to-agent gap
The research on multi-agent context is where my manual system starts to look primitive. When a team of specialized agents needs to collaborate (a research agent feeding a writing agent feeding an editing agent), they need shared context richer than "pass the output text along."
I built exactly this kind of pipeline last week: an n8n workflow where a research agent searches the web and produces a brief, then a separate writer agent turns that into a blog post. The research agent's output is passed to the writer as raw text in the prompt.
The more sophisticated version would have both agents referencing a shared semantic layer: structured entities with metadata and relationships instead of prose. The research agent would tag its findings with confidence levels, source quality, and timeliness. The writer agent would query that structure instead of parsing unstructured text. The difference is like handing someone a stack of printouts versus giving them access to a well-organized database.
Protocols like A2A (Agent-to-Agent) are trying to standardize this: how agents exchange state, what format context takes when it crosses agent boundaries, and how to avoid the "telephone game" problem where information degrades as it passes through multiple agents. My text-in, text-out pipeline works, but it is closer to agents shouting across a room than sharing a whiteboard.
Context governance
The research mentions "guardrails at the context layer" and "observability around context" almost as afterthoughts. I think this is the most important trend for practitioners.
Right now, most people using AI coding assistants have zero visibility into what is in the model's context at any moment. They do not know what got compressed, what got dropped, or what the model is attending to. When things go wrong (the AI suggests a library you ruled out, or forgets a schema constraint), there is no way to debug why it forgot.
My file-based system provides crude governance by making context explicit and auditable. I can read DECISIONS.md and see exactly what the AI should know. If it contradicts a logged decision, I know the file was not read or was not weighted heavily enough. That is not sophisticated observability, but it beats hoping the conversation history is intact somewhere in the attention mechanism.
The enterprise version of this is what some teams are building: logging what was retrieved or summarized for each response, tracking context composition over time, and tuning retrieval and compression policies based on outcome data. Essentially treating context management as an ML pipeline with its own metrics and optimization loop.
I would like to see this reach individual developer tools. Imagine Claude Code showing a sidebar with "here is what I am currently considering from your project files" and letting you pin or remove items. That would make the context engineering workflow I do manually (reading files, re-anchoring, checkpointing) visible and interactive instead of implicit.
Outlook
Here is what I expect context management to look like in a year:
The manual approach I use (markdown files with explicit read/write protocols) becomes a built-in feature of AI coding tools. The tool maintains the structured memory layer automatically, so you no longer manage the files yourself. Your architecture decisions, progress state, and session handovers get tracked without you writing a prompt that says "update PROGRESS.md."
The semantic compression gets good enough that you stop noticing session boundaries. Right now, starting a new session feels like a reset. With good automated compression and retrieval, it should feel continuous, with the tool always holding the right context loaded, whether that is from five minutes ago or five weeks ago.
The agent-to-agent context problem gets solved with standardized protocols and shared memory stores. Multi-agent workflows would then work more like a team with shared understanding than like duct-taped outputs.
But the core principle will not change: context is a resource to be managed, and a bigger window does not remove the need to manage it. The models will keep getting longer windows, and those windows will keep having effective limits well below their theoretical maximums. The teams that treat context engineering as a first-class concern, whether with markdown files or million-dollar infrastructure, will keep getting better results than those who do not.
For now, I will keep my five .md files, because they work. And apparently I have been doing agent memory architecture all along, without the vocabulary for it.