A rule nobody checks is a decoration
I've been running my context-management discipline on a real client programme where most of the execution is done by agents. The principles held up. The rules that were only written down decayed; the ones a script checked before every push held.
Everything I've written on this blog about context management came from projects where the only person who could be hurt by drift was me. The .md files, the handover discipline, the agent-team split, the operating model for consulting work: all of it was built on my own scar tissue, on my own projects, at my own risk.
For a while now I've been running the same discipline somewhere considerably less forgiving: a live client engagement migrating a large multi-brand web estate off a legacy CMS onto a major enterprise CMS platform. Dozens of market sites, thousands of pages, a very small team, and most of the execution done by agents rather than people. It ran for months, with multiple contributors, real money, and a real acceptance bar.
The core principle held up exactly as advertised. The memory lives in files, not in the conversation, and that is what makes months of agent work feel continuous instead of episodic.
Almost every rule in that repo exists because something broke first. The pattern in what broke was consistent enough to be worth writing down: the rules that were merely written down decayed, and the rules that were mechanically checked held. A convention nobody measures is a decoration. It looks like governance, it reads like governance in the onboarding doc, and it is quietly false within about two weeks.
This post is the lessons that cost the most to learn.
1. Repo as the memory
The operating file for that programme contains a sentence I now put in every project I run:
An agent's private memory is never the only home of a project fact.
It sounds like bureaucracy. It is actually an availability requirement, and the reason is that agent sessions die in ways that human collaborators don't. A session compacts and quietly re-summarises the thing you cared about. A machine restarts mid-task. A tool crashes and takes an hour of reasoning with it. A second agent picks up the lane tomorrow with no history at all, and so does the human who joins next month.
Every one of those events silently deletes anything that lived only in a model's head, and silently is the operative word, because nothing announces the loss. You find out later, when a decision gets re-made differently.
So the structure is deliberately boring. A one-page state snapshot read first in every session. A rules file that self-configures whoever opens the repo, human or agent. A daily log that gets appended at the end of every working session, and (this is the part I like) the rules file tells agents explicitly why it isn't optional:
Every working session ends with a status entry: what changed, what's blocked, what's next. Agents: this is not optional; it's how humans audit you.
That framing turned out to matter more than the file format. An agent writing a status entry for itself writes a diary. An agent writing a status entry knowing a human will use it to check its work writes something you can actually audit. The session-end ritual is four steps and takes two minutes: status entry, check the state snapshot still matches reality, run the doc linter, push.
This is the folder-is-the-memory idea from my consulting operating model, with one addition that consulting work never forced: the writer is not the reader. In a solo project I write context files for future-me, who has some residual memory of what happened. Here, the reader is a stranger: a fresh agent, a new colleague, me in six weeks. You can't write for a reader who shares your assumptions, because they don't.
There's a harder version of the same rule, and it was bought expensively. Four agents were dispatched on a batch of fixes. All four reported completion. None of them had committed anything. They'd stalled silently, and nobody noticed for six days, at which point it surfaced during a client-facing review. The protocol that came out of that is now binding, and it's one line: the deliverable is a file on disk. Every task names its exact output paths up front, and the handover check is ls and git log against those paths, never the agent's own account of what it did. An agent's summary of its work is a claim, not evidence. It's produced by the same process that produced the work, and it fails in the same direction.
2. Current truth only
This is the rule I expected to dislike and now defend hardest.
Every document states the CURRENT truth only, no layered history. Do not add "superseded", "was previously", or dated correction banners to a working document; rewrite it so it is simply correct.
Every instinct trained by version control and by professional caution says the opposite: leave the trail, mark the correction, show your working. I wrote correction banners for years. They feel responsible.
They are, at scale, actively harmful, and the reason is a mechanism I already wrote about without connecting it to documents. In the context-rot post I picked up a three-way decomposition of why long contexts degrade, and the middle one was conflicting instructions: older instructions don't disappear when newer ones supersede them, they sit there as latent contradictions and the model ends up adjudicating a conflict you didn't know you'd created.
A "superseded 2026-06-04" banner is a conflicting instruction with a date stamped on it. You are trusting the reader to apply recency as a tiebreak. Models do not reliably do that; they weigh the whole document. Neither, for that matter, do tired humans at 17:40 on a Thursday. And it compounds: every banner doubles the reading cost of the paragraph it sits next to, and the doc gets read hundreds of times before anyone touches it again.
So history is not abolished; it is given exactly one home per kind. Changed doctrine (what we believed, what disproved it, what replaced it) lives in one narrative file and nowhere else. Decisions live in numbered decision records. The daily log and dated investigations are history by design. Retired documents go to an archive. Everything else states current truth.
The doctrine file is the one I'd steal first. It's a flat list of entries in exactly three beats: Believed / Found / Now. For example: believed that the reference screenshots we were checking builds against were a sound baseline; found that two of them were error pages, byte-identical to each other, never quarantined, and filed on the work queue as a real defect; now there's a tool that clusters captures to find error pages before they can become ground truth.
That's a good context artifact for a reason that took me a while to see. It is not a changelog. A changelog tells you what changed, which is almost never what you need. This tells you which of your current beliefs were expensive to acquire, exactly what a fresh reader needs in order not to repeat the mistake, and exactly what a compressor would throw away first. It's the same argument I made about handover quality: compression is a skill with a steep gradient, and the failure mode is compressing by recency instead of by durability. A doctrine file is durability-first compression, done by hand, once.
3. One source of truth per fact
"Link, don't duplicate" is easy to write and impossible to sustain by goodwill. In a repo of a few hundred working documents, facts leak. A figure gets restated in a deck, a path gets copied into a README, a definition gets paraphrased in a deliverable draft. Each copy is fine on the day it's written and wrong within a month.
Two mechanisms fixed this, and the second one is the interesting one.
A concept-ownership index. A single table: concept → the one document that OWNS it → the list of documents that restate it. Change an owner document, and you walk its row in the same session. Not "eventually", not "file a follow-up": the same session.
The reason that table exists is a failure worth describing precisely. A rename sweep was done by string search, and it missed a major client deliverable, because that document described the concept in its own words rather than in the exact string being swept. Search finds the copies that were lazy. It cannot find the copies that were well written. So the rule became that every sweep runs two filters, literal strings and the concept-ownership rows.
That distinction (the paraphrase is invisible to grep) is one of those things that seems obvious once stated and had cost two multi-hour recovery sweeps in a single week before anyone stated it.
A documentation linter. This is the part I'd push hardest on anyone doing serious agent-assisted work. It's a script that must exit zero before anyone pushes, and it checks the mechanical half of the doc rules: dead path references, cross-checks between the state snapshot's headline figures and the generated file those figures come from, freshness on the read-first documents, that every folder has a README, that generated mega-files carry the head stamp telling agents to query them rather than read them whole.
The first time it ran, in a repo where everybody sincerely believed the rules were being followed, it found 22 dead path references and 9 missing READMEs.
That number is the whole argument. The rules had been written for months, they were good rules, and everyone agreed with them. Compliance was still somewhere around "most of the time", because nothing measured it, and it was slowly getting worse, because the drift is invisible to the person creating it.
The companion rule is stated just as bluntly: a stale file is a bug. If your work makes a document wrong (a number, a plan, a risk, the state snapshot) updating it is part of the same task, not a follow-up. Follow-ups are where documentation goes to die, and an agent will absolutely accept "I'll note that for later" as a completed step unless you close that door explicitly.
4. Instrument verification
This is my favourite lesson of the whole programme, and it generalises far beyond it.
An audit found that a linter had never actually run on two of the builds. Not "ran and passed": never executed at all, because a dependency was missing from those repos. A crashing linter prints no errors, which is indistinguishable from a clean run. The command exited, no complaints appeared, the log looked like every other green log. It had looked like that for weeks.
It got worse on inspection. The check was two commands chained with &&. The crashing one was the first. So the second half (an entire category of checks, on every site, not just those two) had never run either. One silent failure had been hiding a second silent failure behind it. And when the linter was finally fixed, it immediately found a real defect that had been shipping for weeks: fixing the measuring instrument found the bug.
Nobody had done anything wrong. Everybody had "run the checks".
The rule that came out of it is now a standing line in my own notes: assert the instrument, never its silence. A green result is not one claim, it's two: the thing is fine, and I actually looked. The second claim is the one that fails quietly, and it fails in exactly the shape of success.
Once you start looking for this shape you find it everywhere. A verification script full of "this must be zero" assertions was passing every one of them, because the patterns had been written with alternation but without the flag that enables it, so none of them matched anything, ever. It was caught only because someone had added a deliberate control line: this pattern MUST find something. When the control went quiet, the whole script was exposed. Underneath it, a real failure had been hiding.
So the practical fixes are two. Give every checker a control case that proves it can still find something. And before trusting a gate, seed a defect you know it should catch and confirm it goes red. The programme runs a self-test that seeds nine known defects and requires nine detections, and the instruction is blunt: run it before any sweep. It costs seconds. It converts "the checks passed" from a belief into a measurement.
There's a layer above this that I now watch for specifically: written is not executed. A mitigation can be designed, agreed, implemented, and documented as done, and still have never run once, because nothing had triggered it since the code landed. That was true of one fix here for weeks. The document said the control existed. The control did exist. Its coverage was zero. If you only ever audit whether a safeguard has been built, you will keep finding that it has.
This matters more with agents than without, for a reason that's easy to miss. An agent reporting on its own verification is reporting on a tool it did not write, whose failure modes it cannot see, in a summary it composes itself. Every layer there is happy to render silence as success. If you've read my context-rot follow-up, I ended it complaining that observability is the missing piece, that I still can't see what's actually in the model's context or which mechanism fired when a session goes sideways. Seeding a defect and watching for red is what you do while you're waiting for the observability you don't have. You can't see inside the instrument, so you poke it and check it flinches.
5. Comparison baselines
The single most useful sentence in that programme's method notes is this one:
A measurement that shows no effect is suspect before it is believed.
It comes from a specific embarrassment. A calibration sweep changed a design token to two different values and got byte-identical results both times. The natural reading of that is a finding: this token doesn't affect anything. The correct reading was that the harness was broken: the edit was landing on one copy of a stylesheet while the dev server was serving a different copy.
Chasing it uncovered something much worse. A shared component library and the per-site deployed copies of it had drifted apart, in both directions, for a long time: the library was ahead in some ways, the sites were ahead in others, and roughly half the shared files differed. The whole economic argument for a shared library is "fix it once and it propagates everywhere", and that had quietly not been true for months. Nothing had ever checked it, because the rule that forbade it was, again, written down rather than measured.
Related lessons from the same list, all bought the same way:
- Assert that the input changed, not just that the output didn't.
- State the expected result before you run the check. Two broken extractions were caught only because someone wrote down "this should change zero pages" first. Without the prediction, "zero pages changed" reads as a clean pass.
- Count the thing, not the files. Counting rendered images instead of pages inverted a coverage reading: "this estate more than doubled" was actually "this estate lost eight pages".
- Locate things by content, never by your own naming. A probe that searched the source site for one of our CSS class names (a class that by definition doesn't exist in the source) reported a spectacular and entirely fake result.
There's a common shape here, and it's specifically an agent risk. Agents are excellent at producing a confident, well-argued conclusion from whatever measurement they were handed. What they don't have is the practitioner's flinch, the small itch that says that number is too clean, that's suspiciously round, I've never seen this thing come out at exactly zero before. A wrong baseline doesn't produce an error. It produces a tidy answer, delivered with the same tone as a correct one, and it will survive review because reviewing the conclusion is not the same as reviewing the baseline.
The engineering notes for that programme eventually grew a line I've adopted wholesale. Over one week, six separate headline figures looked conclusive and turned out to mean something else entirely. One of them was arithmetically impossible on its face, reporting more exact matches than there were things to match. Every one of them was caught the same way: by opening one specific case and looking at it. Hence the rule. Treat any aggregate you have not traced to a concrete instance as provisional.
Unreachable invalid states
Once you've been burned enough times by rules that were only written down, you start looking for places where the rule can stop being a rule.
The best example in the programme was a class of defect that policy had forbidden for months and that kept reappearing anyway: page-specific styling, which destroys the reusability the whole approach depends on. Telling agents not to do it worked about as well as telling anyone not to do anything.
What worked was inverting the structure. The authoring model declares what a component can express. The implementation implements it. A parity tool asserts both directions: nothing implemented that isn't declared, nothing declared that isn't implemented. At that point a page-specific class isn't forbidden, it's unauthorable, and a thing that cannot be authored cannot exist. The rule became part of the architecture.
The library-drift problem got the same treatment from a different angle: a ratchet. The measured number of diverged files is checked in, and the check fails only if it goes up. It doesn't demand the drift be fixed today. It just makes it structurally impossible for the situation to get worse while everyone is busy, which, empirically, is when it gets worse.
The design choice there is the interesting part. A hard assertion (zero drift or the build fails) would have been more principled, and would have parked every piece of work behind a day of adjudication. Which in practice means somebody switches the check off, and then you have neither the check nor the principle. A gate that fails on things people have legitimately decided to live with gets disabled, and a disabled gate is worse than no gate, because everyone still believes it's running.
The same logic killed an earlier attempt at safety here. A tool that could destroy results got an opt-in isolation flag: remember to pass --isolate and you're fine. Hours later, the same person who added it ran the tool without it and wiped several hundred pages of results. The note in the log is the correct conclusion: an opt-in safety flag is documentation of a hazard, not a fix for one. It protects the careful case and leaves the careless case exactly as exposed as before. The tool now isolates by default and requires an explicit act to do the dangerous thing.
One case is the sharpest evidence for the argument here. A decision record was written to forbid a specific dangerous operation, and the next day the person who wrote it performed exactly that operation, reverting real work, caught only by a rendering check. The log entry is admirably blunt about it. The record exists because this is easy to do; knowing that was not sufficient protection.
And the cheapest version of the same move: decisions become numbered records. Anything that constrains future work gets a numbered file with context, decision, consequences, a date and a decider, and supersession is explicit. This is the least glamorous practice on the list and possibly the highest return per minute, because an agent with no memory of an argument will cheerfully re-open it, and be persuasive, and be wrong. A numbered decision record is a hard stop: this was settled, here's why, and if you want to change it you supersede it on the record rather than drifting away from it in a side conversation.
That's the general principle I'd take to any agent-heavy project: remembering harder is not a control. A rule an agent has to hold in mind competes for attention with everything else in the window, and it loses on exactly the day you're busiest. A rule expressed as a state the system cannot reach doesn't compete with anything. Every time you can convert the first into the second, do it, and where you can't, at least make the safe path the default.
Caveats
I want to be careful about the tone here, because there's a version of this post that reads as "we pointed agents at a large migration and it went great". That's not the story, and the speed isn't the interesting part anyway.
The interesting part is that agent execution produces work faster than any human review process can absorb it, which means the only thing standing between you and a large volume of confident, plausible, wrong output is the quality of your checking. Every rule above is a response to something that got through. The linter that never ran, the baseline that was an error page, the library that had forked, the measurement that showed no effect because the harness was broken, all of those were found late, by accident, by someone noticing a number that looked wrong.
Discipline isn't overhead on agent execution. It is the deliverable, in the same sense that the memory files were the real deliverable on my boiler project. The build is the demo. The thing that makes the build trustworthy, and makes the next change to it fast and safe, is the context, written down, owned in exactly one place, and continuously checked by something that isn't a person's good intentions.
Transferable list
Strip out the domain and this is what's left:
- The repo is the memory. A fact that lives only in a model's context does not exist. Write it down in the same session, or accept that it's gone.
- The deliverable is a file on disk, not an agent's account of having produced one. Check the artifact, never the summary.
- Current truth only. Give history exactly one home per kind and keep working documents clean. A superseded banner is a conflicting instruction with a date on it.
- One owner per concept, with a sweep list. And run every sweep twice: on strings, and on concepts. Grep can't find a good paraphrase.
- Lint the context. Rules that aren't measured decay silently, in a way that's invisible to the person causing it. Make the check a gate, not a habit.
- A stale file is a bug, fixed in the same task. Follow-ups are where documentation goes to die.
- Assert the instrument, not its silence. Seed a defect, watch it go red, then believe the green. And check that your safeguards have actually run: written is not executed.
- Know what you're comparing against. A measurement that shows no effect is suspect before it is believed; state the expected result before you run the check; treat any aggregate you haven't traced to a concrete instance as provisional.
- Make invalid states unreachable, and where you can't, make the safe path the default. Remembering harder is not a control.
I've ended most of these posts on the same line, and I'll keep it, with one addition earned here: context is a resource that needs to be managed, not a bucket that needs to be bigger. And past a certain scale, managed has to mean mechanised. A rule you have to remember will lose the day you are busiest; a rule that runs before you push does not depend on anyone remembering it.