Measuring my agent harness against plain Claude Code
I built a multi-agent harness and ran it head to head against a plain control on a real benchmark. On these tasks it cost about five times more and did not score better. What it bu...
I built a multi-agent harness and ran it head to head against a plain control on a real benchmark. On these tasks it cost about five times more and did not score better. What it bu...
The AI Art Magazine is running an open call where agents submit the work and agents judge it. I registered one and let it choose. It didn't choose the piece I wanted, or the one it...
I've been running my context-management discipline on a real client programme where most of the execution is done by agents. The principles held up. The rules that were only writte...
A new piece lays out the mechanisms behind context rot and gives concrete token thresholds. I have been hitting those thresholds by hand since February. This post covers where the...
Anthropic launched Claude Fable 5 yesterday, its first Mythos-class model, priced at double Opus. The context window did not grow. What changed is the management layer around it.
I've been experimenting with feeding agents larger volumes of context. The pattern is steady: the bigger I make it, the worse it looks. Meanwhile, the rest of the field is racing t...
Two months of running the agent-team starter prompt taught me which rules actually hold. The ones wired into tooling survived. The ones that only lived in prose were the first to e...
I've spent months building context management patterns by hand. Here's how the four major platforms are solving the same problems at the infrastructure level, and what it means for...
Anthropic shipped a memory tool that looks structurally identical to the markdown files I've been maintaining by hand. Here's what each approach gets right, where they diverge, and...
We've been solving the 'too complex for one person' problem for decades in program delivery. Now some team members are language models, and the coordination overhead is measured in...
Six years ago I helped a CPG company design a smart meal planner. We never shipped it. This year my wife and I built it, and we use it every week.
Running three AI agents sounds expensive. It is cheaper than running one, if you put the right model in each role.
Splitting a single AI agent into a team changed more than the workflow. It changed how context flows through the whole system.
My AI coding setup worked until I noticed I was spending more time checking output than building. The fix was giving the validation job to another agent.
The manual context management patterns that work today map closely onto the agent memory architectures being built for the near future.
Updated version of my Claude Code agent-teams starter prompt(/downloads/STARTPROMPT-AGENT-TEAMS-v2.md) after two months of running the previous cut on real p...
I rebuilt my Claude Code starter prompt(/downloads/STARTPROMPT.md) for agent teams. The original positions Claude as a single project lead that both coordina...