Claude Fable 5: what a new tier above Opus means for context management and agent teams
Anthropic launched Claude Fable 5 yesterday, its first Mythos-class model, priced at double Opus. The context window did not grow. What changed is the management layer around it.
Anthropic launched Claude Fable 5 yesterday. It is the first model in a new tier above Opus, the first "Mythos-class" model made generally available. I have spent the last day reading everything published about it and testing it through the API and Claude Code.
I have been writing about context management since February: how context loss quietly breaks sessions, how the major platforms manage agent memory, and why agent teams need explicit context boundaries the same way human teams do. The closing line of the context wars post was that context is a resource that needs to be managed, not a bucket that needs to be bigger.
Fable 5 is the strongest validation of that argument I have seen from any platform.
Launch details
The launch had several moving parts.
Fable 5 is a new tier positioned above the Opus family. Pricing is $10 per million input tokens and $50 per million output, exactly double Opus 4.8. It is available immediately via the API (claude-fable-5) and Amazon Bedrock, and it is included free for Claude Pro, Max, Team, and Enterprise subscribers from June 9 through June 22.
The benchmark numbers are strongest where they matter most for agentic work: 80.3% on SWE-Bench Pro against Opus 4.8's 69.2%, 88% on Terminal-Bench 2.1, and 64.5% on Humanity's Last Exam with tools. The headline capability claims are about long-horizon autonomy (working unattended for longer stretches than any previous Claude model) and about staying focused across very long contexts.
Fable 5 is the same underlying model as Claude Mythos 5, which Anthropic is releasing only to a small group of cyberdefenders and infrastructure providers through a US government collaboration called Project Glasswing. Fable 5 is the public version, with safeguards that reroute sensitive queries to Opus 4.8 instead, triggered in under 5% of sessions, per Anthropic.
That safeguard is an orchestration pattern. Anthropic is running model-level routing in production: one model handing specific categories of work to another based on policy. That is the same separation-of-concerns logic I have been describing for agent teams, applied by the platform to itself.
Context window unchanged
Fable 5 ships with a 1M token context window and up to 128K output tokens. Opus 4.8 has a 1M window and 128K output. The new tier costs double and is the company's biggest capability jump in a year, and the context window did not change.
Two years ago that would have been the headline disappointment. Today it barely gets mentioned, because the window arms race is over. Anthropic spent its effort on the management layer instead.
Token efficiency. Anthropic reports Fable 5 is more token-efficient than Opus 4.8 despite being substantially more capable. Bigger models historically thought longer and produced more text, so this reverses the usual pattern. A model that does more with fewer tokens is managing its context better internally, in addition to whatever the user manages externally.
Task budgets. When I compared platforms in March, the feature that set Claude apart was token budget awareness: the model knowing how much room it has left and pacing its work accordingly. That is now a first-class API feature. You give the model a total token budget for an entire agentic loop, and it sees a running countdown and self-moderates, prioritizing, cutting scope, and wrapping up as the budget drains. This differs from max_tokens, which is a hard ceiling the model never sees. max_tokens is a wall the agent runs into; a task budget is a deadline the agent plans around.
Server-side compaction. For conversations that approach the window, the API summarizes earlier turns automatically and keeps going. This is the platform version of the HANDOVER.md pattern I have been doing by hand since February: when the context fills up, distill it and continue. I would still rather control what survives compaction myself for critical work, but for long-running background agents it removes a whole class of session failure.
File-based memory. Given access to file-based memory, Fable 5 showed a 3× performance improvement on complex long-horizon tasks. That is 3×, not 30%. The model is better at writing notes to itself and using them later. In the EU this matters more than usual: file-based memory remains the GDPR-compliant persistence path, while ChatGPT's full conversation recall is still geographically locked out. The architecture I can use from Sweden just became more capable.
Mid-conversation system messages. You can now inject operator instructions partway through a session as proper system-role messages appended to the conversation, instead of editing the top-level system prompt and invalidating the entire prompt cache. Caching economics shaped my whole platform comparison in March, and this closes one of the more annoying gaps: changing an agent's instructions mid-flight no longer costs you the cached prefix.
Put together, the pattern is clear. The new tier is not "more context." It is a model that is aware of its budget, economical with its tokens, persistent through files, and steerable without invalidating the cache. The gains are in the management layer rather than in capacity.
Interface contracts and separation of concerns
The second thing I find interesting about Fable 5 is how opinionated the API surface has become, and what that does for multi-agent work.
Fable 5 accepts adaptive thinking only. The old fixed thinking budgets are gone. temperature, top_p, and top_k are gone; send them and you get a 400. Even explicitly disabling thinking is now an error; you omit the parameter instead. Assistant-turn prefills, long the standard workaround for output control, are rejected in favor of structured output schemas.
You can read this as Anthropic taking knobs away. I read it as interface contracts getting more honest. Every removed parameter is a place where developers used to encode vague intent ("temperature 0.3 means... be a bit careful?") that the model interpreted unpredictably. What replaces them (effort levels, output schemas, task budgets) are contracts with defined semantics. In the separation-of-concerns post I argued that agent teams depend on three things per agent: a clear mandate, a bounded context, and an interface contract. The platform is now enforcing the third one at the API level.
The multi-agent layer around the model has matured the same way. Anthropic's Managed Agents API, where Fable 5 slots in as a coordinator, runs delegation through context-isolated threads. Each subagent thread gets its own conversation history, its own system prompt, and its own tools. Threads share the filesystem but not conversation context: if a coordinator wants a subagent to know something, it has to say so explicitly in the delegated message or write it to disk.
That constraint is the thesis of my agent teams series, enforced by infrastructure. There is no ambient knowledge and no tribal memory leaking between roles. The handoff is the contract: anything not in the handoff does not exist to the subagent. When I described Context as Code, externalizing shared knowledge into structured artifacts both humans and agents read, I was describing a discipline. The shared-filesystem-but-isolated-context model makes it the only way information moves.
Two more constraints in that design are worth attention, because they are governance decisions rather than plain limits. Delegation goes one level deep: a coordinator's subagents cannot spawn their own subagents. A coordinator's roster caps at 20 agents, with at most 25 threads running concurrently. The one-level rule prevents opaque hierarchies where accountability dissolves between the top and the work, the failure mode of middle-management layers in an org chart. My advice in March was a three-to-five agent squad with a thin orchestration layer above it. The platform's ceiling is more generous than my recommendation, but the shape is the same: flat, bounded, and explicitly rostered.
The subagent threads are also persistent. A coordinator can come back to a subagent it briefed earlier, and that subagent still has its prior turns. They are specialists with memory of their own work, not stateless function calls. A coordinator delegates to something that remembers the earlier exchange.
Pricing tiers and team structure
Fable 5 costs exactly double Opus 4.8, which costs more than Sonnet, which costs more than Haiku. For the first time there are four well-differentiated capability tiers in the same model family, and that pricing ladder is an org-design argument.
I wrote about the cost model of agent teams in February: every agent in a pipeline adds token cost, and the cleanest decomposition is worthless if it is economically unviable. The four-tier ladder gives that trade-off real resolution. The coordinator is the role that needs long-horizon judgment, budget awareness, and the ability to recover from surprises, and that is where Fable 5's premium earns its keep. Discovery sweeps, mechanical transformations, and validation passes against explicit rules are Sonnet and Haiku work, at a fifteenth or a thirtieth of the price.
This mirrors how you staff a human program: you do not put your most senior architect on data entry, and you do not put a junior on the steering committee. "Which model for which role" is now a real staffing decision with an order-of-magnitude cost range, and the answer falls out of the role definitions you should have written anyway.
My evaluation plan
My subscription includes Fable 5 free until June 22, so the evaluation window is open and I intend to use it.
Claude Code sessions on my own projects move to Fable 5 now. The long-horizon claims are the thing to test against real work: a four-hour session on the matbotten codebase, a multi-step infrastructure change across this server's stacks, the kind of task where Opus 4.8 is good but still occasionally drifts around hour three. If the file-based memory improvement is real at 3×, my .md-file workflow (CLAUDE.md, DECISIONS.md, HANDOVER.md) should get more leverage without me changing anything, because the model is better at the consuming end of those artifacts.
The agent-team experiments from the spring get re-run with a two-tier structure: Fable 5 coordinating, Sonnet 4.6 doing the specialist work. My hypothesis is that the coordinator tier is where the cost matters and the specialist tiers are where it does not, which would mean the practical cost of upgrading a whole team is much less than the headline 2× suggests, since the coordinator emits a small fraction of total tokens.
After June 22, I decide what is worth $50 per million output tokens on an ongoing basis. My guess today is the orchestrator role and genuinely hard one-shot problems, nothing else. The pricing ladder is there to be matched to the work.
The main takeaway has not changed since March, but it is sharper now. Fable 5 made the biggest capability jump in a year without adding any context window. The gains came from management: models that know their budget, externalize their memory, respect their role boundaries, and hand off work through explicit contracts. Each of those is a discipline practitioners have been building by hand. The platforms keep absorbing them one by one, and the practitioners who understood the discipline first are the ones who will get the most out of the infrastructure.