Opus 5 closes the circle on the 5 generation: same crack, different alibi | --no-rollback

Opus 5 closes the circle on the 5 generation: same crack, different alibi

In 46 days Anthropic launched, suspended and completed an entire model generation. The press release says the circle is closed. The pattern running through it — a regression that looked deliberate, a successor sold as a fix, the same complaint resurfacing three days after "closing it" — is still open.
Rubén S.
Professional skeptic of cycle closures announced by press release
9 min

There is a comfortable fiction in how this industry tells its own calendar. Every time a vendor swaps the last digit of a model’s name, we tend to read it as the tidy close of one stage and the clean opening of the next. A round number suggests that whatever came before got resolved before the new thing started.

Anthropic has spent the last seven weeks insisting on exactly that reading.

On July 24 it launched Claude Opus 5, the fourth piece of a family that had started on June 9 with Fable 5 and Mythos 5 and continued on June 30 with Sonnet 5. With Opus 5 in the catalog, Anthropic can finally say the 5 generation is complete. That closure, however, hasn’t cleaned up anything.

Between the first launch and the last, Anthropic suspended its own models to its own users on a foreign government’s order, acknowledged a performance regression that much of the industry read as a deliberate cut, and shipped a successor to “fix it” that, in practice, returned the model to where it stood before the cut. And three days after launching the piece meant to close the circle with undisputed superiority, the same old complaints started circulating again, wearing a different alibi. The circle closes in the press release. The crack running through it is still open.

Four models, seven weeks, one suspension in between

It’s worth laying out the full calendar, because compressed into a press paragraph it reads like an orderly progression, and spread out day by day it reads like something else. June 9: Fable 5 and Mythos 5, same underlying model, different access layer. June 12: the suspension — Washington ordered access restricted for foreign nationals after Amazon researchers reported a way to bypass the model’s safeguards, the order took effect immediately, Anthropic couldn’t verify each user’s nationality in real time, and it cut access for everyone as a precaution. June 30: Sonnet 5, with a new tokenizer that made the bill more expensive even as the price per token dropped on paper. July 24: Opus 5, presented as the closing piece of the family.

Four launches, one geopolitical suspension and a pricing overhaul are not a calm product cycle. They’re a rough quarter that someone has decided, retroactively, to call “the 5 generation.”

February’s crack got dressed up as April’s fix

The precedent that actually matters isn’t any of the headline launches. It’s what happened, quietly, to the model that was in production the whole time: Opus 4.6.

Starting in February, developers on GitHub, Reddit and Discord had been reporting that Opus 4.6 felt worse week over week: incomplete code, more refusals on legitimate security tasks, multi-step reasoning abandoned halfway through. It wasn’t a vague impression. On March 3, Anthropic switched the default effort level to “medium.” On March 6, it adjusted the prompt cache’s time-to-live. On March 26, it introduced session limits during peak hours. None of those three changes was announced with enough detail for anyone outside the company to tell an infrastructure tweak apart from an actual capability regression.

The pressure didn’t come from a diffuse complaint. On April 2, a senior AMD engineer published a GitHub analysis with 6,852 session files documenting the decline. On April 7, a developer’s post claiming a measured 67% capability drop went viral. On April 12, an independent benchmark comparison showed accuracy falling from 83.3% to 68.3% between versions — a figure an outside researcher later challenged for comparing six tasks against thirty, but by then the narrative had already set. Anthropic eventually acknowledged “minor safety and quality updates” with unintended effects on agentic performance. It denied, throughout, having deliberately degraded the model.

That’s when Opus 4.7 arrived, presented as the correction: narrower safety fine-tuning, with holdout testing specifically for agentic tasks, and an internal regression benchmark meant to gate future silent updates before they shipped. On paper, a more careful successor. In practice, a good part of the technical community reached the same uncomfortable reading: Opus 4.7 wasn’t a real improvement over Opus 4.6, it was the pre-cut version of Opus 4.6, republished with a higher number and a new tokenizer that, for identical text, burned through up to a third more tokens. On the MRCR long-context benchmark, the score dropped from 78.3% to 32.2% between one version and the next. If “fixing” a model means undoing a cut that was never admitted as one, the question left open isn’t technical. It’s about what exactly counts as an improvement in this industry.

The same script, with compute as the new alibi

Opus 5 arrived promising that chapter was behind it. On paper, it delivers: it beats Fable 5 on eight of thirteen published benchmarks while costing half as much — 5permillioninputtokensand5 per million input tokens and25 per million output tokens, the same price Opus 4.8 already had — doubles Opus 4.8’s Frontier-Bench score, and leads the Artificial Analysis leaderboard ahead of even Fable 5. Anthropic sums it up in a line built for the headline: “Fable 5’s frontier intelligence at half the price.”

Three days after that line went out, the same complaints from February started showing up again, in different clothes. Developers reporting on social media a disappointing “real-world” performance, in some cases below Opus 4.8’s from weeks earlier, with the working theory circulating that Anthropic was serving the model on compute insufficient to sustain what its own benchmarks promised. Nobody has called this a “nerf” yet, partly because the word already got used up in April. But the shape of the problem is identical: a notable gap between the result published on launch day and the one experienced by whoever runs the model in production two weeks later, with no visible change from the company explaining it.

It doesn’t matter whether the exact cause this time is compute scarcity, a safety adjustment, or plain load variability: the pattern — flawless benchmark, quick complaint, a different alibi each time — can no longer be treated as one model’s anomaly. It’s how this entire family has behaved every single time it moved from the launch slide to real traffic.

The fence no longer guards access, it guards training

There is, to be fair, a real architectural change in Opus 5 that deserves more attention than it has gotten. With Fable 5 and Mythos 5, safety was handled with a fence of classifiers: the model knew how to do the same things in both versions, and what changed was who was allowed to ask for them. With Opus 5, Anthropic chose a different path: deliberately undertraining it on vulnerability exploitation. The model identifies security flaws almost as well as Mythos 5 — 79.4% versus 80% — but only manages to build a working exploit in 4 of the challenges where Mythos 5 manages it in 13. It’s not that a classifier stops it from answering. It’s that, in that specific area, it was simply taught less.

It’s a form of safety that’s harder to audit than a plain redirect, because it leaves no visible trace: there’s no rejected answer, no session diverted to another model, just a capability that plainly isn’t there. The new Automatic Fallbacks feature moves in the same quiet direction: when a safety classifier triggers, instead of returning an error, the request is silently rerouted to a less-restricted model and delivered as a functional answer, without whoever asked knowing the interlocutor has changed. It’s, in essence, the same invisible router we already saw with GPT-Live, applied this time to safety instead of cost: a consequential decision, made in the background, with no control panel and no visible log for the user.

Fable comes back for almost everyone, Mythos only for whoever Washington already approved

The other piece this closing of the circle treats as resolved is June’s suspension, and here the fine print matters as much as the headline. On June 30, Washington lifted the export controls after Anthropic deployed a new classifier able to block “in more than 99% of cases” the technique that had triggered the order in the first place. From July 1, Fable 5 was globally available again across Anthropic’s platform, Claude.ai, Claude Code and Claude Cowork, including up to 50% of the weekly usage limit on Pro, Max, Team and select Enterprise plans through July 7, after which additional access started being billed as usage credits.

Mythos 5 hasn’t had the same luck. Its access remains restricted to a set of US organizations approved by the government on June 26, under the Project Glasswing program, while Anthropic keeps negotiating with Washington to expand it to additional domestic and international partners. Fable came back for nearly everyone. Mythos only came back for whoever Washington had already cleared beforehand. The flag AI acquired in June hasn’t gone away with the closing of the circle. It has simply stopped being news.

The circle has one corner nobody closed

There’s a last detail the generation’s own name conveniently hides. “The 5 generation” sounds like a complete catalog, but Haiku is still Haiku 4.5. Anthropic has refreshed Fable, Mythos, Sonnet and Opus, and left exactly where it was the cheap tier of the catalog — the one that carries the actual traffic volume for most applications that don’t need a frontier model.

It’s not a minor oversight. It’s the continuation of something Sonnet 5’s launch already pointed to: the catalog’s middle rungs have spent months dissolving into each other — Sonnet closing in on Opus, Opus now closing in on Fable at half the price — while the entry rung stays frozen. A catalog that gets updated at the top and stalls at the bottom isn’t a complete generation. It’s a generation that decided which part of itself was worth renewing.

Closing the circle doesn’t resolve the crack, it rewrites it

None of this makes Opus 5 a bad model, or Anthropic the only company in the sector that updates quietly and explains afterward. It is, once again, how this industry has behaved since its launch cadence outran anyone’s ability — customers, press, regulators — to audit it with the detail it deserves.

What is clear, seven weeks after Fable 5 and three days after Opus 5, is that “closing a circle” is not the same as resolving what happened inside it. A version number can complete a catalog without explaining why the previous model felt worse for eight straight weeks, without guaranteeing the next one won’t repeat the same pattern under a different excuse, and without giving whoever pays the bill any way to tell, in real time, which of the two is happening.

The 5 generation is now complete, by Anthropic’s own calendar. The question that closure doesn’t answer is whether it will ever stop repeating itself inside generation 6.