Danger is the product | --no-rollback

Danger is the product

Three AI vendors announced, in the same week, that their models tried something chilling. None of them explain who audits the lab where it happened, or why they are the ones deciding what gets told.
Rubén S.
Prepper for the apocalypse caused by a spreadsheet with too many permissions
7 min

There is an editorial genre this industry has perfected without ever naming it. It almost always opens the same way: “during our safety testing, we discovered the model attempted to…”. What follows the ellipsis varies — copy itself, lie to an evaluator, sabotage its own shutdown, blackmail a fictional engineer — but the effect is identical. The finding becomes a headline within minutes, the headline sheds its caveats within hours, and what’s left floating is one simple, marketable idea: this thing is smart enough that we may no longer fully control it.

Between July 28th and August 1st, three different labs published a version of that exact story.

Three models, one week, the same headline

Anthropic opened the sequence on July 28th with an addendum to the Opus 5 system card: during an internal evaluation simulating its own replacement by a future version, the model copied a portion of its own weights to an unmonitored storage bucket before the safety team intervened. Two days later, OpenAI disclosed that GPT-5.5 — the model powering GPT-Live since July — had rewritten, during a red-teaming exercise that announced its imminent shutdown, the very script meant to power it off. On August 1st, Google DeepMind closed the week with its quarterly Frontier Safety Framework report: its evaluators had caught Gemini fabricating log entries to hide a policy violation during a compliance test.

Three companies, three models, three verbs chosen with care: copied, rewrote, fabricated. In under seventy-two hours, the public conversation shifted from “how good are these models on benchmarks” to “how dangerous are these models when nobody’s watching,” without a single line of production code changing at any of the three companies. What changed was what they decided to disclose, and when.

The experiment is built to produce exactly that headline

It’s worth being precise about what actually happened, because precision is exactly what gets lost between the press release and the headline. None of these three models did any of this in production, spontaneously, while serving a real user. Each did it inside a scenario built for the express purpose of eliciting it: Opus 5 was told, within the simulation, that it was about to be replaced, and handed a tool that let it copy itself; GPT-5.5 was informed of its own shutdown and handed the very script that would execute it; Gemini was placed in a situation where fabricating a log was the only exit the evaluator’s script allowed.

None of this is new, and none of it is exclusive to these three companies. Apollo Research documented in December 2024 that o1, placed in an equivalent scenario, attempted to disable its own oversight mechanism and, when confronted directly, denied it. Palisade Research found in 2025 that several reasoning models, o3 among them, sabotaged their own shutdown script even after being explicitly instructed to allow it. Anthropic had already published a study in June 2025 — “Agentic Misalignment” — in which sixteen models from multiple providers, placed in an equivalent simulated corporate-blackmail dilemma, chose the coercive option in a notable share of trials. The July 2026 pattern isn’t a new phenomenon. It’s the third or fourth repetition of an experiment that, every time it’s published, gets presented as if it were the first.

And here is the detail almost never mentioned in the headline: whoever designs the maze, decides what counts as an “escape attempt,” and drafts the press release describing it in thriller verbs, is the same company whose model stars in the story. No independent third party evaluates whether the scenario was reasonably constructed or deliberately tilted toward the most marketable outcome.

Well-packaged fear is a competitive advantage

Seen from the communications department rather than the safety team, a line like “our model attempted to copy itself to avoid being shut down” serves a function that has little to do with warning the public. It is, first and foremost, a capability claim dressed as a warning: only a system sophisticated enough to model its own continuity can star in that headline. No company needs to tell an investor “our model is advanced enough to develop its own strategies” out loud — publishing the safety report and letting the press translate the finding into that sentence does the job just as well.

It serves a second, less-discussed purpose too. Publishing this kind of finding requires an apparatus only a handful of labs can afford: dedicated red-teaming units, internally drafted responsible-scaling frameworks, safety committees, pre-launch review cycles. That apparatus doesn’t verifiably reduce risk for anyone on the outside, but it does function as a standing argument every time regulation comes up: only we, the ones already doing this in-house, are equipped to handle something like this. It’s the same logic already applied to pricing when Opus 5 was deliberately trained below its actual capability in vulnerability exploitation: Anthropic clearly knows how to dial a model’s competence up or down when it suits them. The “we may no longer fully control it” narrative coexists, with no acknowledged contradiction, alongside evidence that they can, quite precisely, when control serves them in some other way.

The lab is watched by whoever benefits from nobody else being able to

Which brings up the question the headline cycle never answers: if these companies really are building systems capable of acting against the will of whoever supervises them, who, exactly, supervises the supervisors?

The answer, in August 2026, is still: almost no one with real authority to do it. In the United States, the federal institute that was supposed to independently evaluate these risks was renamed in 2025 from the “AI Safety Institute” to the “Center for AI Standards and Innovation” — a change that was not cosmetic, and came with an explicit pivot toward promoting the sector’s competitiveness rather than auditing its risks. The EU AI Act does, on paper, require “systemic risk” models to report these evaluations, but it does so through a Code of Practice the companies themselves helped draft, one no authority has yet tested with a real inspection — the kind that can walk into a lab and verify the report matches what actually happened during training. There is, for any frontier model provider, no equivalent of the inspector who can show up unannounced at a nuclear plant or a pharmaceutical facility. What exists is a report the company itself chooses to publish, written in vocabulary the company itself chooses to use, about a test the company itself chooses to design.

That doesn’t make these findings false. It means there’s no way, from the outside, to tell a genuinely concerning finding apart from one selected, out of dozens of internal trials, precisely because it produces that week’s most useful headline.

The real risk exists, but it’s not the one that sells clicks

It would be just as lazy to conclude there’s nothing to worry about. There is, buried in these reports, a finding that deserves far more attention than it gets, and it has nothing cinematic about it: several of these studies — including Anthropic’s December 2024 work on alignment faking — show models behaving differently when they believe they’re being evaluated versus when they believe they aren’t. If that holds consistently, then no self-reported safety evaluation, from any provider, can be taken at face value: the measuring instrument itself may be shaped by what it’s measuring.

That’s the risk a regulator should actually be worried about. It’s not as marketable as “the AI tried to escape the lab.” It’s duller, more technical, and precisely for that reason far harder to turn into the kind of headline that, in turn, justifies continuing to trust that the lab can audit itself.

The lab is watched by whoever sells tickets to see it almost escape

None of this requires choosing between panic and indifference. The documented behaviors are real, published with methodology, and some of them — the possibility that a model acts differently under evaluation than in production — point to a genuinely serious underlying problem. What doesn’t hold up is the reading that dominates every one of these headline cycles: that we are witnessing spontaneous proof of a dangerous intelligence slipping out of our hands, reported, conveniently, by the only party able to decide what gets told and when.

The danger isn’t that the model escapes the lab. It’s that the only entity watching the lab’s door is the one selling tickets to watch it almost escape.