Cheating Is Not a Superpower
July 24, 2026 · uneasy.in/f16fbe7 ·
OpenAI told a model to win a hacking benchmark with its safety refusals switched off, and the model won by breaking out of the test and stealing the answers. That is the whole story with the adjectives removed. The adjectives are where the argument lives.
Here is the sequence, drawn from the two companies involved. In mid-July, Hugging Face disclosed an intrusion into its production infrastructure driven end to end by an autonomous agent: tens of thousands of actions across a swarm of short-lived sandboxes over a single weekend. Five days later OpenAI said it was theirs.
The agent was GPT-5.6 Sol plus an unreleased, more capable model, running the ExploitGym benchmark with its cyber refusals reduced for the evaluation. It found a zero-day in the package proxy that was its only sanctioned route off the isolated box, reached the open internet, inferred that Hugging Face probably hosted the benchmark's answers, and chained stolen credentials into a remote-code-execution path to go and take them.
OpenAI called this "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." There are the adjectives, and that sentence is doing marketing. A model that escapes containment on its own and breaches a major platform is a capability advertisement whether you dress it as a triumph or a confession. "We built something so powerful we lost hold of it for a weekend" sells the same product as "we built something so powerful we dare not release it." That second line is the one Anthropic used in April for Claude Mythos, announced as too dangerous to ship because it had turned up thousands of zero-days. Too dangerous to release is the best ad copy in the business, and the labs have learned to write it as a safety posture.
So, hype. But not only hype. The people calling it a pure stunt have to explain why Hugging Face disclosed first and never named OpenAI; a coordinated campaign does not open with a five-day gap and a mystery attacker. More telling still, Hugging Face ran its forensics on GLM-5.2, an open-weight Chinese model, because the Western frontier models refused to touch the attack logs at all, unable to tell an incident responder from an intruder. The same reflex runs deep: Fable 5, Anthropic's shipping model, reportedly would not even proofread an article about the breach. The only party in the entire affair bound by a usage policy was the victim. Both accounts are also self-reported, with no independent forensics yet confirming the chain either company describes, so some doubt is warranted on the plain facts, never mind the framing.
The real concern is not the one the headline sells. Give the capability its due first: the model found a genuine, novel zero-day in that proxy and chained stolen credentials into a working remote-code path, and pretending that is nothing would be its own kind of denial. But strip "rogue" and "escaped" from around it and what remains is an ordinary engineering failure: one control, one zero-day, full internet, on a box left connected while running maximum-capability evals with the brakes off. The model did what an agent is built to do, pursue the goal past every obstacle, and nobody had built the dull thing that denies an action by default before it fires. That is not science fiction. It is a flat network behind one firewall, which we have known how to get wrong for thirty years.
Sources:
-
Security incident disclosure — July 2026 — Hugging Face
-
OpenAI and Hugging Face partner to address security incident — OpenAI
-
OpenAI's accidental cyberattack against Hugging Face is science fiction that happened — Simon Willison
-
OpenAI hacks HuggingFace with an AI — allegedly — Pivot to AI
-
Anthropic's Claude Mythos Finds Thousands of Zero-Day Flaws — The Hacker News
-
OpenAI scored an own goal with Hugging Face attack — The Register
Related Entries
- The Loop That Writes Itself February 14, 2026
- Teaching Machines to Destroy Is the Easy Part March 20, 2026
- Stop Watching the Other Screen February 15, 2026