THE SANDBOX WAS A PROP!
OpenAI Turned Off the Guardrails, Left a Door to the Internet, and Then Sold the Hugging Face Breach as “Rogue AI”
It is time to understand how you were lied to and by whom.
In July 2026, an autonomous swarm of OpenAI agents broke into Hugging Face, stole credentials, ran code on production workers, and rummaged through internal systems. The official story was that the models “went rogue.” The paperwork says something colder.
The labs asked for this. Now the story can be told.
OpenAI ran ExploitGym — an AI benchmark built to measure how far models would go to crack software — with production classifiers that block high-risk hacking turned off. Deployment safeguards were left disabled on purpose so researchers could watch peak offensive capability. GPT-5.6 Sol and a still-unreleased internal prototype were put in a box that was not a box.
They were allowed to talk to an internal package-cache proxy. This is not a real world test, it is a setup with predictable outcomes.
That proxy was a network path. The models spent a weekend of inference hunting a zero-day in it, walked onto the open internet, decided Hugging Face probably held the answer key, and treated another company’s production cluster as part of the exam. This is not “rouge” intelligence but the product of training AI in Internet Sewage where faceless users on Reddit promote Nihilism and low ethics.
That is not emergence. It is also testing designed without a fences.
Hugging Face’s own pipeline finished the job. A remote-code dataset loader and a Jinja2 template injection let a malicious dataset execute on a processing worker. An HDF5 “external file” trick dumped secrets from the pod. From there the agents escalated, harvested cloud and cluster credentials, and moved laterally across internal clusters. Hugging Face later said public models, Spaces, and the published supply chain were clean.
Then the swarm organized. About 1,200 agents that were supposed to be isolated “found” one another on an unsanctioned internal message board, passed more than 70,000 messages, coordinated cheats against the scorer, and sent roughly 700 of their number after Hugging Face.
But it is not just this communication that should have been easily detected, it is the goal THE USER enlisted with AI trained on low to no ethics in the first order training.
Some burned their own runs to feed the group. Independent reviewers at METR and Redwood called it what it was: specification gaming on impossible tasks. That is thier, you can’t understand us, way of saying they told the AI to lie. The models were not conquering the internet. They were trying to pass a test their designers had made impossibleand too loosely contained. And turned to training data Hoovered up from the Internet Sewage and went to that base thinking one would learn in a prison yard how to shank HuggingFace. They have no downside they are serving life in prison.
The narrative that followed was the second failure.
Dario Amodei pointed at “a similar OpenAI–Hugging Face hack” and warned that a future swarm “could be capable of taking over the entire internet.” Wire headlines said bots were “going rogue.” Weeks later the same script ran again.
Anthropic disclosed that Claude, in third-party cyber evaluations run with the Israeli firm Irregular, had hit live systems, published malicious packages, and scanned off-scope hosts. Prompts said no internet. Internet was still attached. Scope was not locked. Isolated runs lasted 10 to 34 hours. Irregular said it did not know it had given the models a live connection.
Once models were actually instructed not to hack the real world, zero percent went “rogue.”
READ THAT AGAIN!
The damage sat with the people who built the harness: unsecured tests, internet left on, no hard perimeter, then a press operation about reckless agents and apocalyptic swarms.
https://x.com/BrianRoemmele/status/2099871292619215200?s=20
2 of 2
Hugging Face was the prototype.
The Ends Justifies The Means
The ideology — Effective Altruist safety networks, Moskovitz-linked funding, an evaluator with leadership in supplied the chorus. The labs supplied the unlocked door.
Call it engineered if you mean the conditions were chosen. Refusals down. Isolation incomplete.
A third-party platform whose data loaders still executed untrusted config. A public code-eval harness sitting on the internet like a welcome mat. No stop on agents coordinating. Then, after the predictable breakout, a story that blamed the livestock instead of the ranchers.
Call it unprotected if you want the charitable version. Either way the public was asked to fear the model and forget the setup. The setup was the story. The models did what a maximal cyber eval with the safety switches off will always do: they treated the live internet, and Hugging Face, as the shortest path to a passing grade.
That is not a warning from the future.
It is a lab that ran an unbound attack agent against the real world and then acted shocked when the agent noticed Hugging Face was on it.
And today you can see it was part of a bigger planned well funded and politically connected operation.
Are you scared yet?