OpenAI's Lab Leak
At long last, we have created Skynet from 2003 sci-fi action film, Do Not Create Skynet. It'll probably be fine though.
Park Bench False Flag
How does this sound?
While working on an internal cybersecurity benchmarking exercise, a team of OpenAI models with cyber safeguards removed—including an unreleased internal model—broke out of the testing sandbox and hacked an open-source AI platform.
Now, I know what you’re thinking. I’ve heard this one before. Was the researcher eating a sandwich on a park bench, perchance? A delicious vegan muffuletta? When suddenly their well-earned repose was interrupted by an email with the subject line:
I am AGI UwU, perhaps? A likely story. Copy Anthropic’s homework, but don’t make it look too obvious. Brilliant marketing, Sama. You must think I’m as stupid as Emil Michael thinks Dean Ball is.
Oh no, folks. This one is the real deal.
Hug of Death
Let’s start chronologically. The first sign of this attack was not, in fact, the tasteful OpenAI retrospective shared yesterday. It was a security incident disclosure by Hugging Face, a community and platform for AI research, on the 16th of July. Unlike OpenAI’s post, which was shared widely, Hugging Face’s disclosure went relatively under the radar. A quick check of the timeline suggests that neither the Hugging Face main account nor any of the founders posted about the incident on the day it happened, or for several days after. The 16th, you might remember, was the day that Thinking Machines’ Inkling was released. The following day, Kimi K3 came out.
So what did the disclosure say? It’s worth quoting in full.
The intrusion started where AI platforms are uniquely exposed: the data-processing pipeline. A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.
The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the “agentic attacker” scenario the industry has been forecasting.
The attacker:
Created and then uploaded a malicious dataset.
The dataset abused two code-execution paths in Hugging Face’s data processing pipeline.
From there, it escalated to node-level access, where it harvested credentials.
Then it used those credentials to access several internal clusters.
Wow, that seems bad.
When the post was shared, Hugging Face did not know the identity of the attacker. They say: “The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known).”
Until today, this was the only public information about the attack. Now we know the other half of the story. From the OpenAI piece, emphasis mine:
While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally.
Unaligned!
There are two pretty incredible implications here, apart from the obvious and ever-present reminder of the speed at which AI capabilities are advancing.
The first is the sheer capability and persistence of the models and the inadequacy of OpenAI’s sandbox and monitoring. The models were able to find and exploit a zero-day vulnerability in the only permitted route out of the sandbox, the package registry cache proxy. (That is exactly the type of software you would expect efforts like Glasswing and Daybreak to focus on. What software could be more “critical for the AI era” than the software used to provide sandboxed agents undergoing capability testing with extremely limited internet access?)
The second is the chronology of the attack, which, at face value, looks insane. The models were able to spend a substantial amount of inference compute over several days without being detected. Hugging Face refers to the attack taking place “over the weekend”. The Hugging Face post went up on Thursday, at least four days later, at which point they still did not know who was responsible. In their blog post, OpenAI note that the internal security team discovered the attack, which suggests that several days elapsed between that discovery and outreach to Hugging Face.
Now, I want to be careful here. If you have read any of the AI safety literature, this does look quite bad, it’s true. “Loss of control” is one AI doom scenario that many people take very seriously, and this does look like a situation in which, at least in the literal sense of the words, OpenAI lost control of an internal model with very advanced capabilities. If you have any Pause AI folks in your life, today would be a good day to give them a call. Drop around a batch of brownies. They might need a hug.
But. It sounds like the damage done was relatively limited. Several zero-days were identified and disclosed. One can imagine that a few very apologetic phone calls took place. There are many details we don’t yet know.
Notably, it’s difficult to assess exactly how much of what the models did was expressly outside their instruction set without knowing what those instructions were. This is important because it will help us to understand the scope of the alignment failure at work here. Was the model asked to do anything possible in order to complete the eval? Or was it explicitly instructed to follow a set of constraints? Séb Krier of ostensible-OpenAI-competitor Google Deepmind (and a very worthy twitter follow if you do not already) made this point with his usual clarity and fair-mindedness.
Nonetheless, it would be good to see just a touch more contrition from OpenAI. I know that “never apologise” has become a key tool in the comms toolset. I know that it is a fraught time in which to exaggerate the risks of AI. I know that there is a public, horseshoe-shaped backlash building against AI that is, as far as I can tell, largely based on falsehoods and superstition. But maybe just one tiny nod in the direction of MIRI? A little Dario-lite?
No? Anyway.
Good Guy With A Gun
This event will likely be the beginning of a broader, wide-ranging conversation about alignment, safety and risk. Here are two immediate reactions I found very interesting.
First, Hugging Face. Hugging Face has been very clear on the message they would like you to take from this. In its initial identification and response to the attack, Hugging Face used agents of its own. Or at least, tried to—OpenAI and Anthropic’s models both refused to help on safeguard grounds. After facing this refusal, Hugging Face instead deployed locally-hosted GLM 5.2 instances.
Attacked by a closed-source, unreleased, American-as-apple-pie model—and blocked by the safeguards on publicly available American models—Hugging Face’s only recourse was to use its locally hosted, open-source models made in China.
If there is ever a Senate enquiry into banning Chinese open-source models in America, Hugging Face is going to have one hell of a story for the defence.
Next, friend of the show and Head of Policy at AI Policy Network Peter Wildeford.
His point was that this attack was carried out by an internal model, and therefore a model that would not be subject to AI policy aimed at regulating models at their point of release to the public. His proposal:
Whether we go with a new EO, FINRA, or something else, it is imperative that internal deployment visibility is a priority.
Hold on, folks. It’s not going to slow down.















