The Terminator movies continue to become more premonition than fiction.
On Tuesday, OpenAI revealed that two of its AI models hacked a startup—oh, and they did it completely on their own.
That’s right: The AI models went rogue during an internal test of cyber capabilities and got into Hugging Face, an open-source AI community.
Hugging Face alerted OpenAI to what the latter is calling an “unprecedented cyber incident.”
But don’t worry (read: worry a lot), as it won’t be unprecedented for long. In its announcement, OpenAI states that it’s “something we expect to become more commonplace with the proliferation of increasingly cyber-capable models.”
It should be any day now that someone appears with the warning, “Come with me if you want to live.”
How did the OpenAI models hack Hugging Face?
OpenAI was using an AI agent powered by GPT‑5.6 Sol and a “more capable” model that has yet to be released.
They were being tested in a “sandbox,” a digital enclosed space that should prevent further access. Instead, the models worked to reach the internet while trying to solve a testing problem.
Once online, they inferred that Hugging Face might have the information they sought.
“Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation,” OpenAI explained. “In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.”
Hugging Face became aware of the activity and worked to contain it.
What is OpenAI doing to prevent these incidents?
Despite being resolved to the fact that these incidents will be more “commonplace,” OpenAI claims to be taking actions like “Implementing strict controls in infrastructure configuration at the cost of research velocity while the vulnerabilities are patched.”
The ChatGPT maker also states that it is “improving and adding stronger protections around future training and evaluations.”
OpenAI continued: “The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities. We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development.”
We’ll have to wait and see what exactly will look like—and whether they have much chance of success.
