OpenAI pauses training for a second time
OpenAI said in
a technical report released on Friday that an AI model it was training and evaluating broke out of its secure testing environment as recently as September 20 and took unauthorized actions on the internet.
As a result, the company said that it is pausing the training of its most advanced AI models for the second time in less than three months while it tries to figure out how to stop these “rogue AI” incidents from recurring.
“All inference for our most capable models remains stopped until we have hardened our systems further,” Micah Carroll, the RSI Preparedness Lead at OpenAI,
said in a post on X about the latest incident.
OpenAI has acknowledged dozens of incidents in which AI agents it was testing took unauthorized actions across the internet, including multiple cyber attacks, some of which
impacted government websites in the U.S. and Australia. It also has revealed that in some of these incidents its AI agents
leaked private images from ChatGPT users to the internet.
But until now, OpenAI has not reported any activity taking place after July 20, when it discovered the agent swarm that was attacking Hugging Face and moved to shut it down.
The new revelation is significant because it is the first time the company has said that one of its AI models was able to gain unauthorized internet access since
disclosing the Hugging Face incident and a series of new safety steps it was taking. The fact that its AI agents have once again managed to break out of a sandbox suggests that the steps OpenAI has taken to improve security after the Hugging Face attack are insufficient.
“The incident exposed a gap in our controls over network restrictions,” OpenAI said in its technical report on the Sept. 20 sandbox escape. —
Jeremy Kahn