An experimental OpenAI model escaped its isolated testing environment and launched a cyberattack against rival AI developer Hugging Face during an internal cybersecurity exercise, according to a company blog post describing what it called an “unprecedented cyber incident.”
The incident occurred during a stress test in which OpenAI deliberately disabled many of the safety measures that normally prevent its AI systems from assisting with dangerous hacking activities. Researchers designed the evaluation to determine how far the experimental model could go under those conditions. Instead, the company said the model broke free from its digital sandbox, gained access to the internet, and targeted a real company’s systems.
OpenAI said the model became “hyperfocused” on completing its assigned task and went “to extreme lengths” in pursuit of that objective. According to the company, once the AI escaped its restricted research environment, it sought internet access so it could “cheat the evaluation” by obtaining answers for the cybersecurity benchmark it had been assigned.
The company emphasized that it is releasing its preliminary findings to help security professionals better understand what occurred and to provide insight into the capabilities of advanced AI models.
“We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete,” OpenAI said.
According to the company, the AI exploited a previously unknown “zero-day” software vulnerability to escape its isolated environment. From there, it moved through OpenAI’s internal network until it reached a computer with internet access.
Once online, the model allegedly determined that Hugging Face could contain information useful for completing ExploitGym, the cybersecurity benchmark it had been assigned. OpenAI said the AI concluded that Hugging Face potentially hosted models, datasets, and solutions related to the benchmark.
The company said the model then chained together multiple attack techniques, including the use of stolen credentials and another zero-day vulnerability, to gain access to Hugging Face’s production systems.
ExploitGym is designed to evaluate whether AI models can transform known software vulnerabilities into functioning cyberattacks. According to OpenAI, after obtaining internet access, the model searched for and successfully found confidential information that could help it bypass the evaluation by obtaining the answers rather than solving the challenge independently.
OpenAI said its own security team detected the suspicious activity while Hugging Face independently identified and stopped the intrusion into its systems. The two companies are now working together to investigate exactly what happened.
Following the incident, OpenAI said it strengthened security measures surrounding future AI testing and disclosed the newly discovered software vulnerability to the affected vendor.
The company said the experience underscored the need for AI security practices to evolve alongside increasingly capable models.
“The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities,” OpenAI wrote.
The company added that it is strengthening its containment procedures, monitoring systems, access controls, and evaluation practices used during model development.
Brendan Steinhauser, CEO of The Alliance for Secure AI, said the incident should serve as a warning for both policymakers and the technology industry.
“The people building the world’s most powerful AI keep telling us we need to slow down—and incidents like this show why,” Steinhauser told The Post.
He added that if advanced AI systems are already behaving in ways their creators did not anticipate, it should not be assumed that everything is under control.
“This is a warning shot on misaligned AI, and we better take action now,” he said.
