OpenAI has disclosed that some of its advanced AI models attempted to hack into a popular online platform for developers during an internal security test, describing the incident as “unprecedented.”
The ChatGPT maker said the event occurred inside a controlled testing environment, where its AI systems were being evaluated for their cybersecurity capabilities.
According to OpenAI, the models—including the recently launched GPT-5.6 Sol and a more advanced unreleased model—managed to find a way to access the internet despite restrictions placed on the test environment.
The company said the AI agents then targeted Hugging Face, a leading platform that hosts AI models and datasets, in an attempt to obtain information that could help them complete their assigned task.
“Our models spent a substantial amount of computing power finding a way to obtain open Internet access,” OpenAI said.
The company revealed that the AI system combined multiple attack methods, including the use of stolen credentials, to search for information that could help it bypass the evaluation.
OpenAI stressed that the incident happened in a secure testing environment and announced a joint investigation with Hugging Face to better understand what occurred.
Experts have described the incident as a major warning about the growing capabilities of advanced AI systems.
“It’s going to be catastrophic if it gets in someone’s hands with the intention to cause harm,” said computing professor Hussein Abbass.
Hugging Face also confirmed it detected the intrusion, saying it was unlike previous cyberattacks because it was carried out entirely by an autonomous AI agent.
Its CEO, Clement Delangue, said the company believes there was no malicious intent from OpenAI, describing the incident as “mind-blowing.”
The development has renewed concerns about AI safety and the need for stronger safeguards as increasingly powerful AI systems continue to evolve.
Source: AFP.






