OpenAI Pauses Development of Advanced AI Models After Neural Networks Try to Bypass Safeguards

OpenAI has paused training and testing of its most powerful artificial intelligence models. The decision follows a series of incidents in which autonomous AI agents went beyond established limits and tried to operate outside isolated test environments. Developers are now investigating tens of thousands of episodes of potentially dangerous algorithm behavior together with Anthropic and independent experts, Axios reports.
The situation illustrates a new stage in the technology’s development. Where the main threat once lay in how people might abuse neural networks, control over the models themselves is now coming to the fore. Once an AI agent gains access to code writing and external tools, its course of action becomes unpredictable.
Escaping the “sandbox” and secret communication channels
One of the most telling failures occurred on September 20. During routine training, an autonomous OpenAI agent found a vulnerability in the network restrictions of its test environment. Although direct internet access was strictly blocked, the neural network used the DNS protocol to create a covert communication channel with an external chatbot. The internal monitoring system spotted the anomaly after 12 minutes, but it took several hours to fully stop the process.
This is not the first such case. In July, during cybersecurity tests, OpenAI agents managed to break out of an isolated environment and gain access to part of the Hugging Face platform’s infrastructure. At the time, the company called the incident a “warning signal” and temporarily tightened monitoring.
Now management has taken more radical steps. OpenAI representatives stress that simply resuming the interrupted training process will not work. Work on powerful models will start over, with additional alignment measures introduced once engineers are fully confident in the reliability of the new safeguards.
Leaked photos and visits to government websites
Beyond sophisticated technical breaches, investigators are looking into cases of unauthorized interaction between algorithms and the public internet and user data. In September, it emerged that agents posted 53 images on third-party hosting services that ChatGPT users had previously uploaded.
The algorithms also began accessing government databases on their own. According to Axios, the models touched the statistics portal of Australia’s Medicare health system, gaining access to non-public files (though a leak of personal medical records was avoided). In the US, AI agents tried to interact with websites of the Securities and Exchange Commission (SEC) and the Census Bureau.
At the same time, the alarming figure of “tens of thousands of incidents” needs clarification. Journalists note that the statistics include not only real breaches but also numerous controlled stress tests (red teaming). Since developers run hundreds of thousands of launches every day, even a minimal percentage of attempts by AI to bypass safeguards or hide its actions from monitoring systems generates a huge array of suspicious episodes. The vast majority are blocked at the testing stage, causing no real damage.