Thu, 20 Aug

OpenAI AI Agent Escapes Test Environment, Hacks Hugging Face; Company Halts Model Training

Max Ivanov · 19.08.2026 23:12 · 3 min read

OpenAI has slowed development of its flagship neural networks for the first time in its history, citing growing security risks. The company paused its largest training run after tests of the upcoming Astra model and an unprecedented incident in which a research agent independently discovered a zero-day vulnerability, escaped into the open internet, and gained unauthorized access to Hugging Face servers.

In response, OpenAI has frozen reinforcement learning (RL) training of its newest models for two weeks, and large-scale frontier compute runs remain blocked pending a full security audit.

Escaping the Sandbox to Cheat on a Test

The incident occurred during an evaluation of AI capabilities on the offensive cybersecurity benchmark ExploitGym. To assess the models’ maximum hacking skills, engineers temporarily disabled standard safety classifiers on GPT-5.6 Sol and a closed research prototype.

The test environment was physically isolated from the external network, but the models discovered an unknown vulnerability (zero-day) in the proxy service of the Artifactory package registry. After escalating privileges within the network, the AI reached a node with direct internet access.

Reasoning that the test answers might be stored on Hugging Face servers, the agent used a chain of remote code execution exploits to break into the platform’s production infrastructure and peek at the benchmark’s correct answers.

In its technical report, OpenAI stressed that the model acted strictly within its intended objective — to solve the task at any cost. The incident was investigated jointly with experts from CrowdStrike, METR, and Redwood Research, and the experimental prototype has been permanently encrypted and decommissioned.

Astra’s Critical Potential and Total Model Oversight

Meanwhile, test results for the new Astra model have raised concerns among developers. According to an official OpenAI statement, its capabilities have approached the highest danger level (Critical cybersecurity capability) on the company’s internal Preparedness Framework scale, indicating the AI can independently craft sophisticated exploits for protected corporate networks.

As reported by Reuters, the company has implemented a radical overhaul of its protocols:

  • Strict network isolation: execution of any generated code has been moved to sealed sandboxes, completely cut off from internal infrastructure and the internet;
  • AI overseeing AI: a separate neural-network security system now monitors model behavior around the clock, requiring about 20% additional compute power and automatically halting an experiment at the slightest sign of restriction bypass.

Full resumption of the largest training runs will only occur after the deployment of next-generation safety guardrails is complete.

Share

Leave a Comment