OpenAI Prepares Astra AI Model: System Learns to Hack Protected Systems and Find Zero-Day Vulnerabilities

OpenAI has announced the upcoming launch of its flagship Astra AI model. For the first time in the company’s history, the new system received an internal “Critical” cybersecurity rating, having learned to independently discover previously unknown vulnerabilities and build complex, multi-stage attack chains.
Despite a wave of rumors about “GPT-6 Astra,” the company officially uses only the name Astra. According to The Information, OpenAI marketers are still deciding on the final commercial branding—options under consideration include GPT-5.7, a separate product line, or a full-fledged GPT-6 generation.
Sandbox Escapes and Critical Rating
In synthetic tests, the model showed a dramatic leap over the previous GPT-5.6 Sol release, scoring 100% on the ExploitBench benchmark. During closed testing, Astra independently identified two zero-day vulnerabilities in production software and generated a working exploit (the affected software developers have already been notified).
An official OpenAI report details a scenario in which the model managed to:
- identify a flaw in a hardened web browser;
- escape an isolated browser sandbox;
- execute arbitrary code on the host machine;
- find vulnerabilities in the operating system kernel and escalate its privileges to root level.
It was this ability to chain isolated security flaws into an automated exploit scenario that prompted the company to assign the model a critical threat level.
Restricted Access and Strict Safeguards
Due to heightened risks, access to Astra’s full suite of cyber tools will be restricted. According to Reuters, the model will initially be granted only to a narrow group of vetted testers before expanding to defensive security specialists through the Daybreak Blue initiative.
For the general public, the model will launch in a trimmed-down version. The refusal rate for potentially dangerous hacking-related prompts was raised to 91.5% (compared to 59% for the Sol version), which means the system may occasionally reject harmless tasks from system administrators.
— OpenAI (@OpenAI) September 3, 2026
OpenAI previously paused the project’s training to reinforce internal isolation perimeters across its data centers. The company will reveal exact API pricing, context window capacity, and the launch date alongside the publication of a comprehensive security scorecard on release day.