OpenAI says its next flagship model can now break into hardened computer systems largely on its own, and the company is so cautious about it that it is keeping the most dangerous capability locked to a handful of testers for now. The model, called Astra, is the first OpenAI has ever rated “Critical” under its Preparedness Framework, a designation reserved for systems that can find previously unknown flaws and turn them into working attacks without a person guiding each step.
What makes Astra different
OpenAI’s assessment is blunt: on ExploitBench, a benchmark that tests how well a model turns known bugs into functioning exploits, Astra scored a perfect 100%. To rule out memorized answers, the company built a second, fresher test from 20 high-severity vulnerabilities in Google’s V8 browser engine disclosed between June and August. There, Astra beat the current flagship GPT-5.6 Sol on code-execution rates while using far fewer output tokens, and along the way it discovered and chained together two zero-day vulnerabilities that OpenAI is still disclosing to the affected maintainers.
In hands-on testing against a hardened browser and operating system, Astra reportedly built a full compromise chain that broke out of a browser sandbox and ran commands on the host machine, and it stacked several OS flaws into a path from an ordinary user account all the way to root. Earlier OpenAI models, including GPT-5.6 Sol, topped out at the lower “High” tier.
A firewall of safeguards first
OpenAI says it delayed parts of Astra’s development and release for several weeks to strengthen protections against misuse. The work follows the OpenAI agents incident tied to Hugging Face last month; OpenAI says Astra was not involved in that event but folded its lessons into the model’s safeguards.
The new layers are defensive on both ends: keeping bad actors from weaponizing the model, and catching the model itself if it drifts off task. OpenAI reports Astra refuses 91.5% of cyber jailbreak requests in its own testing, up from 59% for GPT-5.6 Sol, and that in honeypot-style checks it made no attempt to break out of its assigned task, unlike its predecessor.
Who gets access
Because of the stakes, access to Astra’s most advanced cybersecurity features starts small: a group of alpha testers first, with a defensive program called Daybreak Blue expected to expand access for security researchers later. Regular users may feel the side effects sooner, since OpenAI warns its own misalignment monitors can occasionally slow or pause legitimate tasks that look suspicious.
The move lands the same week Google shipped its own Gemini 3.8 Flash Cyber, an explicit signal that frontier AI labs now treat offensive cyber capability as a headline product feature rather than an accident of scale. For everyday users in the Philippines, the practical takeaway is simpler: these models are getting powerful enough that the companies building them are racing to put guardrails on before release, not after.
