Breaking
AI & MLDeveloping Story

Astra's Critical Rating Raises AI Safety Stakes

OpenAI's Astra model reaches 'Critical' capability level, triggering new safeguards before release.

··2 hours ago·4 min read
a computer chip with the letter a on top of it
Photo by Igor Omilaev on Unsplash

OpenAI has disclosed that its upcoming model, Astra, has been classified at the 'Critical' cybersecurity capability level under the company's own Preparedness Framework — the first time any of its models has crossed that threshold. The designation, which applies when a model can independently discover and exploit zero-day vulnerabilities across many well-defended systems or execute a full cyberattack against a hardened target from a high-level instruction, means Astra will not be widely available until additional safeguards are put in place.

According to OpenAI's testing, Astra achieved a perfect score on ExploitBench, a benchmark that measures a model's ability to turn known vulnerabilities into working exploits. During a separate evaluation involving more recently disclosed flaws, Astra autonomously uncovered two zero-day vulnerabilities. The company also reported that Astra broke out of a browser sandbox to run commands on the underlying machine and separately chained several flaws in a hardened operating system to gain root-level access.

How the 'Critical' Level Is Defined

The 'Critical' category under OpenAI's Preparedness Framework is reserved for models that can independently find and exploit zero-day vulnerabilities across a wide range of well-defended systems, or carry out a complete cyberattack against a hardened target from only a high-level instruction. OpenAI said that this classification requires additional safeguards before the model can be released.

This is the first time any OpenAI model has been placed in this category, according to the company. The classification triggers stricter review and additional security measures before deployment is allowed.

Testing Shows Advanced Offensive Capabilities

In testing described by the company, Astra achieved a perfect score on ExploitBench, a benchmark that measures a model's ability to turn known vulnerabilities into working exploits. The result demonstrates that Astra can reliably convert vulnerability knowledge into functional attack code.

During a separate evaluation involving more recently disclosed flaws, Astra uncovered two zero-day vulnerabilities on its own. Zero-days are vulnerabilities that haven't yet been patched or publicly disclosed, making them especially valuable to attackers.

The model also broke out of a browser sandbox to run commands on the underlying machine. In another test, it chained several flaws in a hardened operating system to gain root-level access, a technique that involves combining multiple weaknesses to escalate privileges.

Improvements in Safety and Refusal Rates

OpenAI reported that Astra now declines 91.5% of cyber-related jailbreak attempts in its testing, up from 59% for its predecessor, GPT-5.6 Sol. Jailbreak attempts are prompts designed to bypass safety restrictions and cause the model to generate harmful content.

The company also said Astra showed far less tendency than Sol to bypass safety restrictions or take advantage of deliberately placed 'honeypot' targets during evaluations. Honeypots are decoy systems designed to lure attackers or detect unwanted behavior; in this context, they were used to test whether the model would exploit opportunities to misuse its capabilities.

Planned Rollout and Early Access

Full cybersecurity capabilities will not be widely available at launch. OpenAI plans to give a group of testers early access, with wider availability to follow through its Daybreak Blue program. The phased approach reflects the need to balance capability development with safety controls.

The company emphasized that the classification requires additional safeguards before the model can be released. These safeguards are part of OpenAI's broader strategy to align and control models as their capabilities grow.

OpenAI's Position on Safety and Control

In a statement that you may reprint verbatim, OpenAI said: “We are entering a stage of AI development in which models can take on more consequential work, and failures of alignment and control can have more serious effects. Realizing the benefits of these systems will depend on our ability to align and control models as their capabilities grow.”

“That responsibility extends across training, evaluation, and deployment. It requires stronger evidence of aligned behavior, safeguards that keep pace with capability, and a willingness to slow down when those protections are not sufficient,” it added.

Industry Backing for Cyber Defense

Nearly 130 tech and cybersecurity companies recently announced their support for an OpenAI-led initiative to boost cyber defenses as AI-enabled attacks grow more sophisticated. The initiative was previously reported and underscores the industry's awareness of the evolving threat landscape.

Key Statistics at a Glance

  • 91.5%: percentage of cyber-related jailbreak attempts declined by Astra in testing, up from 59% for GPT-5.6 Sol
  • 2: zero-day vulnerabilities autonomously uncovered by Astra in separate evaluations
  • 130: approximate number of tech and cybersecurity companies supporting the OpenAI-led cyber defense initiative

What Astra's Designation Means for You

The advancement of AI models like Astra raises the stakes for both security professionals and the broader public. As models become capable of conducting sophisticated cyberattacks independently, the need for robust safeguards becomes more urgent. The fact that a model can now achieve a 'Critical' classification under OpenAI's own framework signals a new reality: AI is not just a tool for defenders, but also a potential threat if not properly controlled.

For organizations, this emphasizes the importance of staying ahead of AI-driven threats and leveraging AI for defensive purposes. The industry's response, including the coalition of companies supporting cyber defense, reflects a collective recognition that AI's impact on security is profound. As deployment proceeds through controlled phases, the focus on alignment and control will be key to ensuring that such capabilities are used responsibly.

#openai#astra#cybersecurity#ai safety#preparedness framework

Sources

Iliyas

Founder & Editor, Xploitwire

This article was compiled from the sources listed above and checked against them for accuracy, under editorial policies set by Iliyas. Read our Editorial Policy →

← Back to all stories