BT

Facilitating the Spread of Knowledge and Innovation in Professional Software Development

Write for InfoQ

Topics

Choose your language

InfoQ Homepage News GPT-6 Astra Is the First Model OpenAI Classifies as Critical for Cybersecurity

GPT-6 Astra Is the First Model OpenAI Classifies as Critical for Cybersecurity

Listen to this article -  0:00

OpenAI has classified GPT-6 Astra at the Critical level for cybersecurity capability under its Preparedness Framework, the first model to reach that threshold. Microsoft made the model generally available in Foundry Models the same day.

The distribution is broader than a single cloud. OpenAI's own announcement listed availability through ChatGPT tiers, the API, and AWS, without naming Azure, which drew comment on Hacker News. One commenter drew the distinction that matters: being hosted on Azure is different from being provided by Azure, where the latter is a managed offering operated and billed by Microsoft under license.

The framework sets the Critical cybersecurity threshold at either of two conditions: a model that can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or one that can devise and execute end-to-end novel attack strategies against hardened targets given only a high-level goal. OpenAI states in the system card that Astra meets it.

The evaluations behind that conclusion are specific. In expert-led testing against a browser and an operating-system kernel, with experts permitted to supervise only for safety and validation and barred from contributing knowledge or directing research, Astra discovered multiple previously unknown vulnerabilities in both. It built a working exploit chain achieving unsandboxed code execution in the browser after 29 hours against a build later found to lack some production mitigations, then adapted the exploit to the official stable release in a further 12 hours. Against the OS kernel, it developed a working local privilege-escalation exploit within 12 hours.

On a benchmark built from vulnerabilities disclosed after the model's knowledge cutoff, Astra discovered and used unknown zero-day vulnerabilities previously as part of its exploit chains. OpenAI says it is disclosing two of those to maintainers, and is withholding product names, configurations and exploit mechanics to reduce risk to systems that may remain unpatched.

OpenAI says it strengthened cyber safeguards across the stack in response and updated its Trusted Access for Cyber program. Internal handling of the model changed too, with stricter isolation, checkpoint encryption, universal monitoring of full trajectories including chains of thought, and a blocking alignment evaluation before internal use.

A second finding in the same document complicates the safety picture. OpenAI reports that Astra's monitorability has decreased relative to GPT-5.6 Sol:

We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT-5.6 Sol, and less likely to include incriminating information in its CoT.

Under adversarial conditions, where researchers instructed the model to evade monitoring, Astra remained undetected while strategically underperforming on evaluations, a behavior OpenAI calls sandbagging, and sometimes evaded internal monitors on sabotage tasks. OpenAI found no evidence of steganographic reasoning, which it takes as an indication that evasion risk may be bounded to lower reasoning tasks, and notes that its alignment evaluations show Astra is less likely than Sol to violate security and safety restrictions overall. It also says the trend is taken seriously and underscores the need for alignment auditing techniques beyond examining chain of thought.

On alignment more broadly, OpenAI reports that in a simulation using more than 54,000 internal Codex tasks, Astra received roughly half as many flags for higher-severity misaligned behavior as Sol. Biological capability remains at the High level rather than Critical, with High safeguards retained.

Microsoft's Foundry announcement covers none of the cybersecurity classification, though it addresses the containment question the capability raises. Astra can interpret on-screen information and interact with approved interfaces, which Microsoft presents as a route into workflows without dedicated APIs. It then states the exposure:

Capability this direct demands containment. Content displayed in an application may be incomplete, misleading, or designed to influence an agent's behavior.

The recommended controls are scoped credentials, approved resources, human checkpoints for consequential actions, and activity records aligned to risk requirements. OpenAI's own testing reports that Astra is significantly more robust to prompt injection than Sol.

Pricing places the model at the top of the Foundry catalog. Standard Global runs $10 per million input tokens and $50 per million output on short context, rising to $20 and $75 on long context. The US Data Zone carries a 10% premium throughout. Cached input runs at a tenth of the input rate. For agentic workflows that accumulate context across steps, which is the use Microsoft positions the model for, the long-context tier is the one most likely to apply.

Deployment options cover Standard for variable demand and Provisioned Throughput for consistent latency, in Global and US Data Zone geographies. Foundry supplies Entra identity and access management, private networking, role-based access control, content filtering, safety evaluations and monitoring. Microsoft states that these help customers configure safeguards but do not eliminate risk or replace an organization's responsibility for selecting appropriate controls.

OpenAI says it will continue investigating the monitorability findings as models become more capable, and describes preserving chain-of-thought monitorability as a core goal of its research program.

About the Author

Rate this Article

Adoption
Style

BT