The model, whose development was slowed in August due to security concerns, is here. GPT-6 Astra is the company's first model to achieve the highest internal risk level for cyber capabilities – and that's precisely why hardly anyone will get to see it initially.
In early August, OpenAI slowed down work on the model after internal audits indicated exceptional cybersecurity capabilities. This assessment has been confirmed: Following the completion of the security evaluation, Astra is the first model from the company to achieve the "Critical" level in its proprietary Preparedness Framework. The rollout will be phased accordingly.
Key Facts at a Glance
- GPT-6 Astra is the first OpenAI model to reach the critical risk level for cyber capabilities.
- The rating means that the model can find and exploit unknown security vulnerabilities without human guidance.
- Initially, only a limited number of organizations will have access; paid ChatGPT plans will follow in the coming days.
- The focus of the new generation is on computer use: The model operates programs itself, instead of suggesting steps.
- Codex adds a note function that preserves context across the boundaries of the context window.
What's behind the critical rating?
The "Critical" level is clearly defined in the Preparedness Framework. It applies when a model can develop functional zero-day exploits in many secured real-world systems without human intervention – or when it can independently derive and execute a complete attack strategy from a mere target specification.
According to OpenAI, both are true. The model achieved a perfect score in a benchmark for exploit development. Previous models had consistently been rated at the lower "High" level.
The consequence is a series of safeguards that did not previously exist in this form: continuous monitoring of all applications in which the model acts autonomously, examination of the thought process for risky actions, and the ability to terminate ongoing processes. The company had explicitly slowed the pace of model development in the summer in order to establish these precautions.
Access is granted in stages
Initially, only a limited number of organizations will have access, specifically participants in a cybersecurity-focused program. In the coming days, the model will be made available to all paying ChatGPT plans, via the API and Amazon's cloud platform.
In corporate environments, access to the launch is disabled by default and must be enabled by administrators. Usage is included in existing quotas; users who need more can purchase additional credits.
What needs to change on the computer
The practical focus is on computer use. The model is designed to enable users to operate programs independently, complete multi-stage workflows, and generate finished documents, spreadsheets, and presentations – even in the background while other work continues in parallel.
On the Mac, this builds upon an existing foundation. Since February, there has been a dedicated Codex app for macOS, which was merged with the ChatGPT application in July; the earlier version has continued to run as ChatGPT Classic ever since. The new model adds a note-taking feature to Codex: instead of merging long sessions into a single summary, the model retains notes across context boundaries, and older sections remain searchable. This feature is initially optional and is expected to become the default in the coming weeks.
Users of ChatGPT via Apple Intelligence will not initially be affected by the model change. Apple will determine which model will apply, and no announcement has been made regarding this.
Why the rating carries more weight than the benchmark values
Top marks in mathematics, software development, and computer use are the expected marketing ploy for a new generation of models. The remarkable part of this announcement is that a company is publicly classifying its own product as a critical risk and therefore restricting access to it.
For readers with a professional background in IT security, the practical consequence is clear: capabilities of this kind will no longer be available exclusively to government actors. This shifts the demands on protection and detection, regardless of how reliably a single provider's security measures function.
We expect this self-assessment to be followed by other providers, who will publish similar tiered models. The evidence supporting this is the development of recent months: Since the first model was classified as high-risk in February, OpenAI has tightened its safeguards with each new release.
What happens next with the model?
Whether the phased rollout lives up to its promise will become clear in the coming weeks when the model is integrated into the regular tariffs. Until then, it remains an announcement with an unusually large number of reservations from the provider itself.
Would you allow an AI model to run independently on your Mac if the manufacturer also lists it as a critical security risk – or is that a deal-breaker for you? Feel free to disagree in the comments.


