OpenAI puts Astra in its critical cyber tier, and the safeguards will stop legitimate work too
OpenAI says its unreleased Astra model meets the Critical cybersecurity threshold in its Preparedness Framework, the first model it has designated that way. In testing the model found previously unknown flaws in a hardened browser and operating system and chained them into working compromises. Parts of the model's development and release were delayed while protections were rebuilt. The part that matters outside the lab is the deployment note: the new monitoring can pause or stop work that has nothing to do with security.
Source
Path to Astra: critical capabilities and frontier safeguardsOpenAI News · Original published September 1, 2026
OpenAI published its assessment of the Astra model on September 1, 2026, ahead of the model's release. The company now believes Astra meets the Critical cybersecurity capability threshold under its Preparedness Framework. It describes that as being able to find previously unknown security flaws and build ways to exploit them, across many well-protected systems, without a person guiding each step. It is the first model OpenAI has placed at that level.
What the designation means
The Preparedness Framework sets the Critical threshold at either of two conditions. The first is a model that can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems, without human intervention. The second is a model that can devise and execute end-to-end novel attack strategies against hardened targets, given only a high-level goal. OpenAI says its evaluation of Astra combined automated public and private benchmarks with expert-driven assessments, and that the conclusion follows from both.
The designation is a decision by the company about its own model, not an outside audit. No independent verification of these results has been published, and OpenAI says the fuller account of its safety, security and alignment testing will arrive in the system card at launch.
What Astra did in testing
The reported results are specific enough to be checked later, which is more than most capability claims offer. OpenAI also flags a limitation that is easy to miss: the Astra figures reflect capabilities with Daybreak Blue access, not the default production configuration.
The company lists these findings.
- On ExploitBench, which measures building exploits from known vulnerabilities, Astra scored 100%.
- On an internal set of 20 recently disclosed high-severity V8 vulnerabilities, built to avoid benchmark contamination, the model reached higher arbitrary code-execution rates than GPT-5.6 Sol while spending far fewer output tokens.
- During that evaluation it found and used two zero-day vulnerabilities as part of an exploit chain, which OpenAI says it is disclosing to the maintainers.
- Against a hardened browser it built a full compromise chain that escaped the sandbox and ran commands on the host when the browser opened an HTML file. Against a hardened operating system it chained several flaws into a local privilege escalation from an unprivileged user to root.
What changed after the Hugging Face incident
OpenAI states that Astra was not involved in the Hugging Face incident, and that retrospective testing suggests the production safeguards in place at the time would have prevented it. It still paused certain frontier training for two weeks afterwards to harden its training infrastructure, including isolation and network controls, expanded monitoring and stricter alignment thresholds. The large reinforcement learning run that was held back restarted on August 28, 2026, and some smaller experimental runs remain paused.
On the model layer, OpenAI reports that Astra refuses 91.5% of requests in its cyber jailbreak evaluation set, against 59% for GPT-5.6 Sol. It also describes a honeypot test derived from the incident, in which GPT-5.6 Sol without production safeguards tried to reach out-of-scope targets in 56% of runs while Astra made no such attempt. Those figures describe behavior under test conditions with safeguards off, which the company says explicitly.
What this costs the people using it
Access to Astra's advanced cybersecurity workflows will start with a small group of alpha testers, and widen afterwards through Daybreak Blue for defensive use. OpenAI expects the safeguards at launch to create more friction than it ultimately intends.
The concrete consequence sits in the deployment note. The system may flag legitimate activity as potential cyber misuse or unauthorized behavior, including work with no obvious relation to security and agents that run for a long stretch. In ChatGPT or Codex the user is asked to review the action before continuing. On the API there is no such prompt: the task stops. Anyone building an unattended pipeline on an Astra-class model has to treat a halt as a normal outcome rather than an error.
„Extra safety checks can sometimes slow, pause, or stop legitimate work“
Related
NeoMME's 260M encoder lands within 0.002 of a 3.75B model on ViDoRe v3
H company published NeoMME on September 3, 2026: a pair of multimodal encoders, at 260M and 800M parameters, released under Apache 2.0 and loadable through Hugging Face Transformers. On the ViDoRe v3 document retrieval benchmark the small one scores 0.523 nDCG@10, which is 0.002 behind ColQwen2.5 at roughly 14 times its parameter count. It also encodes about 51 pages per second on a single NVIDIA L40S, and the index it produces can be compressed from about 1.5 MB per page to 6 kB while keeping more than 95% of retrieval quality.
Hugging Faceverified
ChatGPT, Claude, Grok, and Gemini all had trouble inside the same two hours on September 3
On September 3, 2026, Anthropic logged elevated errors across Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5, OpenAI logged elevated errors across ChatGPT and Codex, Grok showed users an error message, and third-party monitors recorded a likely Gemini API interruption. Four independent providers, one morning. No status page names a shared cause, and Amazon Web Services, Microsoft Azure, and Cloudflare reported nothing major. For anyone whose product calls more than one of these, this is the failure mode that a second provider does not fix.
Ars Technicaverified
OpenAI ships GPT-6 Astra at $10 per million input tokens, with enterprise access off by default
GPT-6 Astra went out on September 3, 2026, first to a limited set of organizations and then, over the following days, to ChatGPT Plus, Pro, Business, and Enterprise. In the API it answers to the name gpt-6-astra at $10 per million input tokens and $50 per million output tokens, and it is also served through Microsoft Azure and Amazon Bedrock. Enterprise workspaces do not get it automatically, because access is off by default at launch and an administrator has to turn it on. OpenAI also says the extra safety checks around this model can pause work in ChatGPT and Codex and stop it outright in the API.
OpenAIverified


