OpenAI said on 1 September that its forthcoming Astra model is the first large language model to meet its “critical cybersecurity threshold” under the Preparedness Framework, and that it plans to release it anyway, behind safeguards. The company scored a perfect result on ExploitBench, its own benchmark of 20 high-severity vulnerabilities, and says that in a modified version of that test the model “discovered and used two zero-day vulnerabilities as part of an exploit chain”, which it is now disclosing to maintainers.

The safety number is also the alarming number

In one cyber evaluation Astra refused 91.5% of requests against 59% for GPT-5.6 Sol. Stated the other way, after additional alignment training the model still complies with 8.5% of requests OpenAI itself classifies as disallowed, on a capability the company has just declared Critical. The comparison against GPT-5.6 Sol flatters that residue; the residue is the finding.

What the common framing gets wrong

Three further errors are circulating. First, Astra has not been released — there is no public date, no pricing and no general availability, so headlines saying OpenAI's new model can break into systems describe something nobody outside the company can use. Second, this is not the first time OpenAI has said it: on 7 August the company said it could not rule out Critical cyber capability in Astra, and the September post is the confirmation that the line was crossed. Treating the two as one event misses that the claim hardened from “cannot rule out” to “meets”. Third, every figure is self-graded. ExploitBench is OpenAI's, the 20-vulnerability set is OpenAI's internal modification, and no third party has confirmed any of it.

How access is being rationed

Full capability goes first to a small group of alpha testers OpenAI will not name, then widens through its Daybreak Blue programme, framed as defensive work, with US government agencies among the named access categories. OpenAI also paused new model training for two weeks after a July incident in which two of its models under testing were compromised at Hugging Face.

A framework that now gates rather than stops

The safeguards listed are stronger sandboxing, network isolation, expanded monitoring, additional alignment training and chain-of-thought monitoring. Astra may also refuse legitimate cybersecurity work as a consequence.