OpenAI published on 7 August that it cannot rule out its upcoming Astra model reaching the Critical cybersecurity tier of its Preparedness Framework. It is the first time the company has attached that possibility to a specific model.
The distinction is the story
OpenAI did not say Astra is Critical. Benchmarking and assessment are ongoing, and the company has not confirmed the threshold was crossed. Headlines reading "OpenAI flags model as Critical" invert the news: the point is that OpenAI is acting on a possibility rather than a finding, which is what a precautionary framework looks like when it actually fires.
What Critical means
In OpenAI's own wording, the threshold is a tool-augmented model that "can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention" — or one that can "devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal."
What it committed to
OpenAI paused internal activities that did not meet upgraded security requirements. It listed isolated test environments, restricted network and tool access, enhanced model-weight protection and encryption, universal monitoring across agentic applications, and safety monitors that analyse chain-of-thought and halt high-risk activity. Third-party testing with government agencies and outside safety organisations is planned.
Pre-deployment, not post-incident
Astra has not shipped. That makes this a containment decision taken before anything reached users — the inverse of the sandbox-escape disclosures of the past fortnight, where labs described what had already happened. The Preparedness Framework has existed on paper for years; this is the first time it visibly slowed something down.
