The UK's AI Security Institute (AISI) published a measurement on July 17 that should focus the minds of security teams: the gap between freely downloadable open-weight models and the best closed frontier systems on offensive-cyber tasks has narrowed to roughly four to seven months, down from six to ten months a year earlier.
The finding
AISI evaluated leading open models — notably GLM-5.2 and DeepSeek V4-Pro — against closed frontier systems and asked a simple question: how far behind is the open ecosystem, measured in time? The answer is that the buffer defenders once relied on is shrinking fast.
How they measured it
On a "Narrow Cyber Tasks" benchmark of 70 tasks across four difficulty tiers, from non-expert to expert, GLM-5.2 matched closed models from about four months earlier, and DeepSeek V4-Pro matched those from about five months earlier. AISI also ran full attack simulations, not just isolated tasks.
The end-to-end test
On a "Cyber Ranges" scenario called "The Last Ones" — a 32-step attack against a simulated corporate network — GLM-5.2 performed comparably to Claude Opus 4.5, a model released roughly seven months before. Chaining dozens of steps into a coherent intrusion is the hard part, and open models are getting closer at it.
The open-model dilemma
The lag is, in effect, the head start defenders and regulators get before frontier offensive capabilities become impossible to restrict. Watching it fall from about ten months toward four sharpens a policy question AISI's own government is wrestling with: whether open-weight releases should face capability thresholds, and who decides. Both models AISI singled out — GLM-5.2 and DeepSeek V4-Pro — are Chinese open-weight systems, which adds a geopolitical edge: capability that ships as a free download cannot be fenced off by export controls the way advanced chips can.
