DeepSeek moved V4-Flash from preview to an official-version API on Friday morning Beijing time. The interesting part is not the release — it is that the company's own benchmark table puts its cheap model above its expensive one, and the expensive one has still not shipped.
The scores and the prices
From DeepSeek's changelog: Terminal Bench 2.1 82.7, Cybergym 76.7, Toolathlon verified 70.3, DSBench-FullStack 68.7, DeepSWE 54.4, NL2Repo 54.2. Context window 1M, maximum output 384K, thinking mode switchable. Pricing per million tokens: Flash $0.14 in / $0.28 out against Pro at $0.435 / $0.87. Concurrency caps are 2,500 for Flash and 500 for Pro. A peak-hour surcharge doubling rates during Beijing business hours is flagged as coming.
Three things the coverage is getting wrong
First, "official version" is not a new model. DeepSeek's changelog states that V4-Flash-0731 "keeps the same model architecture and size" as the preview and "was only re-post-trained". Architecture and parameter count are unchanged — this is a post-training refresh being reported as a launch. Second, it is API-only. No weights were released, and only the V4-Flash interface was upgraded; the app, the web models and the V4-Pro API were untouched. Anyone reading this as an open-weights drop is wrong. Third, the benchmark table is DeepSeek's own, and the comparison set is flattering by construction: it is scored against DeepSeek's own V4-Pro-Preview — a preview, not a shipped product — and against Zhipu's GLM-5.2. No frontier Western model appears in it.
Why the ladder inverting matters
A cheap tier out-scoring the expensive tier on agent tasks, at roughly a third of the price, is a live signal that current gains on agentic work are coming from post-training rather than from scale. It also leaves V4-Pro in an awkward position: still in preview, and now beaten by its smaller sibling on the maker's own numbers.
One figure not to print
Chinese outlets are citing 284 billion total and 13 billion active parameters. That does not appear on any DeepSeek page we could verify, and should not be repeated as fact.
