DeepSeek spent August warning developers that its API was about to get "significantly" more expensive while declining to publish a rate card, an effective date or an explanation. The rate card is now live, and the date was in the release note all along: "New pricing takes effect at 16:00 UTC, Aug 16, 2026." The flagship V4-Pro has left a single flat price for a peak/off-peak schedule in which peak is exactly double off-peak.
Both rate cards, side by side
Before, per million tokens: $0.435 input on a cache miss, $0.003625 on a cache hit, $0.87 output — one flat rate, no time-of-day tiering. After, from DeepSeek's own pricing page: cache-hit input $0.022 off-peak / $0.044 peak; cache-miss input $0.66 / $1.32; output $1.98 / $3.96. The cheaper V4-Flash lists cache-hit input from $0.007 and output at $0.66 off-peak. Both models carry a 1M-token context and a 384K maximum output.
What the common framing gets wrong
The headline everywhere is a 4.55x jump in output pricing, from $0.87 to $3.96. That is the peak rate only. Off-peak output is $1.98, so the honest range is 2.28x to 4.55x depending on the hour of day. Anyone running batch work outside the peak windows sees less than half the advertised increase.
The larger proportional move is the one going unreported. Cache-hit input rose from $0.003625 to between $0.022 and $0.044 — six to twelve times. For agent workloads, which replay a long system prompt and accumulated context on every turn, cache-hit input is frequently the dominant line on the bill. The effective increase for agentic use is therefore considerably worse than the output multiple everyone is quoting.
A third correction: this is not a price rise on an existing product. V4-Pro left preview on 12-13 August, and the cheap flat rate was preview pricing. "DeepSeek raises prices" implies a GA product got more expensive; what happened is that a preview subsidy ended.
The clock is set to Beijing's working day
Peak is defined as 01:00-04:00 and 06:00-10:00 UTC. Converted to China Standard Time that is 09:00-12:00 and 14:00-18:00 — the Chinese working day, morning and afternoon, with the lunch break carved out. The premium is aimed at domestic business hours, and Western users who schedule around it pay the off-peak rate by accident of geography.
A capacity signal, not just a price
DeepSeek's reputation rests on a single number: cost per token. At $3.96 per million output tokens at peak it is no longer an order of magnitude below Western frontier pricing — it is in the same bracket. Time-of-day pricing is also what an operator introduces to ration peak inference. That is the behaviour of a company that is compute-constrained, not compute-rich.
