OpenAI repriced its API on Thursday and attached an unusual explanation to it: the company says its own frontier model rewrote the kernels that serve its models in production.
The prices
Per million tokens, GPT-5.6 Luna is now $0.20 in / $1.20 out, against $1 and $6 at the 9 July launch — an 80% cut. Terra fell 20%, to $2 and $12. Cached input runs a tenth of the miss price on both. Luna now sits below Gemini 3.1 Flash-Lite and at roughly a fifth of Claude Haiku 4.5's input price.
The flagship got nothing
Headlines have converged on "OpenAI cuts prices up to 80%". Sol is $5 in / $30 out today and was $5 / $30 at launch — a 0% cut. The 80% applies to the smallest model in the family. Axios managed the self-contradicting headline "OpenAI discounts GPT-5.6 Luna and Terra, but not Luna."
The causal claim runs backwards
OpenAI's changelog says that "with Codex, GPT-5.6 Sol autonomously rewrote and optimized our production kernels", in Triton and Gluon, delivering a 20% reduction in end-to-end serving costs and over 15% token-generation efficiency from speculative decoding. Some coverage has turned that into "Sol rewrote its own inference stack to fund the price drop". A 20% cost saving cannot fund an 80% price cut. The Luna repricing is a competitive move against Google and Anthropic; the kernel work is a separate — and arguably more consequential — claim. "Autonomously" is also carrying weight: the phrasing is "with Codex", an agentic harness with human review before anything reaches production.
Fast mode is a flat doubling
Priority Processing is gone, replaced by Fast mode, and requests tagged priority migrate automatically. Fast mode costs exactly double the standard rate on Sol, Terra and Luna alike. The changelog substantiates "up to 2.5x faster" for Sol only. Terra and Luna carry the same 100% premium with no published speed figure attached.
