OpenAI opened a limited preview of Ultrafast on 13 August, an accelerated serving mode for GPT-5.6 Sol that the company says runs up to 14x faster than standard inference, reaching up to 750 output tokens per second.

The part that is new

The speedup comes from a partnership with Cerebras. OpenAI's serving fleet is otherwise Nvidia-based, and this is the first time the company has publicly run a frontier model on a different vendor's accelerator at production scale. Cerebras builds wafer-scale processors that hold model weights in on-chip memory, which is why its inference numbers have always been fast — the open question was whether a frontier lab would ship on them.

What 750 tokens a second changes

At standard frontier speeds, a long agent trajectory is measured in minutes. At this rate a model finishes a paragraph faster than a person reads it, which moves whole product categories from asynchronous to interactive. OpenAI named incident response, customer service, financial analysis and e-commerce — all cases where the wait, not the answer, is the constraint.

The limits as stated

This is a preview for a select set of customers, and OpenAI has said it will widen access "as capacity grows" — an admission that Cerebras capacity, not demand, sets the ceiling. No pricing has been published, and the 14x and 750 figures are upper bounds rather than sustained throughput.

Why it matters beyond latency

Every large buyer of inference has spent two years trying to establish a second source. A frontier lab serving its own flagship on non-Nvidia hardware is the strongest existence proof yet that the alternative silicon works for something other than benchmarks.