OpenAI previews Ultrafast: GPT-5.6 Sol at 14x speed, running on Cerebras
Ultrafast is a new API service tier delivering up to 750 output tokens per second — up to 14x standard processing — for the same model. It is powered by Cerebras, capacity is limited, and OpenAI is vetting customers by workload fit.
The interesting word in the announcement is Cerebras. OpenAI has spent two years securing compute through NVIDIA, AMD and a custom Broadcom programme; serving its flagship model at 14x speed on a third party's wafer-scale hardware is a different kind of dependency, chosen for a property its own fleet does not have.
The specifics: same model, GPT-5.6 Sol, no distillation and no smaller variant — up to 750 output tokens per second against standard processing. Launching first in the API. Named target uses are incident response, customer service, financial market analysis and e-commerce, which share a property: the value of the answer decays fast.
Until now, real-time latency meant accepting a smaller or more specialised model. Decoupling speed from capability removes a trade-off that has shaped product design since the first chat interfaces, and it makes a category of application viable that was previously just slow.
The constraint is honest and stated: Cerebras capacity is limited, so access is granted on workload fit rather than on request. A performance tier gated by someone else's supply is a preview, not a product — and how quickly that changes is the thing to watch.
OpenAI — Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed → · TechCrunch — OpenAI introduces 'Ultrafast,' a new mode that makes GPT-5.6 Sol work at 14x the speed → · Neowin — OpenAI introduces new Ultrafast mode for GPT-5.6 Sol delivering 14x faster tokens →