Speed became a product
OpenAI is serving its flagship at 14x on Cerebras hardware. The interesting part is not the latency — it is that capability and speed have been unbundled.
Ultrafast runs GPT-5.6 Sol at up to 750 output tokens per second — same model, no distillation, powered by Cerebras and capacity-limited.
The trade that just disappeared
Every product built on these APIs since 2023 has made the same compromise: the good model is slow, the fast model is worse, pick one per surface. Whole architectures exist to route around it — cheap model for autocomplete, expensive model for the hard path, two prompt sets, two eval suites, two sets of failure modes to reason about.
Selling latency as a tier on an unchanged model deletes that. Not improves it — deletes it. The applications OpenAI names share one property: incident response, customer service, market analysis, e-commerce are all places where a late answer is a wrong answer.
Somebody else's silicon
The dependency is worth sitting with. OpenAI has spent two years securing compute through NVIDIA, AMD and a custom Broadcom programme, and is serving its flagship's fastest tier on a fourth party's wafer-scale hardware — because that hardware has a property its own fleet does not.
Capacity is limited, so access is granted on workload fit rather than on request. That is a preview, not a product.
Where it points
The unit of purchase keeps getting finer. First you chose a vendor, then a model, then a routing policy, and now a latency tier per call. Each step moves a decision out of architecture and into configuration.
Which is good for builders and awkward for anyone whose moat was that switching required a rewrite. It does not any more.
OpenAI — Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed → · TechCrunch — OpenAI introduces 'Ultrafast' → · Neowin — OpenAI introduces new Ultrafast mode for GPT-5.6 Sol →