Cheaper and more correct, in that order
An 80 percent price cut and a 68 percent reduction in factually wrong answers shipped within a week of each other. Only one of those is a capability story.
GPT-5.6 Sol gained an effort slider and, on an internal evaluation of financial, medical and legal prompts, produced responses containing at least one factual error about 68 percent less often than GPT-5.5 Instant. Luna improved 62 percent and became the free default. A week earlier, Luna dropped 80 percent in price and Terra 20 percent.
Read the accuracy claim exactly as written
It is an internal evaluation. The prompts were selected for requiring factual detail. The metric is the share of responses containing at least one error, against a named predecessor. That is a well-specified, checkable claim and a real improvement.
It is not a statement about accuracy in general, and the chosen domains are the ones where whatever error rate remains does the most damage. A 68 percent reduction from an unstated base is compatible with a great many absolute outcomes.
The slider is the more honest shipment
Exposing compute-per-response as a user-facing control concedes something labs have been reluctant to say: there is no single right point on the quality-versus-latency curve, and the vendor cannot know which point a given question deserves. Handing that dial to the user is a small admission that the product is a tool rather than an oracle.
Why the cuts are asymmetric
If falling inference cost were driving this, both tiers would move together. Eighty percent at the bottom and twenty at the top is a deliberate widening, aimed at workloads that would otherwise leave for open weights or a cheaper rival. A trained model's marginal cost only falls, so low-end price is a strategy variable, not a constraint.
And the workloads it is aimed at are going somewhere specific. Chinese open-weight models took 41 percent of Hugging Face downloads and the top six slots on OpenRouter. An 80 percent cut says OpenAI would rather have that volume at a quarter of the revenue than watch it leave.
Cheaper and more correct is a good quarter for users. It is a harder quarter for anyone whose business model assumed capability would stay scarce.
OpenAI — Improving GPT-5.6 Sol in ChatGPT and expanding access to GPT-5.6 Luna for free users → · OpenAI — GPT-5.6 August updates, deployment safety hub → · CNBC — Hugging Face CEO says China is winning the AI race and dominating on open models →