FP8 is a distribution decision
A model reissued at lower precision is not a research result. It is a statement about who the lab expects to run the thing, and it is the most useful signal in an open-weight release.
DeepSeek's MIT-licensed mixture-of-experts returned in an updated build shipping FP8 weights for cheaper inference. No new capability claim came with it.
The constraint on open weights was never the licence
An open-weight model's real limit is not whether you may run it. It is whether you can afford to. A licence that permits anything and a serving bill that permits nothing produce the same outcome.
Shipping FP8 at release removes a quantisation step every serious deployer was performing anyway — and removes the variance that came from everyone performing it slightly differently, on different tooling, with different results.
What it tells you about the lab
A lab that ships quantised weights has thought about your infrastructure. A lab that ships only full precision has thought about your benchmark. Those are different customers and the release tells you which one the lab is optimising for.
FP8-at-release is now something to check for. Its absence means you will be doing that work yourself.
The pricing pattern underneath
Read it alongside the same lab's hosted-tier movements this month, which went up. Raising hosted prices while lowering self-host cost is not incoherent — it is a statement that the durable position is in distribution rather than in serving.
That is a defensible bet. Hosted inference is a commodity with a price war running through it. Being the model everyone self-hosts is a position that compounds.
Where the size question goes next
The same week produced a 3B vision-language model tuned for edge hardware. Three billion parameters is where the deployment question stops being about cost and becomes about jurisdiction — the image never leaves the device, so an entire category of regulatory problem never arises.
Precision and parameter count are two knobs on the same question: how cheap does this have to be before the answer to "where does it run" becomes "wherever you like". Both moved this week.
The Open Weights — Open-source AI, tracked daily → · Thunder Compute — Best Open Source LLMs (August 2026) →