// blog · analysis · open-source2026-08-16source: Release trackers

FP8 is a distribution decision

A model reissued at lower precision is not a research result. It is a statement about who the lab expects to run the thing, and it is the most useful signal in an open-weight release.

DeepSeek's MIT-licensed mixture-of-experts returned in an updated build shipping FP8 weights for cheaper inference. No new capability claim came with it.

The constraint on open weights was never the licence

An open-weight model's real limit is not whether you may run it. It is whether you can afford to. A licence that permits anything and a serving bill that permits nothing produce the same outcome.

Shipping FP8 at release removes a quantisation step every serious deployer was performing anyway — and removes the variance that came from everyone performing it slightly differently, on different tooling, with different results.

What it tells you about the lab

A lab that ships quantised weights has thought about your infrastructure. A lab that ships only full precision has thought about your benchmark. Those are different customers and the release tells you which one the lab is optimising for.

FP8-at-release is now something to check for. Its absence means you will be doing that work yourself.

The pricing pattern underneath

Read it alongside the same lab's hosted-tier movements this month, which went up. Raising hosted prices while lowering self-host cost is not incoherent — it is a statement that the durable position is in distribution rather than in serving.

That is a defensible bet. Hosted inference is a commodity with a price war running through it. Being the model everyone self-hosts is a position that compounds.

Where the size question goes next

The same week produced a 3B vision-language model tuned for edge hardware. Three billion parameters is where the deployment question stops being about cost and becomes about jurisdiction — the image never leaves the device, so an entire category of regulatory problem never arises.

Precision and parameter count are two knobs on the same question: how cheap does this have to be before the answer to "where does it run" becomes "wherever you like". Both moved this week.

The Open Weights — Open-source AI, tracked daily → · Thunder Compute — Best Open Source LLMs (August 2026) →