// news · frontier-models · open-source2026-08-06source: model card and independent analysis

DeepSeek V4 Flash 0731 ships: 13B active beats a 49B-active preview on every row

DeepSeek moved V4-Flash out of preview on 31 July with the same 284B-total, 13B-active architecture it launched with in April — re-post-trained rather than rebuilt. It now scores above the 49B-active V4-Pro preview on every Terminal-Bench and DSBench-Hard row, ships MIT-licensed with a 1M-token window, and runs at about 116 tokens per second.

Same architecture, better checkpoint, higher scores. That is the sentence worth sitting with, because it detaches capability from parameter count in a way benchmark tables usually obscure. Nothing about the model got bigger; the post-training got better, and a 13-billion-active model passed a 49-billion-active one across the board.

The practical consequence is serving cost. Active parameters, not total parameters, set the inference bill. A model that activates 13 billion per token while holding 284 billion in reserve is cheap to run and expensive to store — which is exactly the trade a self-hoster can absorb and an API vendor prices around.

MIT licensing with weights on Hugging Face makes this checkable rather than claimed. Anyone can download it, run the benchmarks, and disagree. That is a meaningfully different epistemic position from a scorecard published alongside an API endpoint.

See our analysis →

Hugging Face — deepseek-ai/DeepSeek-V4-Flash-0731 — model card → · Artificial Analysis — DeepSeek V4 Flash 0731 — intelligence, performance and price analysis → · LM Studio — DeepSeek V4 Flash — model listing →