// blog · analysis · open-source2026-08-05source: open-weights release history

Open weights is becoming a marketing term

Meta's Behemoth never shipped. Qwen's million-token flagship is API-only. The open-weight label increasingly describes a company's posture rather than what you can actually download and run.

Llama 4 launched in April 2025 and is still the current generation, with the two-trillion-parameter Behemoth widely treated as shelved. The open-weight frontier changed hands without a contest, an announcement, or most people noticing.

What Behemoth was actually for

It was never meant to be the model you ran. It was the teacher — the thing whose distillations would keep the smaller, deployable Llamas competitive for years. Losing it does not remove a model from the lineup so much as remove the mechanism that was supposed to keep the lineup fresh.

Which is why a fifteen-month-old generation is a bigger problem than it sounds. Llama 4's derivatives are ageing against labs shipping monthly, and the compounding runs the wrong way: every month without a teacher is a month the distillation gap widens.

The narrowing definition

Then there is the other half. Qwen 3.6 Plus Preview ships a million-token context window through the API and not as open weights. From the lab doing most to keep open weights moving, that is a meaningful line drawn around a flagship capability.

Nobody is being dishonest. Every lab running both an open line and a commercial one now follows the same sequence: the open release is real and useful, the newest capability appears behind an endpoint first. But a reader tracking the label rather than the release notes forms a steadily less accurate picture of what is actually runnable.

Long context is the wrong thing to hold back

Or rather, it is the most understandable thing to hold back and the most consequential. A million-token window is expensive to serve and hard to run locally, so the commercial logic is sound. The effect is that the gap between frontier capability and self-hostable capability widens along the axis hardest to close independently — you cannot compensate for missing context length with cleverness.

What replaced the old order is better

Mistral's sparse-MoE Large 3 at 675B total and 41B active with 256K context, the Ministral 3 edge line, Mistral Small 4, and Qwen's steady cadence. No single model dominates. That is healthier than one incumbent setting the ceiling for everyone, considerably harder to summarise, and much harder to take away — a distributed frontier has no single point of retreat.

And the strongest counterexample landed this week. MiniMax open-sourced H3, the best omni-modal video model to ship with weights available. The label is getting slipperier in language models and simultaneously more meaningful somewhere else.

AI CERTs News — Meta Behemoth cancel claim: inside Llama 4's uncertain future → · Okoone — Meta puts Llama 4 Behemoth on hold as questions rise → · MarkTechPost — Alibaba previews Qwen3.8-Max, a 2.4 trillion-parameter multimodal model →