The agent comes home
A competent 30B model that fits on a consumer card does not beat frontier systems. It removes the metered API from the most token-hungry workload in software.
Thirty billion parameters, 131K context, under 20GB quantised, tuned for tool use and LLM-as-judge. The specification reads like a shopping list written by someone who has actually run an agent loop.
Why agents broke the pricing model
A conversation is a few thousand tokens. A single agent task is tens of thousands across a dozen tool calls, and most of that volume is intermediate state — an error trace, a directory listing, a failed attempt — that no human will ever read.
Paying frontier per-token rates to have a model read its own scratch work is the defining economic problem of the category, and it is why every coding tool has spent 2026 rewriting its pricing.
Context is the enabling spec
The 131K window is what makes local viable. Agent loops accumulate; a model that must be re-primed every few steps cannot hold a task together, and no amount of raw capability compensates for amnesia. A hundred thousand tokens of working memory on hardware you own is a different product from the same weights behind a rate limit.
The honest ceiling
A distilled 30B will lose to a frontier model on the hard tail. It does not need to win there. The volume in agent systems is routine — classify this, extract that, route this, judge that output — and an enormous amount of it currently runs through paid APIs for no capability reason at all.
The counter-pressure is arriving in the same week: Cursor split its usage pools and added a $120 power seat, which is the hosted market pricing agentic load honestly for the first time. Local models and honest metering are the same story told from two ends.
VentureBeat — Meta returns to open source with Muse Glimmer, an Apache 2.0 licensed 30B parameter AI model optimized for agents → · Developers Digest — AI Coding Tools Pricing Comparison 2026 →