// news · open-source2026-08-04source: felloai / huggingface

Thinking Machines follows Inkling with Inkling-Small — a quarter the size, most of the performance

Two weeks after releasing Inkling, its 975B-parameter multimodal flagship and first public model, Mira Murati's Thinking Machines Lab shipped Inkling-Small on July 31 — a distilled variant at roughly a quarter of the size that retains most of the performance, with weights downloadable on Hugging Face.

The cadence is the story. A new lab's first public model is a statement; following it within two weeks with a distilled variant is a strategy. Flagship-then-distill in one motion says Thinking Machines is building a family, not a demo — the same release grammar the incumbent labs took years to settle on, executed on the first try.

Distillation is where open-weight economics actually live. A 975-billion-parameter multimodal model is a research artifact for most users; a quarter-size variant keeping most of the capability is something enterprises can actually serve. The gap between what a lab can train and what a customer can run is the market, and Inkling-Small is aimed straight at it.

It also confirms how crowded the open frontier has become. Inkling arrives into a field where Kimi K3's 2.8-trillion-parameter weights, GLM-5.2, Gemma 3, and Mistral Medium 3 are all downloadable from the same platform. For a lab founded on frontier pedigree, open distribution is no longer a differentiator — it is table stakes, and the competition is on quality per deployable parameter.

See our analysis →

Fello AI — Best AI models in August 2026 → · Hugging Face — State of open source on Hugging Face: Spring 2026 →