// news · frontier-models · tools2026-08-08source: company announcement and reporting

Muse Spark 1.2 scores 82.9% on Terminal-Bench 2.1 — and Meta is competing on price, not the leaderboard

Meta's coding-focused update to the Muse Spark family scores 82.9 percent on Terminal-Bench 2.1 running inside Muse Code, edging GPT-5.6 Terra and Grok 4.5 while trailing Opus 5 at max effort in Claude Code. Meta is explicit that its differentiation is price rather than capability.

The benchmark position is a near-tie at the top of a crowded field, and Meta is not pretending otherwise. Muse Spark 1.2 is described as Muse Spark 1.1 with significantly scaled-up training compute on coding tasks and broader training-environment diversity — better code generation, debugging and codebase comprehension, general agentic capability held steady.

What is strategically interesting is the admission underneath. When a company ships a frontier-adjacent model and announces that the differentiator is cost, it is conceding that capability has stopped being a defensible moat at the top of the coding market. Four labs within a couple of points of each other on the same benchmark makes that concession look like an accurate reading rather than modesty.

The model was trained alongside its harness, which Meta argues is itself a source of coding performance. That is the same bet Anthropic makes with Claude Code — the harness is becoming part of the product rather than a wrapper around it.

See our analysis →

VentureBeat — Meta enters the AI coding wars with Muse Spark 1.2 and Muse Code → · CNBC — Meta debuts Muse Code to take on Anthropic and OpenAI → · Meta AI — Introducing Muse Spark 1.1 →