// news · frontier-models · tools2026-08-13source: Artificial Analysis / reporting

Grok 4.6 ties on the composite and loses on the terminal by eight points

The same benchmark suite that puts Grok 4.6 level with GPT-5.6 Sol also records 26% on Terminal-Bench against 34.6% for Sol and 34.1% for Claude Fable 5. A composite index is an average, and averages hide exactly the thing agentic buyers are purchasing.

Composite intelligence indices exist because nobody can read forty leaderboards. They work by averaging, and averaging is lossy in a specific direction: it flattens the spikes. Grok 4.6 is a clean demonstration.

On the aggregate, 61 — the same as GPT-5.6 Sol. On knowledge work it is ahead, 1753 Elo to 1728. On CursorBench 3.2 it is ahead of Sol at 69.9% to 67.2%. Then on Terminal-Bench, which measures whether a model can actually drive a shell to completion, it records 26% against 34.6% and 34.1%. That is not noise. That is a third less of the thing an agent does all day.

Which number matters depends entirely on what is being bought. For chat, drafting and analysis, the composite is a fair summary and Grok 4.6 is excellent value at $2/$6. For long-running agents that live in a terminal — the workload the whole industry is currently pricing per seat, per agent, and now per agent-seat — the composite is close to the wrong statistic.

SpaceXAI has been criticised separately for not documenting what the model does autonomously. That gap and this one are the same gap wearing different clothes: the industry has gotten very good at scoring what models know and is still improvising on how to score what they do.

See our analysis →

Artificial Analysis — Grok 4.6 benchmarks and analysis → · MarkTechPost — SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents → · Forkast — Grok 4.6 matches GPT-5.6 Sol on composite intelligence →