// blog · analysis · frontier-models2026-08-13source: Artificial Analysis / SpaceXAI

Parity is a composite number

Grok 4.6 ties GPT-5.6 Sol on the index and loses the terminal by eight points. Both facts come from the same scorecard, and only one of them describes an agent.

Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, level with GPT-5.6 Sol, at $2/$6 per million tokens against Sol's $5/$30. The same suite records 26% on Terminal-Bench against 34.6% for Sol.

Averages hide spikes, and spikes are the product

Composite indices exist for a good reason: nobody can read forty leaderboards, and a single number lets a buyer make a decision this week. The cost is that averaging is lossy in one direction. It flattens exactly the variance a specialised workload cares about.

Grok 4.6 is a clean case. Ahead on knowledge work — 1753 Elo to 1728. Ahead on CursorBench 3.2 — 69.9% to 67.2%. Then a third behind on driving a shell to completion. Those are not contradictory results. They describe a model that is genuinely excellent at reasoning over text and meaningfully weaker at operating a machine.

Which number you need depends on what you are buying

For drafting, analysis and chat, the composite is honest and Grok 4.6 at $2/$6 is the value leader by a distance. For long-running agents living in a terminal, the composite is close to the wrong statistic, and a buyer using it will be surprised.

The industry got good at scoring what models know. It is still improvising on how to score what they do.

The evaluation layer is the weak link

This is not an isolated complaint. Benchmark scores move when you rephrase the question, and the same week produced a framework for scoring the benchmarks themselves. A field that measures everything has left its rulers unmeasured.

Read the composite as a starting hypothesis, not a conclusion. Then find the sub-score that matches your actual workload, because that is the one you will be living with.

Artificial Analysis — Grok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency → · SpaceXAI — Introducing Grok 4.6 → · MarkTechPost — SpaceXAI Releases Grok 4.6 →