Claude Code retakes first in the agent rankings as Codex holds the Terminal-Bench record
A late-July refresh puts Claude Code back at number one on the strength of Opus 5 and per-subagent model control, with Codex at number two still holding the published Terminal-Bench record, and xAI's Grok Build the fastest-rising newcomer.
Per-subagent model control is the feature worth noticing. Letting an orchestrator assign different models to different subtasks turns the agent into a routing problem — cheap models for mechanical work, frontier models for the hard reasoning — and converts model choice from a purchase decision into a runtime one.
That Codex keeps the Terminal-Bench record while sitting second is a useful reminder of how little a single benchmark now determines. Leading on a measured axis and leading in practice have come apart, because practical agentic work depends on harness design, tool integration and recovery from failure as much as on raw model capability.
The churn itself is the story: rankings that reshuffle every few weeks, with a newcomer climbing fast and open-weight self-hosting holding a permanent slot. For teams choosing tooling, the rational posture is to assume today's leader is temporary and to prefer harnesses that let the model underneath be swapped without rewriting the workflow.
Levelop — Best AI coding agents in 2026: ranked by 90 days of use → · MightyBot — Best AI coding agents in 2026, ranked →