The workhorse is the battleground
Google shipped a Flash model three weeks after the last one, with double-digit coding gains and half the price. Frontier prestige has moved to the tier that does the actual work.
Gemini 3.7 Flash took FrontierCode 1.1 Main from 34.4% to 43.6% and DeepSWE v1.1 from 49.0% to 65.3%, at $0.75/$3.75 per million tokens through year end.
Flash is not the small model any more
The mid-tier used to be where you compromised. It is now where the competition is, because agentic workloads call a model thousands of times per task and the flagship's economics do not survive that pattern.
The gains are broad rather than isolated — four benchmarks moving together, including AutomationBench from 17.0% to 30.4%. Concentrated improvement suggests targeting; distributed improvement suggests something real happened in training.
Everyone shipped this week
DeepSeek-V4-Pro on the 13th, GLM-5.3 on the 14th, Gemini 3.7 Flash on the 13th, and a 2.4T Qwen the day before. Four labs, one week.
Cadence is cheap to sustain for a quarter and expensive to sustain for a year.
The number that would actually settle something
Not the benchmark and not the release date — whether the same cadence exists in February. The labs shipping fastest are also the ones with the least public accounting of what shipping costs them, and a release schedule is the easiest thing in this industry to sustain temporarily.
Meanwhile the pricing attached to all this is moving in three directions at once. Whatever is being optimised for right now, it is not margin.
Google — Gemini 3.7 Flash: our most intelligent workhorse model → · VentureBeat — Google's Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut → · LLM Gateway — New AI Model Releases — August 2026 Timeline →