// news · frontier-models2026-08-17source: Release trackers

Three weeks between Flash releases, and evaluation cannot keep up

Gemini 3.7 Flash landed three weeks after 3.6 Flash. When the gap between frontier releases is shorter than the time it takes to evaluate one, the published numbers describe a model that is already superseded.

Google shipped Gemini 3.7 Flash three weeks after 3.6 Flash, with DeepSWE v1.1 moving 49.0% to 65.3% and AutomationBench 17.0% to 30.4% across that gap.

The size of the jump is the story people tell. The interval is the more consequential fact. Serious third-party evaluation — building a harness, running it, checking for contamination, writing it up — does not happen in three weeks. Neither does enterprise procurement, security review, or a migration.

The practical result is that independent numbers now describe a model one or two versions behind whatever is being sold. Buyers are left choosing between vendor-published results, which are fast and interested, and independent results, which are slow and honest. Neither is what a purchasing decision actually needs.

It also changes what a benchmark is for. When a board position lasts three weeks, topping it is a marketing event rather than a durable claim about capability. The field has not yet built the equivalent of a rolling evaluation that assumes the target moves — the closest thing is continuously-updated leaderboards, and those carry their own cost problem.

None of this argues the improvements are unreal. It argues that the reporting apparatus around them was designed for an annual cycle and is being run at a monthly one. Research on cheaper post-training suggests the cadence is going to get faster, not slower.

See our analysis →

Memeburn — Google Launches Gemini 3.7 Flash but Its Low Price Has an Expiry Date → · DataNorth — Google releases Gemini 3.7 Flash for coding and agents →