The cheap model tops the legal and long-context boards, not just the coding ones
Gemini 3.7 Flash reaches 90.7% on Harvey LAB-AA and 97.0% on GDM-MRCR long context. The coding numbers got the attention; these two say more about where the tier boundary went.
Alongside the coding results that drew most of the coverage, Gemini 3.7 Flash tops Harvey LAB-AA at 90.7% and GDM-MRCR long context at 97.0%, measured against Claude Sonnet 5 and GPT-5.6 Terra.
Those two boards matter more than they look. LAB-AA is legal analysis, a domain that has spent two years being told it needs the largest available model because the cost of a wrong answer is high. Long-context retrieval is the other one: the standard argument for paying frontier prices has been that only the big models hold a long document in working memory without losing the thread.
A model in the cheap tier taking both undermines the tier itself. The pitch for a flagship was never only raw capability — it was that certain workloads were categorically unsafe below a certain size. If the sub-dollar model reads a contract as well, the categorical argument becomes a preference.
Two cautions. Benchmark tops are narrow: LAB-AA measures a defined analysis task, not the judgement a firm is actually buying. And the price these results ship at is introductory, reverting on 1 January 2027 — a board position won at $0.75 per million and defended at $1.50 is a different commercial proposition.
The direction is consistent with the rest of the week: capability keeps moving down the price ladder faster than the ladder gets rebuilt.
DataCamp — Gemini 3.7 Flash: Features, Benchmarks & Pricing → · VentureBeat — Google's Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut →