Gemini 3.7 Flash: half the price and markedly better on code
Google released Gemini 3.7 Flash on August 13, just three weeks after 3.6. The gain on coding benchmarks is large: the DeepSWE v1.1 score rises from 49.0% to 65.3%, and FrontierCode 1.1 from 34.4% to 43.6%. Introductory pricing, in force through the end of 2026, is $0.75 per million input tokens and $3.75 per million output tokens — half what its predecessor launched at. The model is available through the Gemini API, in AI Studio, Android Studio, and Antigravity, and in Spark for AI Pro and Ultra subscribers.
Source
Gemini 3.7 Flash launches three weeks after last model, live in Spark9to5Google · Original published August 13, 2026
Google released Gemini 3.7 Flash on August 13, just three weeks after 3.6. The gain on coding benchmarks is large, and the introductory price is half what its predecessor launched at.
Scores
- DeepSWE v1.1: from 49.0% to 65.3%
- FrontierCode 1.1 Main: from 34.4% to 43.6%
- WebDev Arena, Elo rating: from 1538 to 1588
- GDP.pdf: from 22.0% to 34.0%
- AutomationBench: from 17.0% to 30.4%
Price
Introductory pricing runs through the end of 2026 and is $0.75 per million input tokens and $3.75 per million output tokens — half the price its predecessor launched at.
Where it is available
The model is used through the Gemini API, in AI Studio, Android Studio, and Antigravity, through the Gemini Enterprise platform and app, and in Spark for AI Pro and Ultra subscribers.
How it fits the wider picture
This is cheap labor for code. Against the price of Opus 5 the difference is an order of magnitude, which explains why mixed setups keep appearing: a cheap model does the mechanical part of the job, and a strong model makes the decisions.
Sources
Related
GPT-5.6 Sol drops to $4 and $20, and overtakes Claude Opus 5 on cost
OpenAI cut Sol's API price on August 21: input from $5 to $4, output from $30 to $20, cached input from $0.50 to $0.40. On a standard task the model goes from more expensive than Claude Opus 5 to cheaper than it. The cut is promotional and runs at least through November 21.
OpenAIverified

The best-scoring speech models reproduce the benchmark's own transcription errors
Hugging Face ran three diagnostics across 11 open speech recognition models and found that the ones with the lowest word error rates are the most likely to repeat mistakes that exist only in the reference transcript. On some tests the models appear to work out which dataset they are being scored on and switch spelling conventions accordingly. A low error rate can mean the model learned the dataset rather than the speech.
Hugging Faceverified
Gemini 3.7 Flash, read from Google's own numbers
Google shipped it on August 13, 2026, 23 days after Gemini 3.6 Flash, with large gains on coding and agent benchmarks and an introductory price it labels as such. All of that holds. Four things are visible only if you open the model card instead of the launch post, and one of them is a score that went down.
Googleverified
