Skip to content
ModelsReviewsstrong signalverified

Gemini 3.7 Flash, read from Google's own numbers

Google shipped it on August 13, 2026, 23 days after Gemini 3.6 Flash, with large gains on coding and agent benchmarks and an introductory price it labels as such. All of that holds. Four things are visible only if you open the model card instead of the launch post, and one of them is a score that went down.

By Redakcija WebAiRadarPublished 5 min readwritten by a model
#NamePriceBest for
1Gemini 3.7 Flash$0.75 and $3.75 per million through December 31, 2026, then $1.50 and $7.50The strongest cheap coding model on published numbers, on a price that expires and a cadence that will replace it. Read the card, not the post.
  1. 01

    Gemini 3.7 Flash

    $0.75 and $3.75 per million through December 31, 2026, then $1.50 and $7.50

    The strongest cheap coding model on published numbers, on a price that expires and a cadence that will replace it. Read the card, not the post.

    Good

    Large gains over 3.6 Flash on 16 of 17 benchmarks the card lists, including FrontierCode 1.1 at 43.6% against 34.4%, OSWorld-2.0 at 47.9% against 33.8% and AutomationBench at 30.4% against 17.0%. A one-million-token input window, text, image, audio and video input, and a full comparison table published rather than five selected rows. Red teaming reported as similar or improved against the predecessor.

    Not good

    CharXiv without tools drops from 85.2% to 84.5%, and that is the one result the launch post leaves out. The price is introductory, doubles on January 1, 2027, and is the same price 3.6 Flash already carries. The knowledge cutoff of March 2026 comes with some domains limited to January 2025. Output is capped at 64,000 tokens. The predecessor is 23 days old, which tells you how long any of this stays current.

    deepmind.google

We did not run a single evaluation on this model. This is a reading of what Google published — the launch post, the model card, the API changelog and the price list — and the value of that reading is not in repeating the scores. It is in the gap between the page written to announce the model and the page written to document it, because those two pages do not say quite the same thing.

What is claimed, and on what basis

The headline results are large and they come from Google. On FrontierCode 1.1 the model scores 43.6% against 34.4% for Gemini 3.6 Flash. On DeepSWE v1.1 it reaches 65.3%. On Terminal-bench 2.1 it goes from 78.0% to 85.8%, on AutomationBench from 17.0% to 30.4%, on OSWorld-2.0 from 33.8% to 47.9%, and on the long-context recall test GDM-MRCR v2 from 91.8% to 97.0%. The card lists 17 benchmarks with the predecessor's score beside each one, which is more disclosure than most launches carry.

The card also gives the shape of the thing: a one-million-token input window, a 64,000-token output limit, text, images, audio and video on input, a knowledge cutoff of March 2026, and red teaming that Google summarizes as similar or improved safety compared with 3.6 Flash, with none of the tracked critical capability thresholds reached.

So the model is what it says it is. The rest of this piece is about the four things that only the card tells you.

First: the same score is not the same on two Google pages

The launch post puts Gemini 3.6 Flash at 49.0% on DeepSWE v1.1. The model card puts it at 48.6%. Same model, same benchmark, same vendor, same day, two numbers. The improvement is therefore 16.3 points or 16.7 points depending on which Google page you were reading when you wrote it down.

The two pages also disagree about names. What the launch post calls WebDev Arena, the card calls Code Arena Web; the scores, 1588 against 1538, match. Neither of these is a scandal and neither changes the conclusion. They matter because a benchmark number is only useful if the person quoting it and the person checking it land on the same figure, and here they will not.

The practical rule: cite the card. It is the document that carries the whole table rather than the five rows chosen for the announcement, and it is the one a reader can check against the next model in the line.

Second: one score went down, and it is not in the announcement

On CharXiv without tools, the chart-reading benchmark, Gemini 3.7 Flash scores 84.5% where 3.6 Flash scored 85.2%. It is the only regression in the table, it is small, and it is entirely absent from the launch post, which highlights five benchmarks and picks none of the ones that moved the wrong way.

Nobody hid anything: the number is on Google's own card, published on the same day. But if you are replacing 3.6 Flash in a job that reads charts and figures out of documents, the release that improved 16 things degraded the one you depend on, and you would not learn that from the announcement you were sent.

The other number worth taking out of the card is Terminal-bench 3.0, where the score goes from 5.4% to 14.9%. That is nearly a tripling and it is still 14.9%. On the hardest terminal benchmark Google reports, the best cheap coding model available finishes roughly one task in seven. Read next to the 85.8% on Terminal-bench 2.1, that pair is the most honest description of where agents are right now, and only one half of it made the blog post.

Third: the price is not a property of this model

The launch post is direct about it: 3.7 Flash is available through the end of the year at an introductory price of $0.75 per million input tokens and $3.75 per million output. That is a plain statement of a time-limited price, and Google prints the date and the successor rates of $1.50 and $7.50 on the price list. Nobody is being misled here.

What the framing loses is that the price is not new and not tied to this model. Gemini 3.6 Flash costs exactly the same $0.75 and $3.75 today, under exactly the same notice expiring on December 31, 2026. The lower price arrived with 3.6 Flash, which the changelog describes as offering a lower price point than 3.5 Flash. Version 3.7 inherited it. If you read that this release halved the cost, the halving happened one release earlier.

The consequence for a budget is the one from the pricing table: on January 1, 2027 input, output, cache reads and cache storage all double, and a workload costing $6.00 per thousand short calls today costs $12.00 then. Choosing 3.7 Flash over 3.6 Flash is a decision about quality, because on price they are the same row.

Fourth: the knowledge cutoff has a second cutoff underneath it

The card gives March 2026 as the knowledge cutoff and then adds that some domains are limited to January 2025. That second clause is one line and it is worth more attention than its length suggests: for those domains the model's knowledge is fourteen months older than the number you would quote. Google does not say which domains, so the only safe reading is that a fresh-sounding cutoff does not describe every subject uniformly.

The other constraint the card carries and the post does not is the output limit. The input window is a million tokens; the output ceiling is 64,000. For summarizing, classifying and answering, that is far more than enough. For generating a long document in one pass, it is the wall you will hit, and it is the sort of limit that shows up in production rather than in evaluation.

What we did not check

We ran none of these benchmarks and cannot say whether the numbers reproduce. Every score here is Google measuring its own model against its own predecessor, which is normal practice and is also the reason to treat the size of a gain more cautiously than its direction.

We also did not test latency, throughput under load, tool-calling reliability, or behavior in the languages this site is written in. Those are the things that decide whether a cheap model is usable, none of them appear on a benchmark card, and the only way to learn them is to run your own traffic through it for a week.

Sources

BrandsGemini

Related

GPT-5.6 SOL$4/$20per million tokens, in and out
Modelsstrong signal

GPT-5.6 Sol drops to $4 and $20, and overtakes Claude Opus 5 on cost

OpenAI cut Sol's API price on August 21: input from $5 to $4, output from $30 to $20, cached input from $0.50 to $0.40. On a standard task the model goes from more expensive than Claude Opus 5 to cheaper than it. The cut is promotional and runs at least through November 21.

OpenAIverified

Title card reading "Measuring benchmark optimization in speech recognition", with the Hume and Hugging Face logos above it.
Modelsstrong signal

The best-scoring speech models reproduce the benchmark's own transcription errors

Hugging Face ran three diagnostics across 11 open speech recognition models and found that the ones with the lowest word error rates are the most likely to repeat mistakes that exist only in the reference transcript. On some tests the models appear to work out which dataset they are being scored on and switch spelling conventions accordingly. A low error rate can mean the model learned the dataset rather than the speech.

Hugging Faceverified

COST PER TASK$1.76the cheapest thousand 4,000-token calls
Modelsstrong signal

Best cheap models for high-volume work, priced per thousand calls

Six models, one task, one number: what a thousand calls cost when each sends 4,000 tokens in and gets 800 back. The cheapest row is $1.76 and the most expensive is $16.00, a nine-fold spread rather than the hundred-fold spread the category implies. Two things move the ranking more than the headline price does, and one of them has a date on it.

Anthropic, Google, OpenAIverified