Best cheap models for high-volume work, priced per thousand calls
Six models, one task, one number: what a thousand calls cost when each sends 4,000 tokens in and gets 800 back. The cheapest row is $1.76 and the most expensive is $16.00, a nine-fold spread rather than the hundred-fold spread the category implies. Two things move the ranking more than the headline price does, and one of them has a date on it.
| # | Name | Price | Best for |
|---|---|---|---|
| 1 | GPT-5.6 Luna | $1.76 per 1,000 calls, $1.04 cached, $0.88 batched | The cheapest way to run a short task a hundred thousand times, as long as your prompt stays short. |
| 2 | Gemini 2.5 Flash | $3.20 per 1,000 calls, $2.12 cached, $1.60 batched | The quiet option: an older model whose price nobody has announced a change to. |
| 3 | Gemini 3 Flash | $4.40 per 1,000 calls, $2.60 cached, $2.20 batched | Cheap, current, and still labeled preview, which is a term about your migration risk rather than about quality. |
| 4 | Gemini 3.7 Flash | $6.00 per 1,000 calls, $3.30 cached, $3.00 batched — and $12.00 from January 1, 2027 | The best code scores in the cheap tier, at a price that is only this price until the new year. |
| 5 | Claude Haiku 4.5 | $8.00 per 1,000 calls, $4.40 cached, $4.00 batched | The most expensive of the cheap models, and the one whose bill has no cliff, no clock and no date on it. |
| 6 | Claude Sonnet 5 | $16.00 per 1,000 calls, $8.80 cached, $8.00 batched | Here as the ceiling. It is what you pay when the task turns out not to be a cheap-tier task. |
- 01
GPT-5.6 Luna
$1.76 per 1,000 calls, $1.04 cached, $0.88 batchedThe cheapest way to run a short task a hundred thousand times, as long as your prompt stays short.
Good
Input at $0.20 and output at $1.20 per million, the lowest pair in the table. Cache reads at $0.02 per million cut the input side to almost nothing. Batch halves both columns again.
Not good
Every one of those prices is the short-context price. Long context roughly doubles them, to $0.40 and $1.80, and OpenAI does not publish the token count where one becomes the other. You cannot check in advance which price a given prompt will be billed at.
- 02
Gemini 2.5 Flash
$3.20 per 1,000 calls, $2.12 cached, $1.60 batchedThe quiet option: an older model whose price nobody has announced a change to.
Good
At $0.30 and $2.50 per million it is the second cheapest row, and the cheapest with no expiry date attached. Cache reads cost $0.03 per million. Text, image and video share the input price.
Not good
Audio input is billed at $1.00 per million, more than three times the text rate. Cache storage runs at $1.00 per million tokens per hour, so a large kept cache costs money on days you barely call it. It is also two generations behind on code.
- 03
Gemini 3 Flash
$4.40 per 1,000 calls, $2.60 cached, $2.20 batchedCheap, current, and still labeled preview, which is a term about your migration risk rather than about quality.
Good
At $0.50 and $3.00 per million it sits between the two Flash generations on price. Cache reads are $0.05 per million. Text, image and video all enter at the same rate.
Not good
Preview terms can change with no committed migration window, which is the wrong footing for something you wire into production. Audio doubles the input price to $1.00, and cache storage is $1.00 per million tokens per hour.
- 04
Gemini 3.7 Flash
$6.00 per 1,000 calls, $3.30 cached, $3.00 batched — and $12.00 from January 1, 2027The best code scores in the cheap tier, at a price that is only this price until the new year.
Good
The strongest published coding results of anything in this table, and the cheapest cache storage among the Gemini rows at $0.50 per million tokens per hour through 2026. All four prices are printed with their end dates, so nothing about the increase is a surprise.
Not good
On January 1, 2027 input, output, cache reads and cache storage all double. At the new prices this becomes the most expensive row in the table, behind Claude Haiku 4.5. Gemini 3.6 Flash now costs exactly the same, so the launch line about halved pricing describes a price its predecessor already has.
- 05
Claude Haiku 4.5
$8.00 per 1,000 calls, $4.40 cached, $4.00 batchedThe most expensive of the cheap models, and the one whose bill has no cliff, no clock and no date on it.
Good
One dollar in and five out per million, with cache hits at ten cents and no storage fee for keeping a cache alive. Anthropic's own worked example puts ten thousand support conversations at about $37. It uses the older tokenizer, so its per-token price compares directly with models up to Sonnet 4.6.
Not good
At list price it is four and a half times Luna for the same task. It is a small model and it behaves like one on long multi-step reasoning. Pinning inference to the United States adds a 1.1 times multiplier on every token category.
- 06
Claude Sonnet 5
$16.00 per 1,000 calls, $8.80 cached, $8.00 batchedHere as the ceiling. It is what you pay when the task turns out not to be a cheap-tier task.
Good
Two dollars in and ten out per million, and the introductory price is now the standard one: the increase to $3 and $15 scheduled for September 1, 2026 was cancelled. The full million-token context window is billed at standard rates.
Not good
Nine times Luna at list price for the same thousand calls. It uses the newer tokenizer, which produces about 30% more tokens for the same text, so its per-token price is not directly comparable with models up to Sonnet 4.6 — the real gap is wider than the columns suggest.
Price per million tokens is the number every vendor publishes and nobody feels. You do not buy a million tokens. You buy one call that then happens two hundred thousand times a month, and what you want to know is what that costs. So this list prices a single task — 4,000 tokens in, 800 tokens out, the shape of a classification or a short summary — and reports the bill for a thousand of them. Every price was copied from the vendor's own list on August 21, 2026, and every row links back to it.
How the bill is built
One call costs 4,000 input tokens at the input rate plus 800 output tokens at the output rate. Scale that to a thousand calls and the formula collapses to something you can do in your head: four times the input price, plus eight tenths of the output price, in dollars. Gemini 3.7 Flash at $0.75 and $3.75 per million comes to three dollars of input and three dollars of output, so $6.00 per thousand calls.
Run the same arithmetic across the field and the order is GPT-5.6 Luna at $1.76, Gemini 2.5 Flash at $3.20, Gemini 3 Flash at $4.40, Gemini 3.7 Flash at $6.00, Claude Haiku 4.5 at $8.00, and Claude Sonnet 5 at $16.00. Sonnet 5 is in the table as the ceiling, not as a candidate: it is the price you pay when you decide the task deserves a stronger model.
Nothing here measures the answer. A cheap model that needs a second attempt on three calls in ten is not a cheap model with an asterisk; it is a model whose real price is 30% higher, whose median latency doubles on those calls, and whose failures you now have to detect. Measure that on your own traffic before you move a workload for a two-dollar difference.
Caching moves the decision into the output column
In high-volume work the 4,000 input tokens are usually the same 4,000 tokens: a fixed instruction block, a schema, a few examples. All three vendors will read that from cache at roughly a tenth of the input price. Anthropic charges cache hits at 0.1 times the base input rate, so ten cents per million on Haiku 4.5. Google reads a Gemini 3.7 Flash cache at $0.075 per million and a Gemini 2.5 Flash cache at $0.03. OpenAI reads a Luna cache at $0.02.
Apply that to the same thousand calls and the rows become $1.04, $2.12, $2.60, $3.30, $4.40 and $8.80. The order does not change. What changes is what you are comparing: once the input side has shrunk to a few cents, the input column stops mattering and the output column decides the whole ranking. Per thousand calls, output alone costs $0.96 on Luna, $2.00 on Gemini 2.5 Flash, $4.00 on Haiku 4.5 and $8.00 on Sonnet 5.
The practical rule that follows: if you are choosing a model for cached, repetitive work, sort the candidates by output price and ignore the input column until the shortlist is down to two.
Google bills cache storage by the hour, Anthropic bills the write
The two vendors charge for the same feature in structurally different places, and the difference decides which of them is cheap for you. Google charges a storage fee while the cache is alive: $0.50 per million tokens per hour on Gemini 3.7 Flash through the end of 2026, and $1.00 per million tokens per hour on Gemini 3 Flash and Gemini 2.5 Flash. Anthropic charges no storage and instead marks up the write — 1.25 times base input for a five-minute cache, 2 times for a one-hour cache — after which reads cost a tenth of base. Their own documentation puts the break-even at one read for the short cache and two reads for the long one.
Storage is a cost that ignores your call volume, which is exactly what makes it a trap at the low end. A 200,000-token document held in a Gemini 3 Flash cache for a day costs $4.80 whether you hit it twice or twenty thousand times. At two hundred thousand calls a month that disappears into the noise. At two hundred calls a month it is most of your bill, and a model with a write fee and no clock is the cheaper structure.
So the question is not which vendor caches more cheaply. It is whether your cache spends its life being read or being kept.
The cheapest Flash row has an expiry date printed on it
Gemini 3.7 Flash costs $0.75 and $3.75 per million through December 31, 2026. On January 1, 2027 those become $1.50 and $7.50, the cache read goes from $0.075 to $0.15, and cache storage goes from $0.50 to $1.00 per million tokens per hour. Google prints all four dates on the price list, which is more than most vendors do, and it is still an increase you have to plan for. Gemini 3.6 Flash carries the identical notice.
Put the January prices into the same task and Gemini 3.7 Flash goes from $6.00 to $12.00 per thousand calls. That is not a small adjustment to a ranking; it moves the model from fourth place to last, past Claude Haiku 4.5 at $8.00 and half again above it. Haiku 4.5 has no dated change on its list, which is not a promise of anything, but it is at least not a scheduled increase.
If you are writing a budget that runs into next year, write two: one for the four months at the current price and one for everything after.
Batch halves everything, and therefore settles nothing
All three vendors publish an asynchronous tier at half price on both columns. Claude Haiku 4.5 batches at $0.50 and $2.50 per million, Gemini 3.7 Flash at $0.375 and $1.875 through the end of 2026, GPT-5.6 Luna at $0.10 and $0.60. Every row in the table halves and the order holds, which is the useful finding: batch is the one lever that is identical across the field, so it changes your total and never your choice.
Two of the halved numbers are worth carrying around. Batched Claude Sonnet 5 costs $8.00 per thousand calls, exactly what unbatched Claude Haiku 4.5 costs. Batched Haiku 4.5 costs $4.00, less than Gemini 3.7 Flash at list price. If the work can wait a few hours — overnight enrichment, nightly classification, anything with a queue in front of it — the tier you assumed was out of reach is already inside the budget you have.
The cost of that discount is the wait, and it is a real cost. Batch is not a cheaper API, it is a different one, and a queue you cannot drain on demand is not a place to put anything a person is waiting for.
What this list does not settle
Long context is priced three different ways, and only one of them is legible. Anthropic includes the full million-token window at standard rates and says so: a 900,000-token request is billed per token like a 9,000-token one. Google doubles Gemini 3.1 Pro above 200,000 tokens in the prompt and prints the threshold. OpenAI publishes a long-context column at roughly double the short-context one for Luna, Terra and Sol, and does not say where short stops and long starts. The cheapest row in this table is therefore cheapest at a price that turns into $0.40 and $1.80 per million at a boundary you cannot look up.
Three other things sit outside the arithmetic. Gemini 3 Flash is a preview model and preview terms change without a migration path. Pinning inference to a region costs a 1.1 times multiplier on the Claude API and a 10% uplift on OpenAI models released on or after March 5, 2026. And free tiers, rate limits and the queue you actually get on a new account are not on any of these pages.
What to choose
For classification, routing, extraction and tagging at real volume, start at GPT-5.6 Luna and only move up when you can show a failure rate that costs more than the difference. For work where the output is long enough to matter and you want a vendor whose long-context bill has no cliff in it, Claude Haiku 4.5 at $8.00 is the honest middle, and $4.00 batched. For anything you plan to still be running in February, do the sum at January prices before you commit.
Sources
Related
GPT-5.6 Sol drops to $4 and $20, and overtakes Claude Opus 5 on cost
OpenAI cut Sol's API price on August 21: input from $5 to $4, output from $30 to $20, cached input from $0.50 to $0.40. On a standard task the model goes from more expensive than Claude Opus 5 to cheaper than it. The cut is promotional and runs at least through November 21.
OpenAIverified

The best-scoring speech models reproduce the benchmark's own transcription errors
Hugging Face ran three diagnostics across 11 open speech recognition models and found that the ones with the lowest word error rates are the most likely to repeat mistakes that exist only in the reference transcript. On some tests the models appear to work out which dataset they are being scored on and switch spelling conventions accordingly. A low error rate can mean the model learned the dataset rather than the speech.
Hugging Faceverified
Gemini 3.7 Flash, read from Google's own numbers
Google shipped it on August 13, 2026, 23 days after Gemini 3.6 Flash, with large gains on coding and agent benchmarks and an introductory price it labels as such. All of that holds. Four things are visible only if you open the model card instead of the launch post, and one of them is a score that went down.
Googleverified