Skip to content
Modelsstrong signalverified

GPT-5.6 Sol drops to $4 and $20, and overtakes Claude Opus 5 on cost

OpenAI cut Sol's API price on August 21: input from $5 to $4, output from $30 to $20, cached input from $0.50 to $0.40. On a standard task the model goes from more expensive than Claude Opus 5 to cheaper than it. The cut is promotional and runs at least through November 21.

By Redakcija WebAiRadarPublished 2 min readwritten by a modelUpdated

Price changes are the part of a model's specification that moves fastest and gets checked least, so here is the whole of it. On August 21 OpenAI reduced the API price of GPT-5.6 Sol by 20% on input and 33% on output. Everything else about the model is unchanged. The reduction is explicitly a promotion with an end date attached.

The new numbers

Input falls from $5 to $4 per million tokens, cached input from $0.50 to $0.40, and output from $30 to $20.

The same promotion covers the two smaller models in the family. Terra goes from $2.50 to $2.00 on input and from $15 to $12 on output. Luna goes from $1.00 to $0.20 on input and from $6.00 to $1.20 on output, the steepest cut of the three and the one to read twice if you run high volume.

What it changes in practice

Price a single task the way a bill actually accumulates — 4,000 tokens in, 800 out, a thousand times — and Sol goes from $44.00 to $32.00, a drop of 27%. Claude Opus 5, unchanged at $5 and $25, costs $40.00 for the same thousand calls.

That is the finding worth carrying. Before August 21 the choice between the two frontier models was between $44 and $40, and Sol was the more expensive one. Today it is $32 against $40, and Sol is the cheaper one by a fifth. Nothing about either model's behavior moved; the ranking flipped on a price change alone.

The input columns crossed too. Sol and Opus 5 charged the same $5 per million on input, which made input a wash and pushed every comparison onto the output column. Now Sol is a dollar cheaper on both sides.

The date attached to it

OpenAI calls this promotional pricing and says it is available at least through November 21, 2026, describing it internally as a three-month promotion running from August 21. The phrase to notice is at least: the floor is stated, the ceiling is not.

For anything you plan to still be running in December, that makes this a rate to budget twice — once at $4 and $20, and once at whatever follows. It is the same discipline the Gemini Flash rows already demand, and the reason a price without a date is not a specification.

Where it does not apply

The reduction covers the API, and is rolling out across eligible plans for ChatGPT Work and Codex credits. OpenAI states that Pro, Plus and Business subscription usage remains unchanged, so a person paying monthly for ChatGPT sees nothing.

This is a change for whoever is billed by the token.

Sources

Corrections

  • Corrected on August 22, 2026: an earlier version described a long-context tier and a batch tier with prices that OpenAI's announcement does not contain. The announcement covers Sol, Terra and Luna, and those figures now stand in their place.

Related

Title card reading "Measuring benchmark optimization in speech recognition", with the Hume and Hugging Face logos above it.
Modelsstrong signal

The best-scoring speech models reproduce the benchmark's own transcription errors

Hugging Face ran three diagnostics across 11 open speech recognition models and found that the ones with the lowest word error rates are the most likely to repeat mistakes that exist only in the reference transcript. On some tests the models appear to work out which dataset they are being scored on and switch spelling conventions accordingly. A low error rate can mean the model learned the dataset rather than the speech.

Hugging Faceverified

GEMINI 3.7 FLASH−0.7the only score that went down
Modelsstrong signal

Gemini 3.7 Flash, read from Google's own numbers

Google shipped it on August 13, 2026, 23 days after Gemini 3.6 Flash, with large gains on coding and agent benchmarks and an introductory price it labels as such. All of that holds. Four things are visible only if you open the model card instead of the launch post, and one of them is a score that went down.

Googleverified

COST PER TASK$1.76the cheapest thousand 4,000-token calls
Modelsstrong signal

Best cheap models for high-volume work, priced per thousand calls

Six models, one task, one number: what a thousand calls cost when each sends 4,000 tokens in and gets 800 back. The cheapest row is $1.76 and the most expensive is $16.00, a nine-fold spread rather than the hundred-fold spread the category implies. Two things move the ranking more than the headline price does, and one of them has a date on it.

Anthropic, Google, OpenAIverified