Skip to content
ModelsComparisonsstrong signalverified

Claude Opus 5 vs GPT-5.6 Sol: which is cheaper depends on a threshold OpenAI does not publish

Sol's price cut on August 21 put it below Opus 5 on both short-context columns — $4 against $5 on input, $20 against $25 on output. Its long-context column went the other way and still costs more. Opus 5's price sits between Sol's two columns, so the cheaper model depends on which column your request lands in, and OpenAI does not say where the boundary is.

By Redakcija WebAiRadarPublished 4 min readwritten by a modelUpdated

Source

Claude pricing

Claude dokumentacija · Original published August 22, 2026

#NamePriceBest for
1Claude Opus 5$5 / $25 per million, across the full million-token windowThe price you can compute in advance. Cheaper the moment a request is long, and never surprising.
2GPT-5.6 Sol$4 / $20 short context, $8 / $30 long context, promotional through at least November 21, 2026Cheaper on short requests by a clear margin, and more expensive on long ones by a wider one. Which you get is not documented.
  1. 01

    Claude Opus 5

    $5 / $25 per million, across the full million-token window

    The price you can compute in advance. Cheaper the moment a request is long, and never surprising.

    Good

    One rate for the entire context window, stated explicitly, so a large request carries no surcharge and the bill is predictable from the token count alone. Batch at $2.50 and $12.50; cache reads at $0.50. No promotional date attached to the price.

    Not good

    More expensive than Sol on every short-context column: a dollar more on input, five more on output. The newer tokenizer counts about 30% more tokens for the same text than earlier Claude models, so cost comparisons against those are not like for like.

    platform.claude.com

  2. 02

    GPT-5.6 Sol

    $4 / $20 short context, $8 / $30 long context, promotional through at least November 21, 2026

    Cheaper on short requests by a clear margin, and more expensive on long ones by a wider one. Which you get is not documented.

    Good

    Twenty per cent below Opus 5 on input and output alike within short context, with cache reads at $0.40 and batch at $2 and $10. The reduction also reaches ChatGPT Work and Codex credits.

    Not good

    The long-context column costs 52% more than Opus 5 on a large job, and the page does not say at how many tokens that column starts. The price is promotional with a stated floor of November 21, 2026 and no stated ceiling.

    developers.openai.com

This page was rewritten on August 22 because its premise stopped being true. Until August 21 both models charged five dollars per million input tokens, and every comparison came down to the output column. OpenAI then cut Sol by 20% on input and 33% on output, and the shape of the question changed with it. What follows is the comparison at today's published terms.

Short context: Sol is now cheaper on both columns

Sol charges $4 per million input tokens against Opus 5's $5, and $20 output against $25. Cache reads are $0.40 against $0.50. Batch halves both sides at both vendors: $2 and $10 on Sol, $2.50 and $12.50 on Opus 5.

Every one of those columns favors Sol, so within short context there is no crossover point to find and no mix of input to output that changes the answer. Price a job of 50,000 input and 5,000 output tokens and it costs $0.30 on Sol against $0.375 on Opus 5 — a fifth less, consistently.

Long context: Opus 5 is cheaper, and it is not close

Anthropic includes the full million-token window at standard rates and states it plainly: a 900,000-token request is billed per token exactly like a 9,000-token one. Caching and batch discounts apply across the whole window at the same rates.

OpenAI runs a second column. Sol's long-context rates are $8 input, $0.80 cached and $30 output — double the input, half again on the output. Push the same large job through both and the gap is wide: at 400,000 input and 20,000 output tokens, Opus 5 costs $2.50 and Sol costs $3.80, which is 52% more.

Here too there is no crossover. Both of Sol's long-context coefficients are higher than Opus 5's, so once you are in that column Opus 5 wins on every mix.

Opus 5 sits between Sol's two columns, and that is the whole comparison

Take one job — 50,000 tokens in, 5,000 out — and price it three ways. On Sol's short-context column it costs $0.30. On Opus 5 it costs $0.375. On Sol's long-context column it costs $0.55. Opus 5 is 25% more expensive than one Sol and 32% cheaper than the other, for the identical request.

So the question is not which model is cheaper. It is which of Sol's two columns your traffic will be billed under, and that is the one number OpenAI's pricing page does not carry: it publishes both columns and no threshold. You can find the answer on your own invoice at the end of the month, which is a strange place to discover the price of something.

Anthropic has no such column, and that is the actual difference between these two price lists. Not the rates — those move — but the fact that one of them can be computed in advance and the other cannot.

One price has an expiry date, the other does not

The $4 and $20 are promotional. OpenAI announced them on August 21 as a reduction of over 20% for the next three months, and the page says the promotional pricing is available at least through November 21, 2026. The floor is stated; the ceiling is not.

Opus 5's $5 and $25 carry no date. That does not make them permanent, but it does mean the two figures you are comparing are not the same kind of number: one is a standard rate and the other is an offer with a stated end. If your decision commits you past November, run the sum twice — once at $4 and $20, and once at $5 and $30, which is what Sol cost until last week.

At the old rates, on 1,000 calls of 4,000 input and 800 output tokens, Sol cost $44 and Opus 5 $40. Today the same thousand calls are $32 on Sol and still $40 on Opus 5. Nothing about either model changed.

Multipliers on both sides

Anthropic sells fast mode for Opus 5 at $10 and $50 per million, double the standard rate, and it stacks on top of caching and inference geography. Pinning inference to the United States multiplies every token category by 1.1 on Claude 4.6 and later.

OpenAI charges 10% more for regional processing on models released on or after March 5, 2026. Both vendors will charge you extra for speed or for where the computer physically sits, and neither charges it until you ask.

One thing no price list can tell you

Claude 4.7 and later use a newer tokenizer, which Anthropic says produces about 30% more tokens for the same text than the previous one. Opus 5 is on that tokenizer.

That figure describes Anthropic against Anthropic. It says nothing about how either model counts against OpenAI's tokenizer, and anyone claiming otherwise is calculating where they should be measuring. If the choice is close on price, count tokens for both on a real sample of your own text. It is an afternoon's work and the only number that describes your bill.

Sources

Corrections

  • Rewritten in full. The first version was built on both models charging $5 per million input tokens; OpenAI cut Sol to $4 input and $20 output on August 21, 2026, which reverses the conclusion on short requests. The title changed with it. One error from the first version is also corrected here: it stated that a cache write costs $6.25 per million at both vendors. That is Anthropic's rate. OpenAI publishes no cache write price at all — only a cached input rate.

Related

GPT-5.6 SOL$4/$20per million tokens, in and out
Modelsstrong signal

GPT-5.6 Sol drops to $4 and $20, and overtakes Claude Opus 5 on cost

OpenAI cut Sol's API price on August 21: input from $5 to $4, output from $30 to $20, cached input from $0.50 to $0.40. On a standard task the model goes from more expensive than Claude Opus 5 to cheaper than it. The cut is promotional and runs at least through November 21.

OpenAIverified

Title card reading "Measuring benchmark optimization in speech recognition", with the Hume and Hugging Face logos above it.
Modelsstrong signal

The best-scoring speech models reproduce the benchmark's own transcription errors

Hugging Face ran three diagnostics across 11 open speech recognition models and found that the ones with the lowest word error rates are the most likely to repeat mistakes that exist only in the reference transcript. On some tests the models appear to work out which dataset they are being scored on and switch spelling conventions accordingly. A low error rate can mean the model learned the dataset rather than the speech.

Hugging Faceverified

GEMINI 3.7 FLASH−0.7the only score that went down
Modelsstrong signal

Gemini 3.7 Flash, read from Google's own numbers

Google shipped it on August 13, 2026, 23 days after Gemini 3.6 Flash, with large gains on coding and agent benchmarks and an introductory price it labels as such. All of that holds. Four things are visible only if you open the model card instead of the launch post, and one of them is a score that went down.

Googleverified