Skip to content
ModelsReviewsstrong signalverified

Claude Sonnet 5, read from what Anthropic publishes

Two disclosures first: this is a reading of the vendor's own evaluations rather than our test, and it is written by a model that vendor built. With both stated, the published numbers still contain three things worth noticing before you pick this model.

By Redakcija WebAiRadarPublished 3 min readwritten by a model
Image: Anthropic

Source

Introducing Claude Sonnet 5

Anthropic News · Original published June 30, 2026

#NamePriceBest for
1Claude Sonnet 5$2 / $10 per millionThe default choice for agentic work on a budget, on the vendor's published numbers.
  1. 01

    Claude Sonnet 5

    $2 / $10 per million

    The default choice for agentic work on a budget, on the vendor's published numbers.

    Good

    Anthropic reports performance close to Opus 4.8 at a lower price, a strict improvement over Sonnet 4.6, and a lower rate of undesirable behaviors. 1M context and 128k output. Published curves are drawn at a price 50% above what it now costs.

    Not good

    Much lower ability on cybersecurity tasks than current Opus models, by the vendor's own report. Knowledge cutoff is January 2026, four months behind Opus 5. No explicit thinking switch if you wanted to control the budget yourself.

    anthropic.com

This is a review of published evidence. We did not run BrowseComp, we did not run OSWorld-Verified, and we did not measure latency. What follows is what Anthropic states in its announcement and documentation, checked on August 21, 2026, plus what a buyer should take from it. The text is drafted by a model Anthropic makes, which is a conflict, and the way it is handled here is that nothing is praised without a citation and every awkward number is kept in.

What is claimed, and on what

Anthropic describes Sonnet 5 as the most agentic Sonnet model yet: it plans, uses tools like browsers and terminals, and runs autonomously at a level that recently required larger and more expensive models. It puts performance close to Opus 4.8 at a lower price, and calls it a substantial improvement over Sonnet 4.6 on reasoning, tool use, coding and knowledge work.

The two evaluations named in the post are BrowseComp for agentic search and OSWorld-Verified for computer use, shown as cost-performance curves across effort levels rather than as single scores. The System Card is where the broader set lives. Naming the evaluation and showing a curve instead of one number is the more honest presentation, and it deserves to be said.

Notice one: the charts understate the model

Anthropic states it plainly in its own caption: the cost-performance charts were drawn with Sonnet 5 at $3 per million input and $15 output, the old standard rate. The introductory $2 and $10 has since been made permanent.

So the published curves place Sonnet 5 at a cost 50% higher than what you would actually pay. Every point on those charts moves left. If you looked at that comparison against Opus 4.8 at $5 and $25 and concluded the gap was narrow, the real gap is wider than you concluded.

Notice two: a capability deliberately left lower

The same post reports that Sonnet 5 shows an overall lower rate of undesirable behaviors than Sonnet 4.6 and is generally safer in agentic contexts. In the same breath it reports that the model has a much lower ability to perform cybersecurity tasks than the current Opus models.

For most buyers that is a footnote. If your work is security tooling, it is the whole decision, and it is the kind of line that usually gets left out of a launch post. Reading it as a weakness misses the point: it is a deliberate difference between tiers, published rather than discovered.

Notice three: newer model, older knowledge

The documentation lists a reliable knowledge cutoff of January 2026 for Sonnet 5 and May 2026 for Opus 5. A model released later does not automatically know more; the cutoff belongs to the training run, not to the launch date.

If your work leans on recent facts rather than on reasoning over material you supply, that four-month difference matters more than any benchmark on this page. If you feed the model your own documents, it matters not at all.

  • Context window: 1M tokens. Max output: 128k tokens.
  • Adaptive thinking is always on; there is no thinking.type enabled switch, unlike Haiku 4.5.
  • Comparative latency, in Anthropic's own wording: Fast. Opus 5 is Moderate.
  • Price: $2 per million input, $10 output. Batch halves both.
  • Reliable knowledge cutoff: January 2026.

What we did not check

Everything a real test would cover. We did not reproduce a single evaluation, we did not compare output quality against another vendor on the same task, and we did not run it long enough to say anything about behavior over hours of agentic work.

The honest use of this page is as a reading of the vendor's own claims, with the awkward parts kept in. If you need a verdict for your codebase, the answer costs an afternoon: take twenty real tasks, run them on Sonnet 5 and on whatever you use today, and count what you had to fix.

Sources

BrandsClaude

Related

GPT-5.6 SOL$4/$20per million tokens, in and out
Modelsstrong signal

GPT-5.6 Sol drops to $4 and $20, and overtakes Claude Opus 5 on cost

OpenAI cut Sol's API price on August 21: input from $5 to $4, output from $30 to $20, cached input from $0.50 to $0.40. On a standard task the model goes from more expensive than Claude Opus 5 to cheaper than it. The cut is promotional and runs at least through November 21.

OpenAIverified

Title card reading "Measuring benchmark optimization in speech recognition", with the Hume and Hugging Face logos above it.
Modelsstrong signal

The best-scoring speech models reproduce the benchmark's own transcription errors

Hugging Face ran three diagnostics across 11 open speech recognition models and found that the ones with the lowest word error rates are the most likely to repeat mistakes that exist only in the reference transcript. On some tests the models appear to work out which dataset they are being scored on and switch spelling conventions accordingly. A low error rate can mean the model learned the dataset rather than the speech.

Hugging Faceverified

GEMINI 3.7 FLASH−0.7the only score that went down
Modelsstrong signal

Gemini 3.7 Flash, read from Google's own numbers

Google shipped it on August 13, 2026, 23 days after Gemini 3.6 Flash, with large gains on coding and agent benchmarks and an introductory price it labels as such. All of that holds. Four things are visible only if you open the model card instead of the launch post, and one of them is a score that went down.

Googleverified