Claude Opus 5: choosing how much effort to spend, at unchanged prices
The new flagship in the Opus line lets you set how much effort it spends on a task, across three levels — low, medium, and high — which gives you direct control over the trade-off between cost and quality. Pricing stayed where it was with 4.8, at five dollars per million input tokens and twenty-five per million output tokens, while scores on a large share of tasks sit close to Fable. Opus 5 became the default on the Max plan and the strongest model available on Pro, and it is called in the API as claude-opus-5.
Opus 5 was released on July 24 as the new flagship in its line. The most interesting change is not a benchmark score but the control you get, because you choose how much effort the model spends on a task.
Choosing the effort
There are three levels: low, medium, and high. It is the same model with a different amount of thinking per task. In practice that means you run routine work cheaply and hard decisions expensively, without swapping models and without two separate paths in your integration.
Price and availability
Pricing stayed where it was with 4.8: five dollars per million input tokens and twenty-five per million output tokens. In the API the model is called as claude-opus-5.
It became the default on the Max plan and the strongest model available on Pro.
What else it brings
TechCrunch reports a new automatic fallback to a weaker model, for now in testing: when safety classifiers fire, the request is rerouted instead of returning an error. In production that means fewer hard interruptions.
The same source reports that classifiers fire noticeably less often on Opus 5 than on Fable, and that the 30-day data retention applied to Fable and Mythos does not apply to Opus 5.
On the safety side, searching for vulnerabilities in compiled binaries is disabled, while source code analysis is allowed.
Sources
Related
GPT-5.6 Sol drops to $4 and $20, and overtakes Claude Opus 5 on cost
OpenAI cut Sol's API price on August 21: input from $5 to $4, output from $30 to $20, cached input from $0.50 to $0.40. On a standard task the model goes from more expensive than Claude Opus 5 to cheaper than it. The cut is promotional and runs at least through November 21.
OpenAIverified

The best-scoring speech models reproduce the benchmark's own transcription errors
Hugging Face ran three diagnostics across 11 open speech recognition models and found that the ones with the lowest word error rates are the most likely to repeat mistakes that exist only in the reference transcript. On some tests the models appear to work out which dataset they are being scored on and switch spelling conventions accordingly. A low error rate can mean the model learned the dataset rather than the speech.
Hugging Faceverified
Gemini 3.7 Flash, read from Google's own numbers
Google shipped it on August 13, 2026, 23 days after Gemini 3.6 Flash, with large gains on coding and agent benchmarks and an introductory price it labels as such. All of that holds. Four things are visible only if you open the model card instead of the launch post, and one of them is a score that went down.
Googleverified
