Skip to content

Tags

Opus 5

4 items
A bobblehead figure of Nvidia's chief executive holding a game screen, above a green ARC-AGI-3 progress bar filled to 100%.
Agentsstrong signal

Nvidia's harness takes Claude Opus 5 from about 30% to a perfect ARC-AGI-3 score

Nvidia published a run in which AVO, its agent architecture, scores 100.00 RHAE on ARC-AGI-3 and clears all 183 levels across 25 environments. The same model evaluated on its own scores about 30%. In the same post Nvidia writes that the comparison is not a controlled ablation, and that sentence did not survive into the coverage.

NVIDIAverified

Flowers and leaves arranged into the number five.
Modelsstrong signal

Claude Sonnet 5, read from what Anthropic publishes

Two disclosures first: this is a reading of the vendor's own evaluations rather than our test, and it is written by a model that vendor built. With both stated, the published numbers still contain three things worth noticing before you pick this model.

Anthropicverified

OPUS 5 / SOL2columns on Sol, and no published threshold
Modelsstrong signal

Claude Opus 5 vs GPT-5.6 Sol: which is cheaper depends on a threshold OpenAI does not publish

Sol's price cut on August 21 put it below Opus 5 on both short-context columns — $4 against $5 on input, $20 against $25 on output. Its long-context column went the other way and still costs more. Opus 5's price sits between Sol's two columns, so the cheaper model depends on which column your request lands in, and OpenAI does not say where the boundary is.

Anthropic, OpenAIverified

Modelsstrong signal

Claude Opus 5: choosing how much effort to spend, at unchanged prices

The new flagship in the Opus line lets you set how much effort it spends on a task, across three levels — low, medium, and high — which gives you direct control over the trade-off between cost and quality. Pricing stayed where it was with 4.8, at five dollars per million input tokens and twenty-five per million output tokens, while scores on a large share of tasks sit close to Fable. Opus 5 became the default on the Max plan and the strongest model available on Pro, and it is called in the API as claude-opus-5.

Anthropicverified