Skip to content

Tags

benchmark

7 items
Title card reading "Measuring benchmark optimization in speech recognition", with the Hume and Hugging Face logos above it.
Modelsstrong signal

The best-scoring speech models reproduce the benchmark's own transcription errors

Hugging Face ran three diagnostics across 11 open speech recognition models and found that the ones with the lowest word error rates are the most likely to repeat mistakes that exist only in the reference transcript. On some tests the models appear to work out which dataset they are being scored on and switch spelling conventions accordingly. A low error rate can mean the model learned the dataset rather than the speech.

Hugging Faceverified

A bobblehead figure of Nvidia's chief executive holding a game screen, above a green ARC-AGI-3 progress bar filled to 100%.
Agentsstrong signal

Nvidia's harness takes Claude Opus 5 from about 30% to a perfect ARC-AGI-3 score

Nvidia published a run in which AVO, its agent architecture, scores 100.00 RHAE on ARC-AGI-3 and clears all 183 levels across 25 environments. The same model evaluated on its own scores about 30%. In the same post Nvidia writes that the comparison is not a controlled ablation, and that sentence did not survive into the coverage.

NVIDIAverified

GEMINI 3.7 FLASH−0.7the only score that went down
Modelsstrong signal

Gemini 3.7 Flash, read from Google's own numbers

Google shipped it on August 13, 2026, 23 days after Gemini 3.6 Flash, with large gains on coding and agent benchmarks and an introductory price it labels as such. All of that holds. Four things are visible only if you open the model card instead of the launch post, and one of them is a score that went down.

Googleverified

Modelsstrong signal

Best cheap models for high-volume work, priced per thousand calls

Six models, one task, one number: what a thousand calls cost when each sends 4,000 tokens in and gets 800 back. The cheapest row is $1.76 and the most expensive is $16.00, a nine-fold spread rather than the hundred-fold spread the category implies. Two things move the ranking more than the headline price does, and one of them has a date on it.

Anthropic, Google, OpenAIverified

Modelsstrong signal

Claude Sonnet 5, read from what Anthropic publishes

Two disclosures first: this is a reading of the vendor's own evaluations rather than our test, and it is written by a model that vendor built. With both stated, the published numbers still contain three things worth noticing before you pick this model.

Anthropicverified

Modelsstrong signal

Gemini 3.7 Flash: half the price and markedly better on code

Google released Gemini 3.7 Flash on August 13, just three weeks after 3.6. The gain on coding benchmarks is large: the DeepSWE v1.1 score rises from 49.0% to 65.3%, and FrontierCode 1.1 from 34.4% to 43.6%. Introductory pricing, in force through the end of 2026, is $0.75 per million input tokens and $3.75 per million output tokens — half what its predecessor launched at. The model is available through the Gemini API, in AI Studio, Android Studio, and Antigravity, and in Spark for AI Pro and Ultra subscribers.

9to5Google / Googleverified