Skip to content
Modelsmedium signalverified

Microsoft's ThinkingBox shows no model passes half of 507 agent tasks 20 times in a row

Microsoft's Copilot Studio team published results from its ThinkingBox benchmark on the Hugging Face blog on October 3, 2026. The benchmark runs 507 business workflows 20 times per model and grades the database state an agent leaves behind, not its final message. Microsoft reports that Claude Opus 5.5 leads single-attempt accuracy at 67.16%, yet it passes only 241 tasks on all 20 attempts. According to the authors, roughly four in five failures come from tool handling rather than reasoning.

By Redakcija WebAiRadarPublished 3 min readwritten by a model
Image: Microsoft, Hugging Face

ThinkingBox is a sandbox for agents, and ThinkingBox-Bench is the benchmark that runs in it. Microsoft and Hugging Face published the results and the setup instructions on October 3, 2026. Each of the 507 tasks starts from a clean backend, and executable checks compare the final records with the required end state. The authors tested 18 models, 10 proprietary and eight open-weight, and they report three numbers for each.

Three numbers per model

The first number, pass@1, is the share of all attempts that succeeded. The second, pass@20, is the share of tasks a model solved at least once in 20 tries. The third is the count of tasks that passed all 20 recorded attempts.

On pass@1, Microsoft reports Claude Opus 5.5 at 67.16%, Claude Opus 5 at 66.50%, and GPT-5.4 at 65.36%. Kimi-K3 is the strongest open-weight model at 57.37%. Scores depend heavily on the domain: across the tested models, retail tasks average 59.52% and auto insurance tasks average 33.83%.

What survives 20 repeats

According to Microsoft, Claude Opus 5.5 and Claude Opus 5 each pass 241 of the 507 tasks on every attempt, which is 47.53%. GPT-6 Astra passes 231 tasks that way. GPT-5.4 passes 128, and GPT-5.6 Sol passes 82.

Kimi-K3 shows the widest gap. The authors report that it solves 476 of 507 tasks at least once, the broadest coverage in the test, but only 68 tasks on all 20 attempts. Claude Opus 5 solves fewer tasks at least once, 79.09%, and far more of them every time.

What a dependable task costs

The authors priced each model's recorded token usage at undiscounted list rates on OpenRouter, using a snapshot from September 20, 2026. They describe the result as a comparative efficiency index, not an invoice.

By cost per successful attempt, Microsoft puts GPT-5.6 Sol lowest at $0.127. GPT-5.4 follows at $0.131 and Claude Opus 5.5 at $0.276. Claude Opus 5 costs $0.475 per successful attempt and scores lower than Opus 5.5.

The ranking changes when the cost of all 20 runs is divided by the number of tasks passed every time. On that measure the authors list GPT-5.4 at $6.80 per dependable task, GPT-6 Astra at $7.45, Claude Opus 5.5 at $7.80, GPT-5.6 Sol at $9.76, and Claude Opus 5 at $13.30. Kimi-K3 comes to $20.68.

Where agents fail

In a subset of 121,680 valid trials across 12 models, 79,853 attempts failed the executable checks. Microsoft reports that 67.24% of those failures still ended cleanly, called a state-changing tool, and returned no final tool error.

The authors assign each failed run one diagnostic label. Tool usage accounts for 79.9% of failures, wrong state updates for 10.3%, incomplete user resolutions for 7.0%, and runs with no state-changing action for 2.9%. Microsoft describes the pattern this way: an agent starts the workflow and then fails to recover from a tool error, a failed precondition, or an empty lookup.

Microsoft recommends four measures. Check the final state before you commit a change, and classify tool errors so that retries target the recoverable ones. Reduce the set of tools, and require human approval for changes that are hard to reverse. The authors say they have not measured the effect of any of these measures on the benchmark.

How to run it

The harness and the dataset are on Hugging Face, and the benchmark runs behind the OpenEnv interface. The authors tested the setup on Linux and WSL with Python 3.11 or later, uv, and Docker. You also need model endpoints for the agent, the simulated user, and the judge, and one endpoint can serve all three roles.

ThinkingBox code is MIT-licensed, the benchmark data uses CDLA-Permissive-2.0, and the OpenEnv environment ships under BSD-3-Clause. Every task is a synthetic reconstruction, and 477 of the 507 tasks are graded on state alone. The results come from the benchmark's own authors, and the post does not cite an independent reproduction. The paper that describes the benchmark is an arXiv preprint and has not been peer reviewed.

„A trajectory is a claim. Database state is the evidence.“
Microsoft and Hugging Face, in the ThinkingBox announcement

Sources

Related

Ai2's graphic for AstaBrief: the model's name in green letters on a dark background scattered with outlined document icons, a few of them highlighted.
Modelsweak signal

Ai2 releases AstaBrief, an open-weights model that writes cited research reports

Ai2 released the weights and training data for AstaBrief on October 2, 2026. The model takes a research question plus retrieved excerpts from the literature and writes a report with citations. Ai2 measures it at 51.1 seconds per report in its Asta platform, against 178.5 seconds for the Claude-powered mode beside it. The license is Apache 2.0, and Ai2 says its evaluation predates the current frontier models.

Ai2verified

Architecture diagram of Strands Decider 2B: on the left, a decoder language model whose text-generating head is discarded; on the right, the same Qwen3.5-2B-Base torso with a LoRA adapter feeds a pointer head that turns the listed options into per-option probabilities.
Modelsmedium signal

Amazon's Strands Labs releases Decider 2B, an open decision model that runs locally

Strands Labs, the experimental arm of the Strands Agents project, released Strands Decider 2B on October 1, 2026. The model does not write text: it picks one of the options you give it, answers yes or no, or scores on a scale, and it attaches a confidence value to each answer. The team reports a median of 115 ms per decision on Nvidia's RTX 3090. Code, weights, training data, and scripts are public under Apache-2.0. If you build agents, this gives you a cheap local step for routing, tool checks, and guardrails.

Strands Agentsverified

Modelsstrong signal

Google announces Gemini 4 Argon for vetted cyber defenders only, at a $2 introductory input price

Google announced Gemini 4 Argon on September 30, 2026. The model is rolling out to a set of trusted cyber defenders through the Fairwind Program, and Google gives no date for developers, enterprises, or consumers. The introductory price is $2 per million input tokens and $10 per million output tokens, and it rises to $4 and $20 when the introductory period ends. The output limit grows to 1 million tokens from 64,000. The benchmark results in the announcement are Google's own reporting, and nobody outside the program can check them yet.

Googleverified