Amazon's Strands Labs releases Decider 2B, an open decision model that runs locally
Strands Labs, the experimental arm of the Strands Agents project, released Strands Decider 2B on October 1, 2026. The model does not write text: it picks one of the options you give it, answers yes or no, or scores on a scale, and it attaches a confidence value to each answer. The team reports a median of 115 ms per decision on Nvidia's RTX 3090. Code, weights, training data, and scripts are public under Apache-2.0. If you build agents, this gives you a cheap local step for routing, tool checks, and guardrails.
Source
Amazon releases its own Jev clone as decision models flood the webTechCrunch AI · Original published October 1, 2026
Decision models are a class that has drawn attention since TypeSafe AI launched Jev in September 2026. Strands Decider 2B is the Strands Agents project's entry in that class. TechCrunch reports it as a release from Amazon Web Services. You can run it, and retrain it, on hardware you already own.
What a decision model does
A large language model generates text. A decision model cannot. Strands Decider reads a piece of text, which the project calls the state, and answers questions about it in one of three forms:
- Choice: the model picks one of the options listed in the request, such as which team should handle a support ticket.
- Yes or no: the model returns a value between 0 and 1, where a value closer to 1 leans toward yes.
- Score: the model places the text on an ordered scale that you define.
How it is built and what that rules out
The model has 1.9 billion parameters. The team started from Qwen3.5-2B-Base, removed the part that produces text, and replaced it with a small pointer head of about a million parameters. That head scores every option in a single pass through the model, with no generation loop. The base model is adapted with a rank-16 LoRA adapter.
The authors state the cost of that design in the announcement. Because all answers come out of one parallel pass, the model is significantly worse at complex problems than reasoning models. It cannot write code, hold a conversation, or summarize documents.
What the team measured
All figures below are the project's own measurements, published in the repository. JevBench, the benchmark behind them, is maintained by a third party.
- Accuracy on the JevBench v1 public set: 0.723, or 167 of 231 tasks.
- Calibration on the same set: a Brier score of 0.342 and an expected calibration error of 0.052.
- By task difficulty: 1.000 on easy tasks, 0.875 on standard tasks, and 0.505 on hard tasks.
- Latency on Nvidia's RTX 3090 under WSL2: a median of 115 ms per question and a 95th percentile of 299 ms.
- Latency on Apple's M3 Pro: a median of 153 ms for tasks under 300 tokens and 234 ms across all tasks, measured with the model already loaded.
- Confidence: on short classification tasks the model has not seen, answers with a confidence of 0.9 or more are right about 95% of the time.
What the numbers do not show
The announcement says the model ranks third of 33 models in its class on the JevBench board, and first of 30 when models slightly over 2 billion parameters are left out. The repository adds a caution of its own. The public set has only 231 tasks, and six repeated training runs of an earlier recipe produced a standard deviation of 3.2 tasks. The team therefore says a difference of fewer than about 10 tasks between two single runs should be treated as unresolved.
How to run it
The package installs with pip install strands-decider. The command strands-decider ask downloads the model StrandsAgents/strands-decider-2B-hobson-v19 from Hugging Face and answers questions about the text passed with --state. You can ask several questions in one call, which is cheaper because the text is read only once.
The command strands-decider serve starts a local server that accepts JSON requests at /v1/systemone. The repository also includes a worked example inside a Strands agent. In it, the model checks a tool call before it runs, so the agent asks the user for a city that was never named instead of guessing one.
Code, weights, training data, and training scripts are public under the Apache-2.0 license. The team says the full training recipe takes about 11 hours on one RTX 3090.
„One forward pass, no generation, no decoding loop.“
Sources
Related
Microsoft's ThinkingBox shows no model passes half of 507 agent tasks 20 times in a row
Microsoft's Copilot Studio team published results from its ThinkingBox benchmark on the Hugging Face blog on October 3, 2026. The benchmark runs 507 business workflows 20 times per model and grades the database state an agent leaves behind, not its final message. Microsoft reports that Claude Opus 5.5 leads single-attempt accuracy at 67.16%, yet it passes only 241 tasks on all 20 attempts. According to the authors, roughly four in five failures come from tool handling rather than reasoning.
Microsoft, Hugging Faceverified

Ai2 releases AstaBrief, an open-weights model that writes cited research reports
Ai2 released the weights and training data for AstaBrief on October 2, 2026. The model takes a research question plus retrieved excerpts from the literature and writes a report with citations. Ai2 measures it at 51.1 seconds per report in its Asta platform, against 178.5 seconds for the Claude-powered mode beside it. The license is Apache 2.0, and Ai2 says its evaluation predates the current frontier models.
Ai2verified
Google announces Gemini 4 Argon for vetted cyber defenders only, at a $2 introductory input price
Google announced Gemini 4 Argon on September 30, 2026. The model is rolling out to a set of trusted cyber defenders through the Fairwind Program, and Google gives no date for developers, enterprises, or consumers. The introductory price is $2 per million input tokens and $10 per million output tokens, and it rises to $4 and $20 when the introductory period ends. The output limit grows to 1 million tokens from 64,000. The benchmark results in the announcement are Google's own reporting, and nobody outside the program can check them yet.
Googleverified


