Ai2 releases AstaBrief, an open-weights model that writes cited research reports
Ai2 released the weights and training data for AstaBrief on October 2, 2026. The model takes a research question plus retrieved excerpts from the literature and writes a report with citations. Ai2 measures it at 51.1 seconds per report in its Asta platform, against 178.5 seconds for the Claude-powered mode beside it. The license is Apache 2.0, and Ai2 says its evaluation predates the current frontier models.
Source
Open-sourcing AstaBrief, the fast report-generation model in AstaHugging Face Blog · Original published October 2, 2026
Ai2 is the Allen Institute for AI, and Asta is its platform for scientific work. AstaBrief already runs there as the Fast mode of the Generate a report feature, next to a Thinking mode that Ai2 says is powered by Claude. The weights are on Hugging Face in the repository allenai/AstaBrief_8B, so you can run the model on your own hardware.
What the model does
AstaBrief does one job. You give it a research question and excerpts that a retrieval step has already found. It writes the full report in a single pass, with citations. It does not search the literature, so retrieval is a step you have to supply.
Ai2 built it on the open model Qwen3-8B, using supervised fine-tuning followed by direct preference optimization (DPO). The model card recommends the prompt format the checkpoint was trained with and warns that a different format may degrade the output. The card lists English as the model's language.
What Ai2 measured
According to Ai2's measurement across the full Asta pipeline, Fast mode averages 51.1 seconds per report and Thinking mode averages 178.5 seconds. Ai2 rounds the difference to about 3.5 times.
On quality, Ai2 says AstaBrief was competitive with the Claude-powered pipeline and with Ai2's DR Tulu model in the evaluations it used during development. The main test set was SQABench-CS2, which holds 200 computer science research questions written by users. In a separate human study, three researchers ranked reports from the three systems on 14 questions. DR Tulu came first on overall preference, and two of the three researchers preferred AstaBrief on citation accuracy.
Ai2 also reports early usage. It says 374 Asta users have tried Fast mode, and 23% of them kept using it without switching back to Thinking mode. Positive feedback came to 84.2% for Fast mode and 85.2% for Thinking mode. Ai2 adds that the feedback is too sparse for strong conclusions.
The limits Ai2 states
Most of the training and evaluation was completed in 2025. Ai2 says it has not rerun the full evaluation against the current frontier models. It describes the results as evidence about its training and system design choices, not as a claim about where the model ranks now.
The model card says the model is intended for research and educational use, in line with Ai2's Responsible Use Guidelines. The license itself is Apache 2.0.
Running it on your own infrastructure
Ai2 says open weights let an institution run AstaBrief on its own infrastructure, including behind its own firewall. It names research questions that reveal sensitive or unpublished work as the case where that is needed. Alongside the weights, Ai2 published an example workflow for generating reports from your own PDFs.
When the model page was checked on October 2, 2026, it listed no provider offering the model as a hosted service. Neither the post nor the model card states how much memory the model needs.
Sources
Related
Microsoft's ThinkingBox shows no model passes half of 507 agent tasks 20 times in a row
Microsoft's Copilot Studio team published results from its ThinkingBox benchmark on the Hugging Face blog on October 3, 2026. The benchmark runs 507 business workflows 20 times per model and grades the database state an agent leaves behind, not its final message. Microsoft reports that Claude Opus 5.5 leads single-attempt accuracy at 67.16%, yet it passes only 241 tasks on all 20 attempts. According to the authors, roughly four in five failures come from tool handling rather than reasoning.
Microsoft, Hugging Faceverified

Amazon's Strands Labs releases Decider 2B, an open decision model that runs locally
Strands Labs, the experimental arm of the Strands Agents project, released Strands Decider 2B on October 1, 2026. The model does not write text: it picks one of the options you give it, answers yes or no, or scores on a scale, and it attaches a confidence value to each answer. The team reports a median of 115 ms per decision on Nvidia's RTX 3090. Code, weights, training data, and scripts are public under Apache-2.0. If you build agents, this gives you a cheap local step for routing, tool checks, and guardrails.
Strands Agentsverified
Google announces Gemini 4 Argon for vetted cyber defenders only, at a $2 introductory input price
Google announced Gemini 4 Argon on September 30, 2026. The model is rolling out to a set of trusted cyber defenders through the Fairwind Program, and Google gives no date for developers, enterprises, or consumers. The introductory price is $2 per million input tokens and $10 per million output tokens, and it rises to $4 and $20 when the introductory period ends. The output limit grows to 1 million tokens from 64,000. The benchmark results in the announcement are Google's own reporting, and nobody outside the program can check them yet.
Googleverified


