A preprint finds six commercial LLM routers no better than a random pick between two models
A preprint submitted to arXiv on October 2, 2026, tested six commercial LLM routers in 14 settings. According to the authors, none of them beat a router that picks at random between Gemini 3.7 Flash and Opus 5 at the same cost, and one trailed it by 10.5 percentage points. The authors, who work at Fastino Labs, trace the gap to the way routers are evaluated and to rosters that hold too many models. The paper has not been peer reviewed.
LLM routers send each query to one of several models, so that a cheap model handles what it can and an expensive one handles the rest. The paper, titled Dynamic LLM Routers are Often Misguided, checks whether commercial products deliver on that aim. Its six authors compared each product with the simplest possible baseline and report that the baseline was at least as accurate.
How the test was run
The authors tested routers from three general platforms, OpenRouter, Microsoft Azure, and vLLM, and from three routing specialists, Not Diamond, Nadir, and Orca. The products were tested in 14 settings in total, such as a cost-oriented and a quality-oriented mode of the same product.
The evaluation set holds 800 queries from 17 benchmarks, in eight categories of 100 queries each. The categories are coding, instruction-following, knowledge, math, question answering, office work, tool use, and Humanity's Last Exam. For every query, the authors recorded which model a router selected. They then generated and graded that model's answer themselves, with default OpenRouter settings and prices, so that every router was measured the same way.
The baseline is a router that ignores the content of the query. It picks Gemini 3.7 Flash with a fixed probability and Anthropic's Opus 5 otherwise, in the configuration the paper labels Opus 5 (high). The probability is set so that the baseline costs the same as the router under test.
What the comparison shows
None of the 14 settings was more accurate than the random baseline at matched cost, the authors report. Orca's adaptive setting trailed it by 10.5 percentage points. Not Diamond's cost setting and one setting of OpenRouter's auto-beta router trailed by 8.8 points each, and Azure's balanced setting by 8.4 points.
The smallest gap belongs to vLLM Semantic Router, at 1.0 point with a confidence interval of 1.8 points either way. It is the only product whose result is not significantly below the baseline, and the authors note that it was fit on their own training data.
The authors also repeated the comparison with a baseline limited to the models each router had access to. They report that no commercial router outperformed it in that case either, so newer models in the baseline do not explain the result.
Four patterns behind the gap
The paper attributes the gap to four patterns and says every router it tested shows several of them.
- Difficulty blindness: the choice of model is at best weakly related to how hard a query is, so the hardest queries reach the stronger model less often than they should.
- Length reversal: some routers spend less on queries with longer answers, although such queries tend to be harder.
- Semantic matching: routers send most queries from the same source to one model, regardless of how hard each query is.
- Roster suboptimality: rosters include many models that are not on the cost-accuracy frontier, or leave out models that are.
Why two models are enough
According to the authors, the standard way of evaluating routers rewards the first three patterns. Router benchmarks weigh cost against accuracy. Under that measure, escalating a query pays most when the expensive model is much more likely to succeed and the answer is short. On the hardest queries both models usually fail, so escalation gains little.
On rosters, the paper reports that Gemini 3.7 Flash trails Opus 5 (high) by 1.6 percentage points in a set of 26 tested models, at a 96% discount. Even a router with perfect knowledge of query difficulty would gain at most 1 percentage point from a roster larger than two models.
The authors also built their own two-model router, which avoids all four patterns. Under the standard methodology its accuracy stayed within noise of random routing at every budget. Their conclusion is that the roster, not the router, determines nearly all of the accuracy.
What the paper does not show
The authors list the limits themselves. The test covers single-turn queries only, and switching models in the middle of a conversation has a cost of its own because of caching. All queries are in English and come from benchmarks, which may not resemble real traffic, and each answer was generated once.
The paper has not been peer reviewed, and the authors state that it is under review at NAACL. It measures accuracy at matched cost and no other reason to use a router, and the results are the authors' own measurement.
„The roster, not the router, determines nearly all of the accuracy.“
Sources
Related

OpenAI will watermark ChatGPT and Codex text in the EU, and API customers can opt in
OpenAI published its plan for text watermarking on October 5, 2026, in response to the EU AI Act. Over the coming weeks, eligible ChatGPT and Codex text output in the European Union will carry an invisible watermark called textGrain. In the API, watermarking is available worldwide from the same day for select models, and it stays off unless you turn it on. The detector is not public: OpenAI is limiting it to approved researchers and expert organizations.
OpenAIverified

Microsoft's ThinkingBox shows no model passes half of 507 agent tasks 20 times in a row
Microsoft's Copilot Studio team published results from its ThinkingBox benchmark on the Hugging Face blog on October 3, 2026. The benchmark runs 507 business workflows 20 times per model and grades the database state an agent leaves behind, not its final message. Microsoft reports that Claude Opus 5.5 leads single-attempt accuracy at 67.16%, yet it passes only 241 tasks on all 20 attempts. According to the authors, roughly four in five failures come from tool handling rather than reasoning.
Microsoft, Hugging Faceverified

Ai2 releases AstaBrief, an open-weights model that writes cited research reports
Ai2 released the weights and training data for AstaBrief on October 2, 2026. The model takes a research question plus retrieved excerpts from the literature and writes a report with citations. Ai2 measures it at 51.1 seconds per report in its Asta platform, against 178.5 seconds for the Claude-powered mode beside it. The license is Apache 2.0, and Ai2 says its evaluation predates the current frontier models.
Ai2verified