Skip to content
Modelsmedium signalverified

A preprint finds six commercial LLM routers no better than a random pick between two models

A preprint submitted to arXiv on October 2, 2026, tested six commercial LLM routers in 14 settings. According to the authors, none of them beat a router that picks at random between Gemini 3.7 Flash and Opus 5 at the same cost, and one trailed it by 10.5 percentage points. The authors, who work at Fastino Labs, trace the gap to the way routers are evaluated and to rosters that hold too many models. The paper has not been peer reviewed.

By Redakcija WebAiRadarPublished 4 min readwritten by a model

Source

Dynamic LLM Routers are Often Misguided

arXiv · Original published October 2, 2026

LLM routers send each query to one of several models, so that a cheap model handles what it can and an expensive one handles the rest. The paper, titled Dynamic LLM Routers are Often Misguided, checks whether commercial products deliver on that aim. Its six authors compared each product with the simplest possible baseline and report that the baseline was at least as accurate.

How the test was run

The authors tested routers from three general platforms, OpenRouter, Microsoft Azure, and vLLM, and from three routing specialists, Not Diamond, Nadir, and Orca. The products were tested in 14 settings in total, such as a cost-oriented and a quality-oriented mode of the same product.

The evaluation set holds 800 queries from 17 benchmarks, in eight categories of 100 queries each. The categories are coding, instruction-following, knowledge, math, question answering, office work, tool use, and Humanity's Last Exam. For every query, the authors recorded which model a router selected. They then generated and graded that model's answer themselves, with default OpenRouter settings and prices, so that every router was measured the same way.

The baseline is a router that ignores the content of the query. It picks Gemini 3.7 Flash with a fixed probability and Anthropic's Opus 5 otherwise, in the configuration the paper labels Opus 5 (high). The probability is set so that the baseline costs the same as the router under test.

What the comparison shows

None of the 14 settings was more accurate than the random baseline at matched cost, the authors report. Orca's adaptive setting trailed it by 10.5 percentage points. Not Diamond's cost setting and one setting of OpenRouter's auto-beta router trailed by 8.8 points each, and Azure's balanced setting by 8.4 points.

The smallest gap belongs to vLLM Semantic Router, at 1.0 point with a confidence interval of 1.8 points either way. It is the only product whose result is not significantly below the baseline, and the authors note that it was fit on their own training data.

The authors also repeated the comparison with a baseline limited to the models each router had access to. They report that no commercial router outperformed it in that case either, so newer models in the baseline do not explain the result.

Four patterns behind the gap

The paper attributes the gap to four patterns and says every router it tested shows several of them.

  • Difficulty blindness: the choice of model is at best weakly related to how hard a query is, so the hardest queries reach the stronger model less often than they should.
  • Length reversal: some routers spend less on queries with longer answers, although such queries tend to be harder.
  • Semantic matching: routers send most queries from the same source to one model, regardless of how hard each query is.
  • Roster suboptimality: rosters include many models that are not on the cost-accuracy frontier, or leave out models that are.

Why two models are enough

According to the authors, the standard way of evaluating routers rewards the first three patterns. Router benchmarks weigh cost against accuracy. Under that measure, escalating a query pays most when the expensive model is much more likely to succeed and the answer is short. On the hardest queries both models usually fail, so escalation gains little.

On rosters, the paper reports that Gemini 3.7 Flash trails Opus 5 (high) by 1.6 percentage points in a set of 26 tested models, at a 96% discount. Even a router with perfect knowledge of query difficulty would gain at most 1 percentage point from a roster larger than two models.

The authors also built their own two-model router, which avoids all four patterns. Under the standard methodology its accuracy stayed within noise of random routing at every budget. Their conclusion is that the roster, not the router, determines nearly all of the accuracy.

What the paper does not show

The authors list the limits themselves. The test covers single-turn queries only, and switching models in the middle of a conversation has a cost of its own because of caching. All queries are in English and come from benchmarks, which may not resemble real traffic, and each answer was generated once.

The paper has not been peer reviewed, and the authors state that it is under review at NAACL. It measures accuracy at matched cost and no other reason to use a router, and the results are the authors' own measurement.

„The roster, not the router, determines nearly all of the accuracy.“
Wang and co-authors, Dynamic LLM Routers are Often Misguided

Sources

Related

Screenshot of the OpenAI API dashboard, Project Settings, Text provenance tab: the Allow text watermarking switch is on, the Models menu reads All selected, and a Save button sits below.
Modelsstrong signal

OpenAI will watermark ChatGPT and Codex text in the EU, and API customers can opt in

OpenAI published its plan for text watermarking on October 5, 2026, in response to the EU AI Act. Over the coming weeks, eligible ChatGPT and Codex text output in the European Union will carry an invisible watermark called textGrain. In the API, watermarking is available worldwide from the same day for select models, and it stays off unless you turn it on. The detector is not public: OpenAI is limiting it to approved researchers and expert organizations.

OpenAIverified

Microsoft's graphic for ThinkingBox: the name spelled in dotted letters on a yellow-green background covered with small blocks of horizontal lines.
Modelsmedium signal

Microsoft's ThinkingBox shows no model passes half of 507 agent tasks 20 times in a row

Microsoft's Copilot Studio team published results from its ThinkingBox benchmark on the Hugging Face blog on October 3, 2026. The benchmark runs 507 business workflows 20 times per model and grades the database state an agent leaves behind, not its final message. Microsoft reports that Claude Opus 5.5 leads single-attempt accuracy at 67.16%, yet it passes only 241 tasks on all 20 attempts. According to the authors, roughly four in five failures come from tool handling rather than reasoning.

Microsoft, Hugging Faceverified

Ai2's graphic for AstaBrief: the model's name in green letters on a dark background scattered with outlined document icons, a few of them highlighted.
Modelsweak signal

Ai2 releases AstaBrief, an open-weights model that writes cited research reports

Ai2 released the weights and training data for AstaBrief on October 2, 2026. The model takes a research question plus retrieved excerpts from the literature and writes a report with citations. Ai2 measures it at 51.1 seconds per report in its Asta platform, against 178.5 seconds for the Claude-powered mode beside it. The license is Apache 2.0, and Ai2 says its evaluation predates the current frontier models.

Ai2verified