AI

你的模型路由器,输给了抛硬币(中英切换)

你花钱接的模型路由器,大概率输给了一枚硬币。

路由器卖的故事很顺:简单题给便宜模型,难题升级给贵模型,账单砍一半,分数不掉。问题是,这个故事很少跟最笨的对照组比——同样的钱,只在两个选对的模型之间随机抛硬币。比了之后,故事就不太好讲了。

14 个商业路由配置,没有一个跑赢抛硬币

14 个配置,0 个赢了硬币

Fastino Labs《Dynamic LLM Routers are Often Misguided》(arXiv:2610.02762) 把 6 家商业路由(OpenRouter、Microsoft Azure、vLLM、Not Diamond、Nadir、Orca)拉出来,跑了 14 个配置,用 26 个模型、17 个基准、8 类任务、800 道评测题,所有成本和准确率按同一套口径重算。

对照组只有一句话:在 Gemini 3.7 Flash 和 Opus 5 (high) 之间按概率随机选,调到和路由器同一个花费。结果 14 个配置的差值全是负数;除了用论文训练集拟合过的 vLLM Semantic Router(−1.0 pp,噪声内),其余都显著落后。最差的 Orca (adaptive):从 187 个模型的名单里实际挑了 96 个,准确率 71.0%,比同价硬币低 10.5 个百分点。Not Diamond (cost) 和 OpenRouter auto-β (high) 都低 8.8。

就算把硬币的两面换成各家路由当时能用的模型(GPT 5.6 Luna / Sol),也没有一家显著跑赢。不是对照组作弊用了新模型。

路由器在看什么?不是难度

选对两个模型,比聪明的路由值钱

论文把病因拆成四条,最扎心的是第一条:没有一个商业路由的选择和题目难度的相关超过 γ = +0.20。它们更像是在认「这题来自哪个数据集」,然后把整个来源丢给同一个模型——同一来源里难题简单题一视同仁。

更有意思的是,这不全是工程师偷懒,是考试本身在奖励错的东西。业界惯用的「成本—准确率帕累托」按实际花费算账:最难的题两个模型都做不出,升级也白搭,所以最划算的是升级中等难度题——60–80 分位那一段比最难的 80–100 段高 6.3 个百分点。更离谱的是「挑短题」:在 gpt-oss-20b → Opus 5 这一对上,只升级答案最短的 20%,能做到 59.5%、$1.03/千次,比升级最难的 20% 只低 1.6 个百分点,花费却只有 1.5%。作者管这叫 accuracy hacking——在聪明模型最便宜的时候才叫它。

真正值钱的是名单

再看另一头。Gemini 3.7 Flash 只比 Opus 5 (high) 低 1.6 个百分点,价格便宜 96%。 就算路由器开天眼、事先知道每道题的真实难度、还在评测集上挑名单和阈值,两个模型的名单离任意大小的最优名单也不超过 1 pp。模型「各有所长」这个让名单越拉越长的假设,实测也很弱:给每个模型加 8 维专长参数,拟合误差(NLL)只从 0.27 降到 0.24。

作者最后自己做了一个避开四个毛病的两模型路由——结果也只和硬币打平(−0.2 到 −0.7 pp,噪声内)。原因一句话:名单选对了,就没剩多少可路由的了。

所以争议点在这:大多数团队把预算花在「更聪明的分发」上,真正决定分数的却是「名单里放了谁」。一个从 187 个模型里挑的路由,不如一个只认识两个模型的硬币。

当然有边界:论文只测单轮、英文、基准题,真实流量和多轮对话(切模型会丢 KV 缓存)可能更复杂;未来模型若真开始强烈专精,长名单也许会回来。但在今天,举证责任在路由器这边。

今晚:先让你的路由器和硬币比一次

拿出路由日志,做一件事:挑出你名单里性价比最高的便宜模型和最强的那个,按你当前的单次花费算出随机比例 p,在同一批留出题上跑「p 概率选便宜、1−p 选贵」。再顺手查两件事:同一来源内部,被升级的题是不是更难的那批;被升级的是不是只是答案短的那批。如果路由器连硬币都打不过,先删名单,再谈路由。

判断很简单:选对两个模型,比聪明的路由器值钱——名单才是准确率,路由只是零头。

The model router you pay for probably loses to a coin flip.

The router pitch is smooth: easy queries go to the cheap model, hard ones escalate to the expensive one, the bill halves and the score holds. The pitch is rarely tested against the dumbest possible control — spend the same money flipping a weighted coin between two well-chosen models. Once you run that control, the pitch gets harder to tell.

No commercial router beats a coin flip

14 settings, zero wins against the coin

Fastino Labs, Dynamic LLM Routers are Often Misguided (arXiv:2610.02762) put 6 commercial routers (OpenRouter, Microsoft Azure, vLLM, Not Diamond, Nadir, Orca) through 14 settings, using 26 models, 17 benchmarks, 8 task categories and 800 evaluation queries, with every cost and accuracy re-measured the same way.

The control is one sentence: pick Gemini 3.7 Flash or Opus 5 (high) at random, weighted to match the router's spend. Every one of the 14 deltas is negative. Apart from vLLM Semantic Router, which was fit on the paper's own training data (−1.0 pp, within noise), all are significantly behind. The worst, Orca (adaptive), picked 96 distinct models from a 187-model roster and landed at 71.0% — 10.5 points below the same-cost coin. Not Diamond (cost) and OpenRouter auto-β (high) trail by 8.8.

Even when the coin's two sides are restricted to models each router could access at the time (GPT 5.6 Luna / Sol), no router significantly wins. The control is not cheating with newer models.

What routers look at: not difficulty

The roster, not the router

The paper traces the gap to four patterns. The sharpest: no commercial router's choices correlate with query difficulty above γ = +0.20. They behave as if they recognize which dataset a query came from and send that whole source to one model — hard and easy queries in the same source treated alike.

The twist is that this is partly what the exam rewards. The standard cost–accuracy Pareto metric uses realized cost. On the hardest queries both models fail, so escalation buys little; escalating the 60–80th difficulty band beats the 80–100th by 6.3 points. Worse, it rewards picking short answers: on a gpt-oss-20b → Opus 5 pair, escalating only the shortest-answer 20% reaches 59.5% at $1.03 per 1k queries, just 1.6 points below escalating the hardest 20% at 1.5% of the price. The authors call it accuracy hacking — calling the smart model only when it is cheap to do so.

The roster is where the value is

Now the other side. Gemini 3.7 Flash trails Opus 5 (high) by only 1.6 points at a 96% discount. Even a router with perfect knowledge of each query's true difficulty, choosing roster and thresholds on the evaluation set itself, gains at most 1 pp from a roster larger than two. The "models specialize" assumption that justifies long rosters is weak in practice: adding 8 specialization dimensions only moves fit error (NLL) from 0.27 to 0.24.

The authors then built a two-model router that avoids all four patterns. It ties the coin too (−0.2 to −0.7 pp, within noise). The reason fits in one line: once the roster is right, there is little left to route.

So here is the argument worth having: most teams spend their effort on smarter dispatch, while the score is decided by who is on the list. A router choosing from 187 models loses to a coin that knows two.

Boundaries apply: single-turn, English, benchmark queries only; real traffic and multi-turn chats (switching models drops the KV cache) may differ, and if future models truly specialize, long rosters may earn their keep. Today, the burden of proof sits with the router.

Tonight: race your router against a coin

Pull your routing logs and do one thing: take the best-value cheap model and the strongest model on your list, compute the mix p that matches your current per-query spend, and run "cheap with probability p, strong otherwise" on the same held-out set. While you are there, check two things: within each source, are escalated queries actually the harder ones, and are escalated queries just the short-answer ones. If the router cannot beat the coin, cut the roster before you tune the router.

The judgment is simple: picking two models well beats a clever router — the roster is the accuracy; routing is rounding.