Sparse supervision for economical LLM routing

Routing should
pay for itself.

SAVERouter accounts for the feedback acquisition cost paid before deployment—and learns an effective router from only a sparse set of query-model outcomes.

33–41% of available training feedback used
1.9–9.5× earlier break-even vs. the fastest conventional router
4 heterogeneous routing benchmarks

Results from the paper's main setting across four evaluated benchmarks.

A cheaper route can still be a bad investment.

Conventional evaluations emphasize serving-time quality and cost. But training a router first requires model executions on historical queries. SAVERouter puts that upfront supervision expenditure into the objective and the evaluation.

Before deployment Acquire supervision Upfront expenditure C0
After deployment Save per request Baseline Cb − routed Cr
The real question When does it pay back? Measured by SA-BEP and SA-CR

Spend feedback where it teaches the router most.

SAVERouter adaptively selects exactly K model outcomes per training query, shares capability evidence across related queries, then learns query-level residual corrections for fine-grained routing.

The SAVERouter method: sparse feedback acquisition, hierarchical capability estimation, and query-level refinement
Method overview from the paper.
01

Adaptive acquisition

Grouped empirical-Bayes UCB chooses informative model outcomes under a fixed feedback budget.

02

Capability sharing

A structured group-model prior pools evidence across related queries without requiring dense labels.

03

Local refinement

Contextual residuals retain query-specific variation and construct a quality-cost routing frontier.

Earlier payback without sacrificing routing quality.

The paper evaluates SAVERouter on LLMRouterBench, Mixinstruct, MMR-Bench, and RouterBench. Lower CR, SA-BEP, and SA-CR are better; higher peak score is better.

Main SAVERouter results across four routing benchmarks
Main results table from the paper.

Estimate when your router pays back.

Enter costs in any consistent unit. The calculator visualizes supervision-amortized break-even and total cost over deployment traffic.

Illustrative calculator; paper results use benchmark-specific costs.

SA-BEP7,693deployment queries
SA-CR@H0.3550paid back by horizon
Best single modelRouter + supervision

From thirty seconds to the full benchmark suite.

Choose the path that matches your goal. The smoke test and tutorial require no benchmark downloads; the paper suite uses frozen profiles and pinned data revisions.

Quick check

Smoke test

python -m pip install saverouter
saverouter smoke-test
Setup guide →
Paper results

Full reproduction

bash scripts/reproduce.sh all

Approximately 12–15 GB of free disk space; CUDA is recommended for MMR-Bench and first-time feature extraction.

Reproduction guide →

Build routing systems that account for the bill before deployment.

@article{lai2026routing,
  title   = {Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing},
  author  = {Lai, Guannan and Bian, Gelin and Ma, Hao-Xuan and Jiang, Jun-Peng and Chen, Long and Liu, Jian-Dong and Tan, Zhi-Hao and Ye, Han-Jia},
  journal = {arXiv preprint arXiv:2609.37402},
  year    = {2026}
}