Foundation model for LLM routing

Pretrain once.
Route anywhere.

RouteFM learns a reusable routing capability across heterogeneous environments, then adapts to unseen tasks and candidate pools from a few behavioral observations—without retraining.

Guannan Lai · Han-Jia Ye

Nanjing University · LAMDA

Environment 01
frozenRouteFM
Unseen environment
01 / Motivation

Routing should transfer,
not restart.

Conventional routers are fitted to one workload and one fixed candidate pool. When the environment changes, routing starts over.

RouteFM turns model routing into an in-context adaptation problem. It characterizes anonymous candidates using observed query, quality, and relative-cost tuples, and transfers one frozen router across domains, modalities, candidate pools, and context budgets.

Comparison between task-specific routing and RouteFM's global routing pretraining approach
From local router fitting to global routing pretraining.
02 / Method

Behavior is the interface.

Candidate names label the output; they are never used as model features. RouteFM reasons from evidence instead.

01

Observe

Collect a small context of queries, outcome quality, and relative cost for every candidate.

02

Characterize

Compress behavioral evidence into capability profiles while retaining fine-grained context.

03

Compare

Jointly compare the current candidate pool with a permutation-equivariant Transformer.

04

Route

Predict target-specific quality and cost, then select the candidate for the new query.

RouteFM architecture with capability profiles, context retrieval, and candidate-pool Transformer
RouteFM architecture. See the paper for the full training objective and implementation details.
03 / Evidence

A frozen router,
outside its training world.

On MMR-Bench, excluded entirely from pretraining, RouteFM transfers across modality, task, and candidate-pool shifts.

RouteFM · K=8 0.7323
MMR-Bench quality
Strongest non-RouteFM baseline 0.7100
Same K=8 setting
Absolute gain +2.23 quality points with eight observations per candidate

Results use the paper's dataset-wise five-fold evaluation protocol. Please consult the paper for baselines, uncertainty, and complete experimental settings.

04 / Live demo

Give it context.
Watch it route.

Choose an environment, target query, and observation budget. The candidates remain anonymous—their behavior is all RouteFM sees.

RouteFM-BGE · CPU
Open full screen

The free embedded demo switches among predictions precomputed with the released frozen RouteFM-BGE checkpoint. Its synthetic contexts are illustrative—not benchmark records or live provider outputs. The repository also includes a full Gradio app for arbitrary queries.

05 / Get started

Route in a few lines.

Install the published Python package. The first run downloads the immutable RouteFM-BGE checkpoint and its frozen query encoder.

quickstart.py
$ pip install "routefm-router[bge]"

from routefm import RouteFMRouter

router = RouteFMRouter.from_pretrained(
    encoder="bge", device="cpu"
)
router.set_context(candidates)

decision = router.route(
    "Prove there are infinitely many primes."
)
print(decision.model_name)
06 / Cite

Build on RouteFM.

If RouteFM supports your research, please cite our paper.

@misc{lai2026pretrainoncerouteanywhere,
  title  = {Pretrain Once, Route Anywhere: Towards a
            Foundation Model for LLM Routing},
  author = {Guannan Lai and Han-Jia Ye},
  year   = {2026},
  eprint = {2609.37362},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI}
}