Router-R1: Teaching LLMs Multi-Round Routing and Aggregation via Reinforcement Learning

About

The rapid emergence of diverse large language models (LLMs) has spurred the development of LLM routers that assign user queries to the most suitable model. However, existing LLM routers typically perform a single-round, one-to-one mapping (\textit{i.e.}, assigning each query to a single model in isolation), which limits their capability to tackle complex tasks that demand the complementary strengths of multiple LLMs. In this paper, we present \textbf{Router-R1}, a reinforcement learning (RL)-based framework that formulates multi-LLM routing and aggregation as a sequential decision process. Router-R1 instantiates the router itself as a capable LLM, leveraging its reasoning ability to interleave "think" actions (internal deliberation) with "route" actions (dynamic model invocation), and integrates each response into its evolving context. To facilitate learning, we employ a lightweight rule-based reward comprising format rewards, final outcome rewards, and a novel cost reward for optimizing the balance between performance and cost, opening a pathway toward enhancing performance-cost trade-offs via RL. Router-R1 also conditions only on simple model descriptors such as pricing, latency, and example performance, enabling strong generalization to unseen model selection. Experiments on seven general and multi-hop QA benchmarks show that Router-R1 outperforms several strong baselines, achieving superior performance while maintaining robust generalization and cost management.

Haozhen Zhang, Tao Feng, Jiaxuan You• 2025

Related benchmarks

Task	Dataset	Result
Code Generation	HumanEval	Pass@187.5	1043
Mathematical Reasoning	MATH	Accuracy76.56	535
Mathematical Reasoning	MathQA	Accuracy82.81	354
Mathematical Reasoning	AIME 2025	Accuracy10	227
Question Answering	SQuAD 2.0	F172.45	215
Multi-hop Question Answering	2Wiki	--	215
Multi-hop Question Answering	MuSiQue	--	209
Mathematical Reasoning	AMC'23 (test)	Accuracy95	152
Multi-hop QA	HotpotQA	Exact Match36.8	143
Question Answering	HotpotQA	F179.84	132

Showing 10 of 59 rows

Other info

Follow for update

@wizwand_team Discord