EntroRouter: Learning Efficient Model Routing via Entropy Regulation
About
Model routing balances solution accuracy and computational cost by selecting among models of varying capabilities. While recent multi-round frameworks interleave reasoning and planning, we identify a structural failure mode termed Trust Region Collapse. We demonstrate that the deep coupling of reasoning and routing, exacerbated by the dominance of strong pre-training priors under sparse supervision, leads to degenerate local optima where capable experts are systematically suppressed. To decouple these processes, we propose $\textbf{EntroRouter}$, a single-round routing framework that treats entropy regulation as a core objective. We first initialize the policy via Soft Supervision, fitting a distribution of suitable models to establish a high-entropy prior for exploration. Subsequently, we stabilize Reinforcement Learning using a Soft Anchor, which utilizes offline capability estimates to orchestrate controlled entropy contraction within a safe trust region. Extensive experiments demonstrate that EntroRouter retains 98.3% of the strongest expert's accuracy while reducing computational costs by 48.25%.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Mathematical Reasoning | AMC | Accuracy96.6 | 40 | |
| Mathematical Reasoning | AIME 25 | Accuracy78.8 | 35 | |
| Question Answering | MMLU-Pro | Accuracy80.6 | 27 | |
| Question Answering | GPQA Diamond | Quality (%)73.2 | 11 | |
| Mathematical Reasoning | HMMT Nov 25 | Accuracy80 | 9 | |
| Mathematical Reasoning | OlympiadBench | Accuracy81.8 | 9 | |
| Mathematical Reasoning | AIME 24 | Accuracy90.6 | 9 | |
| Mathematical Reasoning | MATH 500 | Accuracy98.4 | 9 | |
| Mathematical Reasoning | SMT 2025 | Accuracy80.4 | 9 | |
| Mathematical Reasoning | Math Benchmarks Suite Avg. | Accuracy88.6 | 9 |