Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

EntroRouter: Learning Efficient Model Routing via Entropy Regulation

About

Model routing balances solution accuracy and computational cost by selecting among models of varying capabilities. While recent multi-round frameworks interleave reasoning and planning, we identify a structural failure mode termed Trust Region Collapse. We demonstrate that the deep coupling of reasoning and routing, exacerbated by the dominance of strong pre-training priors under sparse supervision, leads to degenerate local optima where capable experts are systematically suppressed. To decouple these processes, we propose $\textbf{EntroRouter}$, a single-round routing framework that treats entropy regulation as a core objective. We first initialize the policy via Soft Supervision, fitting a distribution of suitable models to establish a high-entropy prior for exploration. Subsequently, we stabilize Reinforcement Learning using a Soft Anchor, which utilizes offline capability estimates to orchestrate controlled entropy contraction within a safe trust region. Extensive experiments demonstrate that EntroRouter retains 98.3% of the strongest expert's accuracy while reducing computational costs by 48.25%.

Kaiyi Zhang, Xueliang Zhao, Zhuocheng Gong, Wei Wu, Yankai Lin• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningAMC
Accuracy96.6
40
Mathematical ReasoningAIME 25
Accuracy78.8
35
Question AnsweringMMLU-Pro
Accuracy80.6
27
Question AnsweringGPQA Diamond
Quality (%)73.2
11
Mathematical ReasoningHMMT Nov 25
Accuracy80
9
Mathematical ReasoningOlympiadBench
Accuracy81.8
9
Mathematical ReasoningAIME 24
Accuracy90.6
9
Mathematical ReasoningMATH 500
Accuracy98.4
9
Mathematical ReasoningSMT 2025
Accuracy80.4
9
Mathematical ReasoningMath Benchmarks Suite Avg.
Accuracy88.6
9
Showing 10 of 11 rows

Other info

Follow for update