Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Policy-Guided Stepwise Model Routing for Cost-Effective Reasoning

About

Inference-time computation has greatly enhanced the performance of large language models (LLMs) on challenging reasoning tasks, but this strategy can incur high inference costs. One solution is to route intermediate chain-of-thought (CoT) states to language models of different sizes; however, existing approaches rely on handcrafted routing strategies that limit performance, or on training large process reward models that may be infeasible in many applications. We formulate stepwise model routing as a constrained decision-making problem, which we solve by training a small control policy using reinforcement learning in conjunction with threshold calibration to tune the performance-efficiency tradeoff. We validate our method on three math benchmarks (GSM8K, MATH500, and OmniMath) on both open and closed models. Our method consistently improves the accuracy-cost tradeoff compared to handcrafted approaches, while achieving a comparable tradeoff to methods that require training large process reward models.

Wenwen Si, Insup Lee, Osbert Bastani• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningMATH500 (test)
Accuracy85
922
Mathematical ReasoningOmniMath (test)--
20
Mathematical ReasoningMATH 500
Accuracy82.8
6
Mathematical ReasoningGSM8K
Accuracy (GSM8K)94.5
6
Mathematical ReasoningOmniMath
Accuracy29.1
6
Mathematical ReasoningGSM8K (test)
Accuracy94.5
6
Showing 6 of 6 rows

Other info

Follow for update