Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Toward Calibrated Mixture-of-Experts Under Distribution Shift

About

Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important for understanding and trusting reported probabilities. Recent work shows that enforcing calibration at the level of individual predictors can improve ensemble accuracy and calibration, with mixture-of-experts (MoE) models showing strong empirical improvements in particular; however, the conditions under which calibration helps MoE are not well understood. In this work, we study how MoE models behave under distribution shift, focusing on how routing mechanisms interact with expert-level calibration. We show that expert calibration is sufficient to ensure calibration of the overall model under a broad class of distribution shifts in hard-routed models, but is insufficient for calibrating soft-routed models. To address this, we propose an adversarial reweighting that penalizes calibration errors of the routed aggregate under distribution shift, and we demonstrate that it improves the accuracy-calibration tradeoff both on average and on difficult subsets of the data, across model classes, prediction tasks, and distribution shifts.

Gina Wong, Drew Prinster, Suchi Saria, Rama Chellappa, Anqi Liu• 2026

Related benchmarks

TaskDatasetResultRank
Image ClassificationPACS (leave-one-domain-out)
P Accuracy93.1
39
Image ClassificationCIFAR-10H--
25
Image ClassificationCIFAR-10H hard subset
Accuracy62.4
14
Toxicity ClassificationCivilComments
Average Accuracy91.5
10
Toxicity ClassificationCivilComments Hard subset - demographic identities
Hard Accuracy86.9
7
Showing 5 of 5 rows

Other info

Follow for update