Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging

About

The "alignment tax" of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration, making models overconfident, less reliable, and model outputs less diverse. We show that this trade-off can be navigated effectively via a simple post-hoc intervention: interpolating between a model's weights before and after alignment. Crucially, this is not a strict trade-off. We find that the process consistently reveals Pareto-optimal interpolations - models that improve accuracy beyond both parents while substantially recovering the calibration lost during alignment. Our work demonstrates that simple model merging provides a computationally efficient method for mitigating the full scope of the alignment tax, yielding models that are more capable and more reliable.

Tiancheng Hu, Benjamin Minixhofer, Nigel Collier• 2025

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningMATH L5
Accuracy0.6314
162
Instruction FollowingIFEval
Accuracy (Greedy)77.63
52
Multi-task Language UnderstandingMMLU-Pro
Accuracy (%)51.91
42
ReasoningBBH
Accuracy67.56
42
Scientific ReasoningGPQA
Accuracy0.3817
42
Mathematical ReasoningMATH (test)
Accuracy0.4808
24
Showing 6 of 6 rows

Other info

Follow for update