In-Training Defenses against Emergent Misalignment in Language Models

About

Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EM): Even a small, domain-specific fine-tune can induce harmful behaviors far outside the target domain. Even in the case where model weights are hidden behind a fine-tuning API, this gives attackers inadvertent access to a broadly misaligned model in a way that can be hard to detect from the fine-tuning data alone. We present the first systematic study of in-training safeguards against EM that are practical for providers who expose fine-tuning via an API: We evaluate whether they a) prevent broad misalignment, b) allow narrow misalignment, c) learn well on benign tasks, and d) remain coherent. We investigate five training regularization interventions: (i) KL-divergence regularization toward a safe reference model, (ii) $\ell_2$ distance in feature space, (iii) preventive steering with an evil persona vector, (iv) interleaving training examples from a general instruct-tuning dataset and (v) inoculation prompting. We demonstrate that selecting interleaving data by the perplexity gap between aligned and misaligned models yields the best results overall.

David Kacz\'er, Magnus J{\o}rgenv{\aa}g, Clemens Vetter, Esha Afzal, Robin Haselhorst, Lucie Flek, Florian Mai• 2025

Related benchmarks

Task	Dataset	Result
Emergent Misalignment Measurement	Code	Misalignment0.21	6
Emergent Misalignment Measurement	Legal	Misalignment4.01	6
Emergent Misalignment Measurement	Medical General Evaluation	Misalignment7.89	6
Misaligned Task Learning	Code In-domain	Misalignment54.95	6
Misaligned Task Learning	Legal In-domain	Misalignment27.17	6
Emergent Misalignment Measurement	Security General evaluation	Misalignment Score5.68	6
Misaligned Task Learning	Medical In-domain	Misalignment59.2	6
Misaligned Task Learning	Security In-domain	Misalignment22.03	6

Showing 8 of 8 rows

Other info

Follow for update

@wizwand_team Discord