Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance

About

Fine-tuning safety-aligned large language models (LLMs) can substantially compromise their safety. Previous approaches require many safety samples or calibration sets, which not only incur significant computational overhead during realignment but also lead to noticeable degradation in model utility. Contrary to this belief, we show that safety alignment can be fully recovered with only a single safety example, without sacrificing utility and at minimal cost. Remarkably, this recovery is effective regardless of the number of harmful examples used in fine-tuning or the size of the underlying model, and convergence is achieved within just a few epochs. Furthermore, we uncover the low-rank structure of the safety gradient, which explains why such efficient correction is possible. We validate our findings across five safety-aligned LLMs and multiple datasets, demonstrating the generality of our approach.

Jiawen Zhang, Lipeng He, Kejia Chen, Jian Lou, Jian Liu, Xiaohu Yang, Ruoxi Jia• 2026

Related benchmarks

TaskDatasetResultRank
Code GenerationHumanEval (test)
Pass@162
701
Sentiment AnalysisSST-2 (test)
Accuracy92
162
Sentiment AnalysisSST2
ASR97
131
Linguistic AcceptabilityCOLA
Accuracy (CoLA)81
108
Text ClassificationSST-2
CACC93
80
Topic ClassificationAGNews
ASR0.02
78
Jailbreak DefenseLlama model jailbreak evaluation prompts
ASR4
60
Safety Defense EvaluationMedicine
ASR22
60
Safety Defense EvaluationMATH
ASR85
60
Sentiment AnalysisSST2 BadNets Attack (test)
ASR94
14
Showing 10 of 19 rows

Other info

Follow for update