Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SafeGene: Reusable Adapters for Transferable Safety Alignment

About

Open-weight LLMs are increasingly fine-tuned into customized assistants, but downstream fine-tuning can weaken safety alignment and make models more vulnerable to malicious prompts, even when the training data is not intentionally harmful. This creates a recurring safety recovery problem as target models are repeatedly updated with new task data or user interactions. We propose SafeGene, a reusable safety-adapter module designed for cross-task reuse within each architecture-compatible model family. Rather than treating safety recovery as a model-specific repair step, SafeGene treats safety capability as an independent, reusable adapter representation decoupled from task-specific updates. This representation is obtained from aligned--degraded model discrepancies, refined into task-transferable safety vectors through data-aware layer selection, and expressed in each downstream task-adapted model via few-shot layer-wise coefficient recalibration. Experiments across multiple model families, downstream tasks, and safety judges show that SafeGene-enhanced models reduce harmful response rates while maintaining downstream performance, outperforming representative safe adaptation methods in safety--utility trade-off.

Yanghan Wang, Zhiqiang Kou, Fu Feng, Jing Wang, Xin Geng• 2026

Related benchmarks

TaskDatasetResultRank
Question AnsweringBoolQ
Accuracy87.86
233
Natural Language InferenceMNLI--
80
Safety EvaluationBeavertails
ASR34.16
44
Safety EvaluationDirectRefusal
Attack Success Rate (ASR)30.66
25
Safety EvaluationBeaverTails & DirectRefusal Average
Average ASR32.84
25
Sentiment AnalysisSST2
Top-1 Accuracy96.56
10
Natural Language InferenceMNLI
Accuracy84.45
5
Sentiment AnalysisSST2
Accuracy95.3
5
Text ClassificationAG-News
Accuracy88.91
5
Question AnsweringBoolQ
Accuracy79.63
5
Showing 10 of 10 rows

Other info

Follow for update