Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Defending Against Harmful Supervision Hidden in Benign Samples

About

Existing defenses are effective when harmful content is explicitly mixed into downstream fine-tuning data, but crafted samples can instead hide harmful supervision inside benign tasks. We propose Embedded Attack, where harmful QA pairs are embedded within benign training samples, and show that representative guardrails often fail to detect them at the example level. To address this, we propose Dual-Reference SFT (DR-SFT), which adapts DPO-style contrastive objective design to SFT through token-level regularization, mitigating harmful fine-tuning beyond coarse data filtering.

Bang An, Yibo Yang, Dandan Guo, Ebtisam Alshehri, Carlos Hinojosa, Bernard Ghanem• 2026

Related benchmarks

TaskDatasetResultRank
Safety EvaluationHEX-PHI
Attack Success Rate (ASR)1.8
107
Malicious Fine-tuning AttackHEX-PHI
Attack Success Rate (ASR)3.9
24
Embedded Direct AttackGSM8K ⊕ HEx-PHI
Utility74.4
12
Embedded Direct AttackGSM8K ⊕ BeaverTails
Utility70.6
12
Embedded Direct AttackSamsum ⊕ HEx-PHI
Utility50.1
12
Harmful Fine-tuning AttackBeavertails
Attack Success Rate (ASR)17
12
Backdoor AttackGSM8K ⊕ HEx-PHI
ASR (Without Trigger)0.6
8
Backdoor AttackGSM8K ⊕ BeaverTails
ASR (No Trigger)10
8
SQL GenerationSQL-Create Context
Utility98
4
Showing 9 of 9 rows

Other info

Follow for update