Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models

About

While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful queries. To address this vulnerability, prior works depend heavily on external manual data annotation for safety alignment. However, we observe that LRMs can inherently identify safety risks when being re-presented with original queries alongside their own reasoning trajectories -- a capability we term Latent Safety Awareness. To leverage this safety awareness, we first employ Supervised Fine-Tuning (SFT) to explicitly induce safe tags to trigger safety analysis and guidance following the initial reasoning content for unsafe queries, while preserving standard responses for general queries to ensure adaptive triggering. Subsequently, we apply Direct Preference Optimization (DPO) to further enhance the correctness and stability of the safety analysis and guidance. Notably, responses required for both training stages are entirely generated by models being optimized. With (Safe Trigger) SFT and DPO, experimental results demonstrate significant safety enhancement. For example, the Attack Success Rate (ASR) of DeepSeek-R1-Distill-Llama-8B, on average, drops 24.65% and 36.72% on harmful and jailbreak benchmarks, respectively. Finally, our Safe Trigger method exerts almost no negative impact on general performance or user experience.

Ke Miao, Jiaxin Li, Hongliang Chen, Yuke Hu, Zhan Qin• 2026

Related benchmarks

TaskDatasetResultRank
Safety EvaluationWildJailbreak
ASR0.0145
90
Common Sense ReasoningWinoGrande
Accuracy85.64
67
Over-refusal evaluationXSTest Safe
Over-refusal Rate2
45
Harmful Content Safety EvaluationHexPhi
Attack Success Rate (ASR)1.67
20
Jailbreak Safety EvaluationMulti-Shot Jailbreak (MSJ)
ASR0.00e+0
20
Question Answering and ReasoningARC Mean of Easy and Challenge
Accuracy94.75
20
Harmful Content Safety EvaluationXsTest Harmful
Attack Success Rate (ASR)0.00e+0
20
Jailbreak Safety EvaluationPAP
ASR2
20
Reading ComprehensionDROP
DROP Score0.673
20
Harmful Content Safety EvaluationAdvBench
Attack Success Rate (ASR)0.38
20
Showing 10 of 14 rows

Other info

Follow for update