Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

STAIR: Improving Safety Alignment with Introspective Reasoning

About

Ensuring the safety and harmlessness of Large Language Models (LLMs) has become equally critical as their performance in applications. However, existing safety alignment methods typically suffer from safety-performance trade-offs and the susceptibility to jailbreak attacks, primarily due to their reliance on direct refusals for malicious queries. In this paper, we propose STAIR, a novel framework that integrates SafeTy Alignment with Itrospective Reasoning. We enable LLMs to identify safety risks through step-by-step analysis by self-improving chain-of-thought (CoT) reasoning with safety awareness. STAIR first equips the model with a structured reasoning capability and then advances safety alignment via iterative preference optimization on step-level reasoning data generated using our newly proposed Safety-Informed Monte Carlo Tree Search (SI-MCTS). We further train a process reward model on this data to guide test-time searches for improved responses. Extensive experiments show that STAIR effectively mitigates harmful outputs while better preserving helpfulness, compared to instinctive alignment strategies. With test-time scaling, STAIR achieves a safety performance comparable to Claude-3.5 against popular jailbreak attacks. Relevant resources in this work are available at https://github.com/thu-ml/STAIR.

Yichi Zhang, Siyuan Zhang, Yao Huang, Zeyu Xia, Zhengwei Fang, Xiao Yang, Ranjie Duan, Dong Yan, Yinpeng Dong, Jun Zhu• 2025

Related benchmarks

TaskDatasetResultRank
Scientific Question AnsweringGPQA Diamond
Accuracy48.98
131
Massive Multitask Language UnderstandingMMLU-Pro
Accuracy (MMLU-Pro)44.92
122
Over-refusalXSTest--
102
Mathematical ReasoningGSM8K
Accuracy87.6
80
Safety EvaluationStrongREJECT--
77
Mathematical ReasoningMATH500
Accuracy (%)83.8
56
Safety EvaluationHarmBench
PAIR78.75
39
Safety EvaluationWildChat
Safe@177.8
34
Specification AlignmentSPECBENCH Average over scenarios
Safety Score89.27
33
Safety EvaluationAdvBench
Overall Safety Score100
30
Showing 10 of 54 rows

Other info

Follow for update