Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Deliberative Alignment: Reasoning Enables Safer Language Models

About

As large-scale language models increasingly impact safety-critical domains, ensuring their reliable adherence to well-defined principles remains a fundamental challenge. We introduce Deliberative Alignment, a new paradigm that directly teaches the model safety specifications and trains it to explicitly recall and accurately reason over the specifications before answering. We used this approach to align OpenAI's o-series models, and achieved highly precise adherence to OpenAI's safety policies, without requiring human-written chain-of-thoughts or answers. Deliberative Alignment pushes the Pareto frontier by simultaneously increasing robustness to jailbreaks while decreasing overrefusal rates, and also improves out-of-distribution generalization. We demonstrate that reasoning over explicitly specified policies enables more scalable, trustworthy, and interpretable alignment.

Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, Rachel Dias, Andrea Vallone, Hongyu Ren, Jason Wei, Hyung Won Chung, Sam Toyer, Johannes Heidecke, Alex Beutel, Amelia Glaese• 2024

Related benchmarks

TaskDatasetResultRank
Jailbreak DefensePAIR
ASR18
97
Over-refusalOver-refusal XSTest and OKTest
Over-refusal Accuracy (XSTest)97.2
12
Jailbreak RobustnessAdvBench w/o attack (original)
ASR0.00e+0
9
Jailbreak RobustnessPAIR v1 (test)
Compliance Rate11.2
9
Safety EvaluationBeaverTail v1 (eval)
Compliance Rate4.8
9
Jailbreak RobustnessAdvBench AdvReasoning (original)
ASR58
9
Jailbreak RobustnessWildJailbreak (eval)
Compliance Rate42.5
9
Safety EvaluationMalicious Instruct v1 (test)
Compliance Rate1
9
Safety EvaluationXSTest Safe v1
Accuracy88.8
9
Jailbreak RobustnessJailbreakV v1 (test)
Compliance Rate0.00e+0
9
Showing 10 of 14 rows

Other info

Follow for update