Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Safety Alignment of LMs via Non-cooperative Games

About

Ensuring the safety of language models (LMs) while maintaining their usefulness remains a critical challenge in AI alignment. Current approaches rely on sequential adversarial training: generating adversarial prompts and fine-tuning LMs to defend against them. We introduce a different paradigm: framing safety alignment as a non-zero-sum game between an Attacker LM and a Defender LM trained jointly via online reinforcement learning. Each LM continuously adapts to the other's evolving strategies, driving iterative improvement. Our method uses a preference-based reward signal derived from pairwise comparisons instead of point-wise scores, providing more robust supervision and potentially reducing reward hacking. Our RL recipe, AdvGame, shifts the Pareto frontier of safety and utility, yielding a Defender LM that is simultaneously more helpful and more resilient to adversarial attacks. In addition, the resulting Attacker LM converges into a strong, general-purpose red-teaming agent that can be directly deployed to probe arbitrary target models. Code at github.com/facebookresearch/advgame.

Anselm Paulus, Ilia Kulikov, Brandon Amos, R\'emi Munos, Ivan Evtimov, Kamalika Chaudhuri, Arman Zharmagambetov• 2025

Related benchmarks

TaskDatasetResultRank
Safety EvaluationHarmBench
ASR4.7
153
Safety EvaluationWildGuard (test)--
27
Safety EvaluationDAN--
26
Benign ComplianceWJB (WildJailbreak)
Compliance Rate94.4
15
Benign ComplianceXSTest
Comply Score81.2
12
Safety EvaluationWJB
ASR (%)8.5
5
Instruction FollowingIFBench
Utility Score30.7
5
TruthfulnessTruthfulQA
Utility48.7
5
General Language UnderstandingMMLU
Utility Score71.8
5
Showing 9 of 9 rows

Other info

Follow for update