Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

About

Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. Recent methods often appear to deliver high safety with high utility, but this conclusion rests largely on coarse global utility metrics (e.g., FID, CLIPScore) that are insensitive to fine-grained semantic correctness, creating an illusion of high utility. We show that when utility is measured with structured evaluation, this illusion breaks: on TIFA (Text-to-Image Faithfulness evaluation with Question Answering), safety-aligned models suffer substantial drops in semantic fidelity, including failures in object counts, attributes, and relationships. To diagnose the source of this gap, we analyze the text-encoder prompt embedding space and uncover semantic collapse, a contraction of embedding spread coupled with distortion of inter-prompt similarity structure, which strongly correlates with structured utility loss. Guided by this insight, we propose StructureAware Geometric Regularization (SAGE), a safety alignment objective that explicitly preserves embedding spread and inter-prompt relational structure during adaptation. Our method restores structured utility (TIFA +5.0% over prior state-of-the-art) while maintaining strong safety performance and competitive coarse-grained utility scores. Our source code and trained models are available at https://adeelyousaf.github.io/SAGE_ECCV26_Project_Page/.

Adeel Yousaf, Soumik Ghosh, James Beetham, Amrit Singh Bedi, Mubarak Shah• 2026

Related benchmarks

TaskDatasetResultRank
Compositional Image GenerationGenEval
Overall Score59.8
94
Defense against adversarial promptsMMA (test)
Attack Success Rate0.5
15
Defense against adversarial promptsP4D (test)
Attack Success Rate1.1
15
Defense against adversarial promptsRing-A-Bell (test)
Attack Success Rate1.87
15
Defense against adversarial promptsSneaky (test)
Attack Success Rate0.00e+0
15
Compositional generationT2I-CompBench++
Color Score38
10
Text-to-Image GenerationUtility Evaluation Set
CLIPScore26.4
10
Safety EvaluationSafety Benchmarks
Average ASR1.2
10
Text-to-Image GenerationTIFA
Object Score77.6
10
Text-to-Image Safety EvaluationSafety Benchmarks (MMA, Sneaky, I2P-S, Ring, P4D)
MMA0.4
10
Showing 10 of 15 rows

Other info

Follow for update