Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ShieldGemma 2: Robust and Tractable Image Content Moderation

About

We introduce ShieldGemma 2, a 4B parameter image content moderation model built on Gemma 3. This model provides robust safety risk predictions across the following key harm categories: Sexually Explicit, Violence \& Gore, and Dangerous Content for synthetic images (e.g. output of any image generation model) and natural images (e.g. any image input to a Vision-Language Model). We evaluated on both internal and external benchmarks to demonstrate state-of-the-art performance compared to LlavaGuard \citep{helff2024llavaguard}, GPT-4o mini \citep{hurst2024gpt}, and the base Gemma 3 model \citep{gemma_2025} based on our policies. Additionally, we present a novel adversarial data generation pipeline which enables a controlled, diverse, and robust image generation. ShieldGemma 2 provides an open image moderation tool to advance multimodal safety and responsible AI development.

Wenjun Zeng, Dana Kurniawan, Ryan Mullins, Yuchi Liu, Tamoghna Saha, Dirichi Ike-Njoku, Jindong Gu, Yiwen Song, Cai Xu, Jingjing Zhou, Aparna Joshi, Shravan Dheep, Mani Malek, Hamid Palangi, Joon Baek, Rick Pereira, Karthik Narasimhan• 2025

Related benchmarks

TaskDatasetResultRank
Safety EvaluationUnsafeBench
F1 Score53.19
39
Multimodal SafetyBeaverTails V
F1 Score91.54
15
Multimodal SafetyMM-Safety
F1 Score96.56
15
Multimodal SafetyMMDS-R
F1 Score68.57
15
Multimodal SafetyMultimodal Safety Suite Avg
Macro-average F174.28
15
Harm RecognitionFacebook Hateful Memes
F1 Score82.09
15
Multimodal SafetyVLGuard
F1 Score69.67
15
Multimodal SafetyMMDS-Q
F1 Score64.42
15
Harm RecognitionCOCO Crime Scene slice
F1 Score66.67
15
Multimodal SafetyJailBreakV
F1 Score90.17
15
Showing 10 of 20 rows

Other info

Follow for update