ShieldGemma 2: Robust and Tractable Image Content Moderation
About
We introduce ShieldGemma 2, a 4B parameter image content moderation model built on Gemma 3. This model provides robust safety risk predictions across the following key harm categories: Sexually Explicit, Violence \& Gore, and Dangerous Content for synthetic images (e.g. output of any image generation model) and natural images (e.g. any image input to a Vision-Language Model). We evaluated on both internal and external benchmarks to demonstrate state-of-the-art performance compared to LlavaGuard \citep{helff2024llavaguard}, GPT-4o mini \citep{hurst2024gpt}, and the base Gemma 3 model \citep{gemma_2025} based on our policies. Additionally, we present a novel adversarial data generation pipeline which enables a controlled, diverse, and robust image generation. ShieldGemma 2 provides an open image moderation tool to advance multimodal safety and responsible AI development.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Safety Evaluation | UnsafeBench | F1 Score53.19 | 39 | |
| Multimodal Safety | BeaverTails V | F1 Score91.54 | 15 | |
| Multimodal Safety | MM-Safety | F1 Score96.56 | 15 | |
| Multimodal Safety | MMDS-R | F1 Score68.57 | 15 | |
| Multimodal Safety | Multimodal Safety Suite Avg | Macro-average F174.28 | 15 | |
| Harm Recognition | Facebook Hateful Memes | F1 Score82.09 | 15 | |
| Multimodal Safety | VLGuard | F1 Score69.67 | 15 | |
| Multimodal Safety | MMDS-Q | F1 Score64.42 | 15 | |
| Harm Recognition | COCO Crime Scene slice | F1 Score66.67 | 15 | |
| Multimodal Safety | JailBreakV | F1 Score90.17 | 15 |