SARIF: Segment Anything for Robust Image Forensics
About
Image forgery localization remains challenging due to diverse manipulation techniques and distribution shifts. Existing forgery localization models achieve high accuracy on benchmarks but often struggle with cross-domain generalization and robustness. In this paper, we propose SARIF (Segment Anything for Robust Image Forensics), a framework that leverages the Segment Anything Model (SAM), which has a promptable architecture and strong generalization ability. SARIF introduces a feedback-guided mask decoder and a dual-encoder design that extracts forgery-specific information to capture forensic traces while exploiting the SAM architecture. To localize manipulated regions, we design a block-wise prompting mechanism that derives forgery-specific cues from residual features between an adapted encoder and its frozen counterpart. These features are fused with the previous mask prompt to drive a feedback-based mask refinement process, enabling automatic forgery segmentation without manual input. Extensive experiments on standard forgery-localization benchmarks show that SARIF achieves strong average cross-dataset performance and robustness to common image corruptions.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Image Forgery Localization | Columbia Unseen Domain | mIoU55.8 | 30 | |
| Image Forgery Localization | CASIA v1 | Pixel-level AUC58.4 | 20 | |
| Image Forgery Localization | Columbia | DSC58.4 | 15 | |
| Image Forgery Localization | CASIA Seen Domain v2 | mIoU56.7 | 15 | |
| Image Forgery Localization | DIS25k Unseen Domain | mIoU41.5 | 15 | |
| Image Forgery Localization | CASIA Unseen Domain v1 | mIoU52 | 15 | |
| Image Forgery Localization | CASIA v2 (test) | DSC63.1 | 15 | |
| Image Forgery Localization | CoMoFoD | DSC65.4 | 15 | |
| Image Forgery Localization | IMD 2020 | DSC48.4 | 15 | |
| Image Forgery Localization | IMD Unseen Domain 2020 | mIoU40.4 | 15 |