Effective Image Tampering Localization via Enhanced Transformer and Co-attention Fusion
About
Powerful manipulation techniques have made digital image forgeries be easily created and widespread without leaving visual anomalies. The blind localization of tampered regions becomes quite significant for image forensics. In this paper, we propose an effective image tampering localization network (EITLNet) based on a two-branch enhanced transformer encoder with attention-based feature fusion. Specifically, a feature enhancement module is designed to enhance the feature representation ability of the transformer encoder. The features extracted from RGB and noise streams are fused effectively by the coordinate attention-based fusion module at multiple scales. Extensive experimental results verify that the proposed scheme achieves the state-of-the-art generalization ability and robustness in various benchmark datasets. Code will be public at https://github.com/multimediaFor/EITLNet.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Image Forgery Localization | Columbia Unseen Domain | mIoU20.9 | 30 | |
| Image Forgery Localization | CASIA v1 | Pixel-level AUC52.9 | 20 | |
| Image Forgery Localization | CASIA Seen Domain v2 | mIoU47.9 | 15 | |
| Image Forgery Localization | CASIA Unseen Domain v1 | mIoU46.5 | 15 | |
| Image Forgery Localization | DIS25k Unseen Domain | mIoU25.6 | 15 | |
| Image Forgery Localization | CASIA v2 (test) | DSC54 | 15 | |
| Image Forgery Localization | IMD Unseen Domain 2020 | mIoU19.7 | 15 | |
| Image Forgery Localization | MSID Unseen Domain | mIoU45.9 | 15 | |
| Image Forgery Localization | MSID | DSC58.8 | 15 | |
| Image Forgery Localization | DIS25k | DSC30.8 | 15 |