Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization

About

Visual manipulation localization (VML) aims to identify tampered regions in images and videos, a task that has become increasingly challenging with the rise of advanced editing tools. Existing methods face two central issues. The first is resolution diversity. Resizing or padding can distort subtle forensic cues and introduce unnecessary computational cost. The second is the difficulty of extending spatial models for images to spatio-temporal inputs in videos, which often results in maintaining separate architectures for the two data types. To address these challenges, we propose RelayFormer, a unified framework that adapts to varying resolutions and naturally handles both static and temporal visual data. RelayFormer partitions inputs into fixed-size sub-images and introduces Global Local Relay (GLR) tokens that propagate structured context through a relay-based attention mechanism. This design enables efficient exchange of global cues, such as semantic or temporal consistency, while preserving fine-grained manipulation artifacts. Unlike prior approaches that depend on uniform resizing or sparse attention, RelayFormer scales to variable resolutions and video sequences with minimal overhead. Experiments across diverse benchmarks demonstrate superior performance and strong efficiency, combining resolution adaptivity without interpolation or excessive padding, unified processing for images and videos, and a favorable balance between accuracy and computational cost. Code is available at~\href{https://github.com/WenOOI/RelayFormer}{https://github.com/WenOOI/RelayFormer}.

Wen Huang, Jiarui Yang, Tao Dai, Jiawei Li, Shaoxiong Zhan, Bin Wang, Shu-Tao Xia• 2025

Related benchmarks

TaskDatasetResultRank
Image Manipulation LocalizationNIST16
F1 Score47.6
93
Image Manipulation LocalizationCoverage
F1 Score70.4
78
Image Manipulation LocalizationColumbia
F1 Score88.3
60
Image Manipulation LocalizationCASIA v1
F1 Score80.6
54
Image Forgery DetectionCocoGlide--
20
Image Manipulation LocalizationIMD 2020
F1 Score38.1
18
Image Forgery DetectionAutoSplice
F1 Score37.9
18
Video Manipulation LocalizationMOSE100 E2FGVI
IoU56.1
8
Video Manipulation LocalizationMOSE100 STTN
IoU54.9
8
Video Manipulation LocalizationMOSE100 FuseFormer
IoU56.1
8
Showing 10 of 11 rows

Other info

Follow for update