UIA-ViT: Unsupervised Inconsistency-Aware Method based on Vision Transformer for Face Forgery Detection

About

Intra-frame inconsistency has been proved to be effective for the generalization of face forgery detection. However, learning to focus on these inconsistency requires extra pixel-level forged location annotations. Acquiring such annotations is non-trivial. Some existing methods generate large-scale synthesized data with location annotations, which is only composed of real images and cannot capture the properties of forgery regions. Others generate forgery location labels by subtracting paired real and fake images, yet such paired data is difficult to collected and the generated label is usually discontinuous. To overcome these limitations, we propose a novel Unsupervised Inconsistency-Aware method based on Vision Transformer, called UIA-ViT, which only makes use of video-level labels and can learn inconsistency-aware feature without pixel-level annotations. Due to the self-attention mechanism, the attention map among patch embeddings naturally represents the consistency relation, making the vision Transformer suitable for the consistency representation learning. Based on vision Transformer, we propose two key components: Unsupervised Patch Consistency Learning (UPCL) and Progressive Consistency Weighted Assemble (PCWA). UPCL is designed for learning the consistency-related representation with progressive optimized pseudo annotations. PCWA enhances the final classification embedding with previous patch embeddings optimized by UPCL to further improve the detection performance. Extensive experiments demonstrate the effectiveness of the proposed method.

Wanyi Zhuang, Qi Chu, Zhentao Tan, Qiankun Liu, Haojie Yuan, Changtao Miao, Zixiang Luo, Nenghai Yu• 2022

Related benchmarks

Task	Dataset	Result
Part Segmentation	ShapeNetPart	mIoU (Instance)86	254
Deepfake Detection	DFDC	AUC71.84	230
Deepfake Detection	DFD	AUC0.947	193
Deepfake Detection	DFDC (test)	AUC89.98	130
Deepfake Detection	CDF v2	AUC0.8241	97
Classification	ScanObjectNN	OA89.3	77
Deepfake Detection	Celeb-DF	ROC-AUC0.8047	48
Frame-level Deepfake Detection	DFD	AUC94.68	42
Deepfake Detection	FF++ video-level 8 (test)	Accuracy95.71	40
object recognition	ModelNet40 5-way	Accuracy97.2	40

Showing 10 of 35 rows

Other info

Follow for update

@wizwand_team Discord