Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

VigilFormer: Deformable Attention for Video Anomaly Detection with Causal Risk Inference

About

Video anomaly detection in surveillance settings must balance detection accuracy against real-time throughput, a tension that existing methods address either through stronger feature extractors or more efficient architectures, but rarely both. We present VigilFormer, a unified framework that combines deformable spatio-temporal attention with causal temporal modeling to detect anomalies in untrimmed surveillance video. The proposed Deformable Spatio-Temporal Encoder (DSTE) attends to a sparse set of informative locations across frames, avoiding the quadratic cost of dense attention while retaining the ability to capture irregular motion patterns. A Causal Anomaly Classifier (CAC) applies dilated causal convolutions over snippet-level features and optimizes a contrastive multiple-instance learning objective that separates anomalous and normal representations without frame-level labels. To meet deployment constraints, an Adaptive Confidence Scheduler (ACS) dynamically skips low-information frames at inference time, reducing redundant computation in static scenes. Evaluated on UCF-Crime, ShanghaiTech, and CUHK Avenue, VigilFormer achieves AUC scores of 87.83%, 97.21%, and 89.74% respectively, at 41.5 FPS on a single GPU, outperforming recent weakly-supervised methods in both accuracy and speed.

Xinze Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Video Anomaly DetectionUCF-Crime
AUC87.83
288
Video Anomaly DetectionCUHK Avenue
Frame AUC89.74
75
Video Anomaly DetectionShanghaiTech
ROC AUC0.9721
61
Showing 3 of 3 rows

Other info

Follow for update