Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Vision Non-Causal Trapezoidal Mamba: Eliminating Directional Scanning in Vision SSMs with Second-Order Dynamics

About

State Space Models (SSMs) have emerged as an alternative to Vision Transformers, yet most vision SSMs inherit directional token scanning from causal sequence modeling. While effective for sequential data, directional scanning introduces spatial bias and orientation-sensitive representations. We present Vision Non-Causal Trapezoidal Mamba (VNCT), a second-order non-causal vision SSM that enables all image tokens to interact in a single pass, eliminating direSctional scanning and achieving low single-image inference latency. VNCT exhibits more orientation-robust representations, showing reduced performance degradation under image rotations and flips, while improving Boundary IoU by up to 3.7 points, leading to more accurate boundary preservation and object localization. Across ImageNet-1K classification, COCO object detection and instance segmentation, and ADE20K semantic segmentation, VNCT consistently outperforms both directional-scanning vision SSMs and first-order non-causal SSMs. These results show that directional scanning is unnecessary for high-performance vision SSMs and that second-order non-causal state-space modeling offers a simple, efficient, and robust alternative for visual recognition.

Anvitha Ramachandran, Dhruv Parikh, Haoyang Fan, Rajgopal Kannan, Viktor Prasanna• 2026

Related benchmarks

TaskDatasetResultRank
Semantic segmentationADE20K (val)--
3089
Instance SegmentationMS-COCO
mAP Mask44.5
123
Object DetectionMS-COCO
APb49.7
63
Semantic segmentationADE20K
mIoU (SS)50.1
16
Inference LatencyImageNet-1K 1.0 (val)
Latency (ms)5.25
16
Semantic segmentationCityscapes (val)
BIoU66.1
4
Image ClassificationImageNet-1k (val)
Top-1 Accuracy (Original)84.2
3
Showing 7 of 7 rows

Other info

Follow for update