RegimeVGGT: Layer-Wise Spatially Preserving Redundancy Removal for Visual Geometry Grounded Transformer
About
Visual Geometry Grounded Transformer (VGGT) recovers dense 3D scene structure from multi-view images in one forward pass, but quadratic cross-frame attention limits its scalability. Existing training-free accelerators reduce computation uniformly along one axis, missing layer heterogeneity. Our spectral, probing, and causal analyses reveal three regimes: shallow layers lack cross-view structure, middle layers drive cross-view alignment, and deep layers are redundant for dense geometry yet their cross-frame attention remains essential for pose. RegimeVGGT applies layer-wise U-shaped compression along two axes: Saliency-Guided Banded Merging protects geometry- and edge-salient tokens, while Selectively Protected K/V Downsampling preserves cross-frame spatial coverage and the pose-critical path through a phase-shifted spatial grid, a reference-frame anchor, and uncompressed camera/register tokens. Training-free, RegimeVGGT achieves a 6.7x speedup over VGGT* at matched reconstruction quality.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Point Cloud Reconstruction | 7 Scenes | Inference Time (s)3 | 58 | |
| Point Cloud Reconstruction | ScanNet-50 (300 frames) | Chamfer Distance (CD)0.449 | 21 | |
| Camera pose estimation | ScanNet-50 (300 frames) | ATE0.095 | 19 | |
| Camera pose estimation | ScanNet 50 (500 frames) | ATE0.096 | 19 | |
| Point Cloud Reconstruction | N-RGBD | Accuracy (Mean)2.4 | 17 | |
| Point Cloud Reconstruction | ScanNet 50 (500 frames) | Chamfer Distance (CD)0.452 | 6 | |
| Point Cloud Reconstruction | ScanNet-50 (100 frames) | CD0.43 | 6 | |
| Point Cloud Reconstruction | ScanNet 50 (1000 frames) | Chamfer Distance (CD)0.465 | 5 | |
| Camera pose estimation | Tanks and Temples (train) | Barn Accuracy93.1 | 4 | |
| Dense Reconstruction | NRGBD Stride 3 | Mean Peak Memory (GB)33.1 | 4 |