Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers

About

We propose Confidence-Guided Token Merging (Co-Me), an acceleration mechanism for visual geometric transformers without retraining or finetuning the base model. Co-Me distilled a light-weight confidence predictor to rank tokens by uncertainty and selectively merge low-confidence ones, effectively reducing computation while maintaining spatial coverage. Compared to similarity-based merging or pruning, the confidence signal in Co-Me reliably indicates regions emphasized by the transformer, enabling substantial acceleration without degrading performance. Co-Me applies seamlessly to various multi-view and streaming visual geometric transformers, achieving speedups that scale with sequence length. When applied to VGGT and Pi3, Co-Me achieves up to 21.5x and 20.4x speedup, making visual geometric transformers practical for real-time 3D perception and reconstruction.

Yutian Chen, Yuheng Qiu, Ruogu Li, Ali Agha, Shayegan Omidshafiei, Jay Patrikar, Sebastian Scherer• 2025

Related benchmarks

TaskDatasetResultRank
3D Reconstruction7 Scenes
Completion15.8
161
Camera pose estimationCO3D v2
AUC@3086.15
132
3D ReconstructionETH3D
F1 Score54.8
50
Pose EstimationHiRoom
AUC@328.33
47
Camera pose estimationScanNet++
AUC @ 30°88.76
32
Camera pose estimation7Scenes
AUC@30.1619
32
Camera pose estimationRE10K
AUC@3075.25
30
Point Cloud ReconstructionScanNet-50 (300 frames)
Chamfer Distance (CD)0.45
21
Camera pose estimationScanNet 50 (500 frames)
ATE0.15
19
Camera pose estimationScanNet-50 (300 frames)
ATE0.16
19
Showing 10 of 22 rows

Other info

Follow for update