Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

GeoMix: Descriptor-Free Visual Localization via Global Context and Multi-Detector Training

About

Descriptor-free visual localization eliminates high-dimensional descriptor storage, preserves scene privacy, and simplifies map maintenance, yet its accuracy still lags far behind descriptor-based pipelines. We identify this gap to insufficient geometric discriminability in geometry-only matching. Without visual appearance, current methods underutilize local geometry cues, lack the global context among keypoints, and overfit to a single keypoint detector. We further observe that descriptor-free matching naturally enables multi-detector training, as heterogeneous keypoints can be optimized in a shared geometry-only space without aligning descriptor spaces. Building on these insights, we propose GeoMix, a descriptor-free 2D-3D matching framework that strengthens geometric discriminability at three levels. Locally, directional and distance-aware embeddings enrich neighborhood aggregation with fine-grained spatial structure. Globally, learnable context nodes aggregate and redistribute scene-wide information via cross-attention to resolve ambiguities beyond local receptive fields. At the training level, Mix-Training exploits this detector-agnostic geometry space to learn representations across multiple keypoint detectors. Extensive experiments on MegaDepth, Cambridge Landmarks, 7Scenes, and Aachen Day-Night show that GeoMix sets a new state of the art among descriptor-free methods, reducing 75th-percentile rotation error by 89\% and translation error by up to 90\% over the previous best, while generalizing zero-shot to unseen detectors and narrowing the gap to descriptor-based pipelines. Code is available at $\href{https://github.com/YejunZhang/Geomix}{\text{this links}}$.

Yejun Zhang, Xinjue Wang, Zihan Wang, Esa Rahtu, Juho Kannala• 2026

Related benchmarks

TaskDatasetResultRank
Visual LocalizationCambridge Landmarks
King's Positional Error (cm)14
59
Visual LocalizationAachen Day-Night (day)
Recall @ (0.25m, 2°)27.8
43
Visual Localization7 Scenes
Chess Median Translation Error (cm)3
33
Visual LocalizationMegaDepth
Reprojection AUC @ 1px21.51
5
Showing 4 of 4 rows

Other info

Follow for update