Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

G2G: Exploiting Intra-Group Geometry for Inter-Group Pose Estimation

About

Recovering the relative 6-DoF pose between two image groups underlies cross-sequence relocalization and multi-camera rig odometry. Each group carries known intra-group geometry from visual odometry or rig calibration, and pretrained multi-view backbones already fuse such geometry into visual features. Yet current models treat all views as an unstructured set, leaving cross-group reasoning as the missing piece. We introduce \ours{}, which keeps the foundation model entirely frozen and adds three lightweight trainable modules to bridge the two groups: a perceiver resampler, a cross-group bridge with merged self-attention, and a multi-frame pose head. The trainable footprint totals about 32M parameters, under 6\% of the full model, and is supervised only by relative poses. Across four datasets that span indoor and outdoor simulation, real-world cross-season capture, and zero-shot sim-to-real transfer, \ours{} attains state-of-the-art accuracy on both tasks, while every baseline is retrained with its full original supervision. Code is available at https://github.com/WeiYuFei0217/G2G.

Yufei Wei, Shuhao Ye, Chenxiao Hu, Yiyuan Pan, Dongyu Feng, Rong Xiong, Yue Wang, Yanmei Jiao• 2026

Related benchmarks

TaskDatasetResultRank
Multi-Camera Rig OdometryHM3D 8-camera rig
Translation Error (m)0.402
6
Multi-Camera Rig OdometryHM3D 4-camera rig
Translation Error (m)0.218
6
Multi-Camera Rig OdometryTartanGround
Translation Error (m)0.311
6
Multi-Camera Rig OdometryNCLT intra-date
Translation Error (m)0.078
6
Multi-Camera Rig OdometryNCLT cross-date
Translation Error (m)0.172
6
Multi-Camera Rig OdometryZJH simulation
Translation Error (m)0.131
6
Rig odometryHM3D 8-cam
mAA@3075.4
6
Rig odometryHM3D 4-cam
mAA@3067.4
6
Rig odometryTartanGround (TG) 4-cam
mAA@3088.5
6
Rig odometryNCLT (intra)
mAA@3091.6
6
Showing 10 of 14 rows

Other info

Follow for update