Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Pose Anything Anywhere:Model-free Object Poses from Arbitrary References

About

Estimating the 6D pose of unseen objects is a fundamental yet challenging problem for open-world robotics and embodied perception. Model-based methods are accurate but depend on CAD assets or heavy onboarding, while most model-free approaches are still limited to pairwise single-anchor matching and thus fail under occlusion and large viewpoint changes with low query-reference overlap. Therefore, we present PANY, a unified model-free framework that seamlessly supports both RGB and RGB-D inputs, operates on one or sparse pose-free reference views, and generalizes effectively to novel objects. Built on a multi-view transformer geometry backbone, PANY moves beyond pairwise matching by learning view-consistent geometry and cross-view alignment cues that remain stable under wide baselines and limited overlap. When additional unposed assist views are available, PANY aggregates them via pose-graph canonical registration to increase geometric coverage and reinforce the final pose. Extensive experiments show that PANY achieves state-of-the-art performance across multiple benchmarks, substantially outperforming existing model-free methods, improving pose accuracy by +12% on YCB-V and over +20% on LM-O. Furthermore, PANY consistently performs well under both single-reference and sparse-reference settings, demonstrating strong robustness in real-world environments.

Hongli Xu, Jiaqi Hu, Junwen Huang, Boyang Zhong, Peter KT Yu, Nassir Navab, Benjamin Busam, Slobodan Ilic• 2026

Related benchmarks

TaskDatasetResultRank
6D Object Pose EstimationLM-O (test)--
22
6D Object Pose EstimationReal275 52
AR81.8
11
6D Object Pose EstimationToyota-Light 16
AR51.1
11
6D Pose EstimationYCB-Video
Overall ADD AUC94.6
7
Pose EstimationLineMOD
ADD-0.1d55.3
7
Showing 5 of 5 rows

Other info

Follow for update