Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Pano3D: Unified 3D Reconstruction and Panoptic Segmentation

About

Recent advances in 3D feedforward reconstruction neural networks have achieved remarkable success in dense reconstruction from images without any camera parameters. Yet, equipping these models with robust semantic understanding remains an open problem. Here we introduce an approach that performs 3D reconstruction and 3D panoptic segmentation in a unified framework. We build on existing 3D reconstruction models and augment them with a set-based mask decoder. The approach is jointly trained with a geometric and semantic loss, which are shown to be mutually beneficial. More precisely, the features are initialized from the geometric information and then finetuned to capture jointly geometry and semantics. We demonstrate the generality of our approach by successfully applying our framework both to online and all-to-all attention reconstruction backbones. Our method achieves state-of-the-art performance in 3D panoptic segmentation across ScanNet, ScanNet200, and ScanNet++ datasets. Ablation studies show that such joint training of a unified model equips 3D feedforward reconstruction neural networks with panoptic segmentation and yields mutually beneficial improvements.

Victor Barberteguy, Ahmet Iscen, Mathilde Caron, Alireza Fathi, G\"ul Varol, Cordelia Schmid• 2026

Related benchmarks

TaskDatasetResultRank
3D Semantic SegmentationScanNet
mIoU65.3
57
3D Semantic SegmentationScanNet++
mIoU (20 classes)29.7
42
Class-agnostic 3D instance segmentationScanNet200
mAP14.6
24
Class-agnostic 3D instance segmentationScanNet++
AP6.9
21
Semantic segmentationScanNet++--
15
3D ReconstructionScanNet++
AbsRel2.44
8
3D Semantic SegmentationScanNet200 open-vocabulary
3D mIoU24.6
6
Class-agnostic 3D instance segmentationScanNet
AP15.7
4
Showing 8 of 8 rows

Other info

Follow for update