Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Prompting Diffusion Models for Zero-Shot Instance Segmentation

About

Several disruptive research directions have recently emerged in computer vision, including foundation models achieving previously unseen zero-shot performance in scene understanding, even interactively, and generative models that synthesize extremely realistic images. The latter have also been shown to be highly effective in scene understanding tasks thanks to their rich priors. However, for promptable segmentation, foundation models struggle with accurately segmenting an object's region, leading to false positives and over-segmentation. Notably, early attempts that leverage generative priors use prompts only during post-processing, yielding suboptimal segments because the process is agnostic to the user input. In this paper, we target these limitations with Prompt2Seg, a spatial conditioning framework for diffusion-based segmentation. Prompt2Seg augments a frozen diffusion segmentation model with a conditioning branch. Our approach takes spatial prompts, represented as 2D Gaussians or confidence maps, as explicit input signals, training the model to respond directly to user intent. Fine-tuned on a deliberately constrained set of object categories drawn from Hypersim and Virtual KITTI 2, Prompt2Seg generalizes zero-shot to a wide range of unseen object types and visual domains. We evaluate on seven datasets ranging from standard benchmarks to more challenging domains, including paintings, egocentric views, and X-ray data. Furthermore, we demonstrate that Prompt2Seg consistently outperforms the underlying diffusion segmentation backbone across all benchmarks. Our results suggest that the rich priors encoded in generative pretraining, combined with principled spatial conditioning, offer a compelling path toward broadly generalizing interactive segmentation without large-scale mask supervision.

Irem Zeynep Alag\"oz, Nils Morbitzer, Andrea Ramazzina, Nassir Navab, Federico Tombari, Stefano Gasperini• 2026

Related benchmarks

TaskDatasetResultRank
Interactive SegmentationPascal VOC--
48
Instance SegmentationEgoHOS
mIoU55.6
13
Interactive Instance SegmentationCOCO Large
mIoU62.3
7
Interactive Instance SegmentationDRAM
mIoU53.1
7
Interactive Instance SegmentationPIDRay
mIoU47.5
7
Interactive Instance SegmentationZeroWaste
mIoU51.4
7
Interactive Instance SegmentationCOCO Medium
mIoU49.7
7
Interactive Instance SegmentationCOCO Small
mIoU12.1
7
Interactive Instance SegmentationHRSOD
mIoU66.6
7
Showing 9 of 9 rows

Other info

Follow for update