ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation
About
Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active subset of classes. We introduce ActiveSAM, a training-free, zero-shot inference framework that turns SAM 3 into an active-vocabulary segmenter. ActiveSAM first canonicalizes and expands class prompts, then estimates an image-conditioned active set from a low-resolution presence preview. Only the retained classes are decoded at full resolution, using bucketed prompt multiplexing with the frozen SAM 3 decoder. The preview stage uses only class-presence evidence and skips unnecessary segmentation-head computation, while the final stage applies margin-aware background calibration to suppress low-confidence pixels. ActiveSAM requires no target-dataset training, no weight updates, and no oracle class-presence labels. Across eight OVSS benchmarks, ActiveSAM improves the speed-accuracy tradeoff of training-free open-vocabulary semantic segmentation, outperforming the current state-of-the-art SegEarth-OV3 by approximately +1.4 mIoU on average while running up to 5.5x faster on large-vocabulary datasets. ActiveSAM also demonstrates the strongest robustness under image corruption that simulates real-world distribution shift, making it well-suited for deployment in noisy-input domains such as autonomous driving and embodied AI. Code is available at https://github.com/VILA-Lab/ActiveSAM.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Open Vocabulary Semantic Segmentation | ADE20K | mIoU40 | 110 | |
| Open Vocabulary Semantic Segmentation | Cityscapes | mIoU69.7 | 105 | |
| Open Vocabulary Semantic Segmentation | COCO Stuff | mIoU44.4 | 84 | |
| Open Vocabulary Semantic Segmentation | Pascal Context 60 | mIoU55.4 | 65 | |
| Open-Vocabulary Segmentation | LoveDA | mIoU41.1 | 33 | |
| Open-Vocabulary Segmentation | iSAID | mIoU28.3 | 33 | |
| Open-Vocabulary Segmentation | Vaihingen | mIoU45.1 | 33 | |
| Open-Vocabulary Segmentation | Potsdam | mIoU40.6 | 33 | |
| Open Vocabulary Semantic Segmentation | COCO Object | mIoU72 | 30 | |
| Open Vocabulary Semantic Segmentation | Pascal VOC 2012 V21 | mIoU83.8 | 28 |