Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

RATS! Patches Talk Through Registers: Emergent Parts in Register Attention Transformers

About

When humans see a bird, they recognize far more than just "bird" -- they see a head, wings, and talons, a structured assembly of reusable parts that can be identified across every bird they have ever seen. We ask whether a self-supervised visual model can discover the same compositional structure on its own. To this end, we propose RATS (Register Attention Transformers), which decomposes the classification token into N learnable register tokens that route patch information through an L->N->N->L bottleneck via a three-step compress-communicate-broadcast attention. The N registers are partitioned across the H attention heads, so that registers assigned to different heads do not interact with each other. Without auxiliary losses or part annotations, each register spontaneously specializes into a proto-semantic region whose emerging structure resembles object parts. RATS surpasses all baselines by +12 mIoU on average across five segmentation benchmarks, with consistent gains on ADE20K (+1.11 mIoU) and COCO (+0.2 AP^m). Its register dictionary further exhibits part-level consistency and semantic proximity across related categories. Our results suggest that RATS may provide a useful architectural prior for structured and interpretable visual representation learning.

Timing Yang, Predrag Neskovic, Jansen Seheult, Wenchao Han, Anand Bhattad, Alan Yuille, Feng Wang• 2026

Related benchmarks

TaskDatasetResultRank
Semantic segmentationADE20K (val)
mIoU46.52
3089
Object DetectionCOCO 2017 (val)--
2930
Instance SegmentationCOCO 2017 (val)
APm0.381
1304
Part SegmentationCOCO
M2O mIoU19.32
13
Part SegmentationADE20K
M2O mIoU30.77
13
Part SegmentationImageNet (IN)
mIoU (M2O)23.47
13
Part SegmentationIN-S919
M2O mIoU75.39
13
Part SegmentationPartImageNet (PartIN)
M2O mIoU48.77
13
Showing 8 of 8 rows

Other info

Follow for update