Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment

About

Personalized object detection aims to adapt a general-purpose detector to recognize user-specific instances from only a few examples. Lightweight models often struggle in this setting due to their weak semantic priors, while large vision-language models (VLMs) offer strong object-level understanding but are too computationally demanding for real-time or on-device applications. We introduce MOCHA (Multi-modal Objects-aware Cross-arcHitecture Alignment), a distillation framework that transfers multimodal region-level knowledge from a frozen VLM teacher into a lightweight vision-only detector. MOCHA extracts fused visual and textual teacher's embeddings and uses them to guide student training through a dual-objective loss that enforces accurate local alignment and global relational consistency across regions. This process enables efficient transfer of semantics without the need for teacher modifications or textual input at inference. MOCHA consistently outperforms prior baselines across four personalized detection benchmarks under strict few-shot regimes, yielding a +10.1 average improvement, with minimal inference cost.

Elena Camuffo, Francesco Barbato, Mete Ozay, Simone Milani, Umberto Michieli• 2025

Related benchmarks

TaskDatasetResultRank
Personalized Object DetectionPerSeg 1-shot
mAP59.1
13
Personalized Object DetectionPOD 5-shot
mAP45.9
13
Personalized Object DetectionCORe50 1-shot
mAP60.9
13
Personalized Object DetectionCORe50 5-shot
mAP70.6
13
Personalized Object DetectioniCubWorld 1-shot
mAP62.8
13
Personalized Object DetectioniCubWorld 5-shot
mAP77.6
13
Personalized Object DetectionPOD 1-shot
Accuracy36.3
5
Personalized Object DetectionPOD 5-shot
Accuracy45.9
5
Personalized SegmentationPerSeg 1-shot
Accuracy59.1
5
Showing 9 of 9 rows

Other info

Follow for update