Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models

About

Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify this limitation as visual context rot: decisive evidence may exist in the full image, yet fail to be reliably selected and used amid redundant visual context. We propose LOCUS (LOcal visual CUe Search), a training framework that teaches MLLMs to internalize local evidence search through a verifiable proxy task. During training, LOCUS provides a local crop as a visual cue and optimizes the model to recover its spatial support in the full image using an IoU-based reward. The visual cue is used only during training, leaving the standard image-question inference interface unchanged. Experiments across fine-grained perception, hallucination, general understanding, and reasoning benchmarks show that LOCUS improves localization-sensitive visual understanding while preserving broad capabilities. Attention analyses further indicate stronger focus on task-relevant evidence regions, suggesting that training-time visual cue search provides an effective route to internalized fine-grained evidence selection.

Zhou Tao, Fang Zhang, Zewen Ding, Shida Wang, Xiaokun Sun, YongXiang Hua, Haoyu Cao, Linli Xu• 2026

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningWeMath
Accuracy78.8
317
Hallucination EvaluationPOPE
Accuracy88.2
281
Mathematical ReasoningMathVerse
Accuracy61.2
266
Logical reasoningLogicVista
Accuracy56.7
163
Visual GroundingRefCOCOg--
52
Perception and ReasoningRealworldQA
Score71
37
Visual PerceptionOCRBench
Score85.4
28
HallucinationHallusionBench--
26
PerceptionMMStar
Score69.7
23
GeneralAI2D
Score82.5
10
Showing 10 of 17 rows

Other info

Follow for update