LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
About
Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify this limitation as visual context rot: decisive evidence may exist in the full image, yet fail to be reliably selected and used amid redundant visual context. We propose LOCUS (LOcal visual CUe Search), a training framework that teaches MLLMs to internalize local evidence search through a verifiable proxy task. During training, LOCUS provides a local crop as a visual cue and optimizes the model to recover its spatial support in the full image using an IoU-based reward. The visual cue is used only during training, leaving the standard image-question inference interface unchanged. Experiments across fine-grained perception, hallucination, general understanding, and reasoning benchmarks show that LOCUS improves localization-sensitive visual understanding while preserving broad capabilities. Attention analyses further indicate stronger focus on task-relevant evidence regions, suggesting that training-time visual cue search provides an effective route to internalized fine-grained evidence selection.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Mathematical Reasoning | WeMath | Accuracy78.8 | 317 | |
| Hallucination Evaluation | POPE | Accuracy88.2 | 281 | |
| Mathematical Reasoning | MathVerse | Accuracy61.2 | 266 | |
| Logical reasoning | LogicVista | Accuracy56.7 | 163 | |
| Visual Grounding | RefCOCOg | -- | 52 | |
| Perception and Reasoning | RealworldQA | Score71 | 37 | |
| Visual Perception | OCRBench | Score85.4 | 28 | |
| Hallucination | HallusionBench | -- | 26 | |
| Perception | MMStar | Score69.7 | 23 | |
| General | AI2D | Score82.5 | 10 |