Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images

About

Recent multimodal large language models (MLLMs) have begun to support Thinking with Images by invoking visual tools such as zooming and cropping during inference. Yet these systems remain brittle in fine-grained visual reasoning because they must decide where to look before they have access to the evidence needed to make that decision correctly. We identify this circular dependency as the Grounding Paradox. To address it, we propose Test-Time Scaling over Perception (TTSP), a framework that treats perception itself as a scalable inference process. TTSP generates multiple exploratory perception traces, filters unreliable traces using entropy-based confidence estimation, distills validated observations into structured knowledge, and iteratively refines subsequent exploration toward unresolved uncertainty. Extensive experiments on high-resolution and general multimodal reasoning benchmarks show that TTSP consistently outperforms strong baselines across backbone sizes, while also exhibiting favorable scalability and token efficiency. Our results suggest that scaling perception at test time is a promising direction for robust multimodal reasoning under perceptual uncertainty.

Zheng Jiang, Yiming Chen, Nan He, Jiahui Chen, Chaoyang Li, Houde Qian, Lifeng Sun• 2026

Related benchmarks

Task	Dataset	Result
Visual Grounded Reasoning	TreeBench	Overall Score51.8	162
Visual Perception and Reasoning	V*Bench	Attribute Score95.7	49
Perception	MME-RealWorld-Lite	Overall Score59	46
High-Resolution Multimodal Reasoning	HR-Bench-8K	Overall Score86.4	40
High-Resolution Multimodal Reasoning	HR-Bench-4K	Overall Score88.3	40
Reasoning	MME-RealWorld-Lite	OCR Score84	37
Visual Question Answering	VisualProbe Easy	Accuracy67.4	9
Visual Question Answering	VisualProbe Medium	Accuracy41.4	9
Visual Question Answering	VisualProbe Hard	Accuracy47.2	9
Visual Question Answering	VisualProbe (Overall)	Accuracy49.7	9

Showing 10 of 10 rows

Other info

Follow for update

@wizwand_team Discord