Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space

About

Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step perception and logical deduction. While explicit visual generation incurs prohibitive computational costs, existing latent approaches often rely on external experts or lack rigorous cognitive logic. In this paper, we introduce ProLaViT (Progressive Latent Visual Thought), a framework empowering MLLMs to perform structured visual derivation in the continuous latent space. Unlike works dependent on heterogeneous external models, ProLaViT leverages an endogenous self-distillation mechanism, utilizing the model's own visual encoder to supervise latent thoughts. To facilitate this, we construct a scalable programmatic synthesis pipeline enabling the model to internalize algorithmic precision without inference time tools. We design two reasoning paradigms: (1) Coarse-to-Fine Causal Chain for spatial tasks, guiding attention from global context to local targets. (2) Dialectical Reasoning Chain for logical tasks, incorporating counter-factual thinking for verification. Furthermore, we propose a Distance-Weighted Diversity Loss to impose topology-aware constraints, preventing feature degeneration by enforcing semantic distinctiveness. Extensive experiments demonstrate that ProLaViT outperforms baselines on vision-centric benchmarks, achieving superior accuracy and interpretability with high efficiency.

Peiming Li, Yifan Wang, Xiaotian Zhang, Zhiyuan Hu, Shiyu Li, Zheng Wei, Yang Tang• 2026

Related benchmarks

TaskDatasetResultRank
Chart Understanding and ReasoningChartQA
Accuracy78.99
143
Visual ReasoningBLINK
Accuracy57.49
116
Visual ReasoningMMVP
Accuracy79
67
Visual ReasoningCV-Bench
Accuracy78.33
47
Visual ReasoningBLINK-J
Accuracy76
23
Visual ReasoningVisualPuzzles
Overall Score75.25
12
Visual ReasoningVSTAR
VStar Accuracy80.11
9
Showing 7 of 7 rows

Other info

Follow for update