Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Latent Visual States for Efficient Multimodal Reasoning

About

The integration of visual evidence has significantly enhanced the capabilities of large multimodal models. However, this integration predominantly relies on generating discrete outputs (etc., code or box coordinates) to invoke external tools, a process that introduces rigid dependencies and substantial latency. To overcome these limitations, we propose {EVA} (LatEnt Visual StAtes), a novel framework that natively generates continuous latent visual representations. These internal representations manifest as an adaptive sequence of Latent\_slot tokens, serving as intermediate visual thoughts during the reasoning process. These Latent\_slot tokens are then trained end-to-end with the discrete text tokens. This co-optimization, notably, causes extreme policy deviation in the 'transition window' following the Latent\_slot tokens. We develop D-GSPO (Decouple-GSPO) to target this root cause by decoupling the optimization of latent and discrete components. To support SFT, we construct EVA-230K, a high-quality text-image interleaved CoT dataset encompassing a diverse range of real-world scenes, documents, charts and OCR tasks. Extensive experiments across multiple benchmarks confirm that EVA achieves significant performance gains while enhancing inference efficiency.

Xiuwei Chen, Wentao Hu, Yongxin Wang, Zisheng Chen, Likui Zhang, Kun Xiang, Jianhua Han, Hui-Ling Zhen, Jingyuan Zou, Hang Xu, Xiaodan Liang• 2026

Related benchmarks

TaskDatasetResultRank
Visual ReasoningJigsaw
Accuracy66.7
44
Multimodal UnderstandingMME-RealWorld-Lite
Overall Score49.8
38
Visual ReasoningV* Attribute, Spatial, Overall
Overall Accuracy80.2
6
Multimodal PerceptionHRbench 4K FSP FCP Overall
Overall Score73.7
5
Multimodal PerceptionHRbench-8K FSP FCP Overall
Overall Score68.4
5
Multimodal UnderstandingMME-Real Perception, Reasoning, Overall
Perception Score63.9
4
Vision-centric ReasoningIQ (test)
Accuracy30
4
Vision-centric ReasoningRelative Reflect
Accuracy39.6
4
Showing 8 of 8 rows

Other info

Follow for update