Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

About

Chain-of-Thought (CoT) prompting has proven highly effective for enhancing complex reasoning in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Yet, it struggles in complex spatial reasoning tasks. Nonetheless, human cognition extends beyond language alone, enabling the remarkable capability to think in both words and images. Inspired by this mechanism, we propose a new reasoning paradigm, Multimodal Visualization-of-Thought (MVoT). It enables visual thinking in MLLMs by generating image visualizations of their reasoning traces. To ensure high-quality visualization, we introduce token discrepancy loss into autoregressive MLLMs. This innovation significantly improves both visual coherence and fidelity. We validate this approach through several dynamic spatial reasoning tasks. Experimental results reveal that MVoT demonstrates competitive performance across tasks. Moreover, it exhibits robust and reliable improvements in the most challenging scenarios where CoT fails. Ultimately, MVoT establishes new possibilities for complex reasoning tasks where visual thinking can effectively complement verbal reasoning.

Chengzu Li, Wenshan Wu, Huanyu Zhang, Yan Xia, Shaoguang Mao, Li Dong, Ivan Vuli\'c, Furu Wei• 2025

Related benchmarks

TaskDatasetResultRank
Spatial decision-makingMaze size 8
Success Rate89.14
22
Visual Spatial PlanningVSP (test)
Average Accuracy12
17
Goal predictionSokoban ID
Classification Accuracy73.3
15
Goal predictionGather ID
Classification Accuracy46.7
15
Goal predictionMaze ID
Classification Accuracy83.3
15
Goal predictionGather OOD
Classification Accuracy36.7
15
Goal predictionMaze (OOD)
Classification Accuracy56.7
15
Goal predictionSokoban OOD
Classification Accuracy61.7
15
Goal predictionFrozenLake OOD
Classification Accuracy80
15
Goal predictionFrozenLake ID
Classification Accuracy86.7
15
Showing 10 of 28 rows

Other info

Follow for update