Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Improving Multimodal Reasoning via Worst Dimension Optimization

About

Multimodal reasoning requires a path that retains integrity over a wide range of constraints, from visual grounding to logic consistency. However, the current Process Reward Models focus on heuristically defined rewards that equally weigh these factors, which may lead to the concealment of individual dimension failures by the dominating factors, without guaranteeing the validity of the reasoning process in general.

Haocheng Lv, Huaping Zhang, Qiuchi Li, Lei Li, Chunxiao Gao• 2026

Related benchmarks

TaskDatasetResultRank
Chart Understanding and ReasoningChartQA
Accuracy87.2
143
Multimodal ReasoningMathVista
Accuracy67.5
89
Multimodal ReasoningMMMU
Accuracy54.2
77
Multimodal Chain-of-Thought ReasoningM3CoT
Accuracy79.7
53
Diagram UnderstandingAI2D
Accuracy84.2
29
Multimodal Mathematical ReasoningMathVista
Accuracy67.5
18
Fine-grained visual understandingMMStar
Accuracy65.2
17
multimodal reasoning across university-level disciplinesMMMU
Accuracy54.2
12
Showing 8 of 8 rows

Other info

Follow for update