Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards

About

Most unified large multimodal models (LMMs) that support both visual understanding and image generation still rely on curated post-training supervision, such as human annotations, preference labels, or external reward models. We ask whether a unified LMM can improve both abilities autonomously using only unlabeled images. We propose a self-evolving training framework with three internal roles: a Proposer that generates visual questions, a Solver that answers and evaluates them, and a Generator that synthesizes images. Training uses only self-derived consistency signals, without human annotations, preference labels, or task-trained external reward/judge models. To stabilize learning, we introduce Solver Token Entropy (STE), a continuous difficulty signal based on token-level prediction uncertainty that remains useful even when sample-level consistency becomes unreliable. For image generation, we design a multi-scale internal evaluation scheme that combines question-answer fidelity scoring with cycle-consistent captioning. This creates a solver-mediated coupling, where better visual understanding enables more reliable generation assessment and stronger internal training signals. The framework preserves the same role decomposition, reward logic, and training schedule across diffusion-based BLIP3o, rectified-flow BAGEL, and autoregressive VARGPT-v1.1 architectures, requiring only each backbone's native prompting and generation interface. Across eight understanding metrics, our method consistently improves over the corresponding base models. On BAGEL, it achieves a $+3.5\%$ absolute gain on MMMU and improves GenEval image generation performance from $82\%$ to $85\%$. Code and models are publicly released.

Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar, Abdelrahman Shaker, Fahad Khan, Hisham Cholakkal, Salman Khan, Rao Muhammad Anwer• 2026

Related benchmarks

TaskDatasetResultRank
Multimodal UnderstandingMM-Vet
MM-Vet Score69.5
664
Real-world Visual Question AnsweringRealworldQA
Accuracy73.9
183
Text-based Visual Question AnsweringTextVQA
TextVQA Accuracy88.5
141
Image GenerationGenEval
Overall GenEval Score87
132
Multi-discipline ReasoningMMMU
Accuracy58.8
83
Multi-modal UnderstandingMMBench
Accuracy87.1
51
Multimodal UnderstandingSEED
SEED Score81.8
51
Multi-modal EvaluationMME
MME Perception Score1.70e+3
43
Showing 8 of 8 rows

Other info

Follow for update