Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Transferability Between Understanding and Generation in Unified Multimodal Models

About

Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact remains understudied. We investigate $\boldsymbol{\mathsf{transferability}}$ in UMMs: whether training a capability on one task improves the same capability on the other without explicit supervision. Through controlled experiments, we empirically find that transferability depends on architecture-models with fully shared transformer backbone and a unified visual encoder exhibit consistent cross-task transfer, while loosely coupled designs show little or none. Leveraging this transferability, we propose a practical training strategy. The most straightforward way to improve a target generative capability (e.g., counting) is to fine-tune generation directly, but this can degrade visual quality due to distribution shift. Instead, we train the corresponding understanding task and let it transfer into generation, which improves capability-specific generative performance while minimizing distribution shift. We validate this across three capabilities-counting, spatial relation, and text recognition/generation-showing that cross-task transferability can be systematically exploited in UMMs.

Jiwon Kang, Heeji Yoon, Jaewoo Jung, Jaewon Min, Minkyeong Jeon, Biyeon Hwang, Sangwon Jung, Seungryong Kim• 2026

Related benchmarks

TaskDatasetResultRank
Multimodal UnderstandingMMMU (test)--
112
Multimodal UnderstandingMMBench (test)
Accuracy90.9
71
Multimodal UnderstandingPOPE (test)
Accuracy87.3
4
Multimodal UnderstandingMME-P (test)
Perception Score1.55e+3
4
Showing 4 of 4 rows

Other info

Follow for update