Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Dual-branch Prompting for Multimodal Machine Translation

About

Multimodal Machine Translation (MMT) typically enhances text-only translation by incorporating aligned visual features. Despite the remarkable progress, state-of-the-art MMT approaches often rely on paired image-text inputs at inference and are sensitive to irrelevant visual noise, which limits their robustness and practical applicability. To address these issues, we propose D2P-MMT, a diffusion-based dual-branch prompting framework for robust vision-guided translation. Specifically, D2P-MMT requires only the source text and a reconstructed image generated by a pre-trained diffusion model, which naturally filters out distracting visual details while preserving semantic cues. During training, the model jointly learns from both authentic and reconstructed images using a dual-branch prompting strategy, encouraging rich cross-modal interactions. To bridge the modality gap and mitigate training-inference discrepancies, we introduce a distributional alignment loss that enforces consistency between the output distributions of the two branches. Extensive experiments on the Multi30K dataset demonstrate that D2P-MMT achieves superior translation performance compared to existing state-of-the-art approaches. Our code is publicly available at https://github.com/MentaY/DDP.

Jie Wang, Zhendong Yang, Liansong Zong, Xiaobo Zhang, Dexian Wang, Ji Zhang• 2025

Related benchmarks

TaskDatasetResultRank
Machine TranslationMulti30K eng → deu 2016 (test)
BLEU43.12
31
Machine TranslationMulti30K eng → deu 2017 (test)
BLEU35.54
31
Machine TranslationMulti30K eng → fra 2017 (test)
BLEU56.62
30
Machine Translation (En-Fr)Multi30K 2016 (test)
BLEU63.7
26
Machine TranslationMSCOCO En-De
BLEU31.01
23
Machine TranslationMSCOCO En-Fr
BLEU46.23
21
Showing 6 of 6 rows

Other info

Follow for update