Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic Compositionality

About

Contrastively trained vision-language models like CLIP, have made remarkable progress in learning joint image-text representations, but still face challenges in compositional understanding. They often exhibit a "bag-of-words" behavior--struggling to capture the object relations, attribute-object bindings, and word order dependencies. This limitation arises not only from the reliance on global, single-vector representations for optimization, but also from the insufficient exploitation and modeling of the rich compositional information inherently present in paired image text data. In this work, we propose MACCO (MAsked Compositional Concept MOdeling), a framework that masks compositional concepts in one modality and reconstructs them conditioned on the full contextual information from the other, enabling the model to capture and align cross-modal compositional structures more effectively. To facilitate this process, we introduce two auxiliary objectives that jointly align and regularize masked features both inter-modally and intra-modally. Extensive experiments on five compositional benchmarks, along with in-depth analyses, demonstrate that our approach not only significantly enhances compositionality in VLMs but also improves their ability to capture syntactic structure and linguistic information. Additionally, the improved compositionality also benefits text-to-image generation and multimodal large language model. Code is available at https://github.com/hiker-lw/MACCO.

Wei Li, Zhen Huang, Xinmei Tian• 2026

Related benchmarks

TaskDatasetResultRank
Compositional ReasoningSugarCrepe--
95
Compositional ReasoningVALSE--
65
Compositional ReasoningVL-Checklist
Attribute Score72.1
47
Multimodal PerceptionMME
Perception Score1.45e+3
45
Compositional ReasoningWinoground
Group Score8.2
33
Multi-modal Hallucination EvaluationAMBER--
28
Vision-Language Compositional ReasoningWhat's-up
Relation Score42.4
10
Compositional ReasoningMMVP
Average Score21.5
3
Visio-linguistics CompositionalityHard Positive Compositional Benchmark (test)
Original Test Accuracy68.3
3
Visio-linguistics CompositionalityHard Positive Compositional Benchmark SWAP (test)
Test Accuracy (Original)66.1
3
Showing 10 of 11 rows

Other info

Follow for update