Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

About

Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched. Existing vision-language CBMs often rely on pre-aligned encoders or global cosine similarity, which obscures fine-grained concept localization and fails to reflect true semantic geometry. In this work, we rethink concept alignment as a dynamic cross-modal transport process instead of static projection and propose the Optimal Transport Flow Concept Bottleneck Model (OTF-CBM). It first learns a data-driven semantic cost via Inverse Optimal Transport to measure cross-modal distances, and then performs unbalanced optimal-transport-based flow matching to model semantic transitions between visual patches and textual concepts. With velocity-based concept activation, OTF-CBM captures interpretable geometric relations without ODE integration. Experiments further show that OTF-CBM achieves superior classification accuracy and concept faithfulness, offering a new geometric and dynamical perspective for interpretable cross-modal reasoning.

Chenyang Zhang, Anqi Dong, Guangming Zhu, Nuoye Xiong, Siyuan Wang, Lin Mei, Liang Zhang• 2026

Related benchmarks

TaskDatasetResultRank
ClassificationCUB
Accuracy89.92
100
ClassificationCIFAR100
Accuracy90.21
90
Image ClassificationPlaces365--
79
ClassificationImageNet
Accuracy85.62
43
ClassificationAWA2
Class Accuracy98.88
41
Image ClassificationCUB (in-distrib.)
Top-1 Accuracy89.9
17
Image ClassificationCUB (Out-Of-Domain)
Accuracy82
7
Image ClassificationPlaces365 In-Domain
Accuracy55.1
7
Image ClassificationPlaces365 (Out-Of-Domain)
Accuracy50.5
7
Showing 9 of 9 rows

Other info

Follow for update