Bridging Vision and Language Concepts through Optimal Transport Semantic Flow
About
Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness fundamentally depends on how well visual and textual representations are aligned or matched. Existing vision-language CBMs often rely on pre-aligned encoders or global cosine similarity, which obscures fine-grained concept localization and fails to reflect true semantic geometry. In this work, we rethink concept alignment as a dynamic cross-modal transport process instead of static projection and propose the Optimal Transport Flow Concept Bottleneck Model (OTF-CBM). It first learns a data-driven semantic cost via Inverse Optimal Transport to measure cross-modal distances, and then performs unbalanced optimal-transport-based flow matching to model semantic transitions between visual patches and textual concepts. With velocity-based concept activation, OTF-CBM captures interpretable geometric relations without ODE integration. Experiments further show that OTF-CBM achieves superior classification accuracy and concept faithfulness, offering a new geometric and dynamical perspective for interpretable cross-modal reasoning.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Classification | CUB | Accuracy89.92 | 100 | |
| Classification | CIFAR100 | Accuracy90.21 | 90 | |
| Image Classification | Places365 | -- | 79 | |
| Classification | ImageNet | Accuracy85.62 | 43 | |
| Classification | AWA2 | Class Accuracy98.88 | 41 | |
| Image Classification | CUB (in-distrib.) | Top-1 Accuracy89.9 | 17 | |
| Image Classification | CUB (Out-Of-Domain) | Accuracy82 | 7 | |
| Image Classification | Places365 In-Domain | Accuracy55.1 | 7 | |
| Image Classification | Places365 (Out-Of-Domain) | Accuracy50.5 | 7 |