CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features

About

Multimodal encoders like CLIP excel in tasks such as zero-shot image classification and cross-modal retrieval. However, they require excessive training data. We propose canonical similarity analysis (CSA), which uses two unimodal encoders to replicate multimodal encoders using limited data. CSA maps unimodal features into a multimodal space, using a new similarity score to retain only the multimodal information. CSA only involves the inference of unimodal encoders and a cubic-complexity matrix decomposition, eliminating the need for extensive GPU-based model training. Experiments show that CSA outperforms CLIP while requiring $50,000\times$ fewer multimodal data pairs to bridge the modalities given pre-trained unimodal encoders on ImageNet classification and misinformative news caption detection. CSA surpasses the state-of-the-art method to map unimodal features to multimodal features. We also demonstrate the ability of CSA with modalities beyond image and text, paving the way for future modality pairs with limited paired multimodal data but abundant unpaired unimodal data, such as lidar and text.

Po-han Li, Sandeep P. Chinchali, Ufuk Topcu• 2024

Related benchmarks

Task	Dataset	Result
Image Classification	CIFAR-10	Accuracy85.39	973
Text-to-Image Retrieval	Flickr30K	R@125.44	607
Classification	Cars	Accuracy1.42	571
Text-to-Image Retrieval	Flickr30k (test)	Recall@130.92	528
Image-to-Text Retrieval	Flickr30k (test)	R@144	472
Image-to-Text Retrieval	Flickr30K	R@134.8	451
Image Classification	CIFAR-100	Accuracy40.8	375
Image Classification	Pets	Accuracy8.45	320
Image Classification	Food101	Accuracy29.82	177
Text-to-Image Retrieval	COCO	--	161

Showing 10 of 27 rows

Other info

Follow for update

@wizwand_team Discord