Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Omnimodal Dataset Distillation via High-order Proxy Alignment

About

Dataset distillation compresses large-scale datasets into compact synthetic sets while preserving training performance, but existing methods are largely restricted to single-modal or bimodal settings. Extending dataset distillation to scenarios involving more than two modalities, i.e., Omnimodal Dataset Distillation, remains underexplored and challenging due to increased heterogeneity and complex cross-modal interactions. In this work, we identify the key determinant that bounds the endpoint discrepancy in the omnimodal setting, which is exacerbated with an increasing number of modalities. To this end, we propose HoPA, a unified method that captures high-order cross-modal alignments via a compact proxy, which is compatible with trajectory matching as well. By abstracting omnimodal alignment with a shared similarity structure, our method avoids the combinatorial complexity of pairwise modality modeling and enables scalable joint distillation across heterogeneous modalities. Theoretical analysis from the spectral perspective reveals the rationality of our proposed method against bimodal dataset distillation techniques. Extensive experiments on various benchmarks demonstrate that the proposed method achieves superior compression-performance trade-offs compared to existing competitors. The source code will be publicly released.

Yuxuan Gao, Xiaohao Liu, Xiaobo Xia, Tongliang Liu• 2026

Related benchmarks

TaskDatasetResultRank
Text-to-Video RetrievalDiDeMo (test)
R@131.7
399
Video-to-Text retrievalDiDeMo (test)
R@131.8
111
Cross-modal retrievalMSR-VTT (test)
R@1 (V→T)37.3
19
Average RetrievalDiDeMo (test)
R@118.6
19
Multimodal RetrievalVGGSound-S (test)
Recall@1 (Video -> Text)6.7
19
Text-to-Audio RetrievalDiDeMo (test)
R@15.4
19
Video-to-Audio RetrievalDiDeMo (test)
R@119.8
19
Audio-to-Text RetrievalDiDeMo (test)
R@15.1
19
Audio-to-Video RetrievalDiDeMo (test)
R@118.6
19
Showing 9 of 9 rows

Other info

Follow for update