See & Sniff: Learning Visuo-Olfactory Representations
About
While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. We introduce SmellNet-V, a scalable visuo-olfactory dataset built on the insight that odor identity is largely invariant to visual transformations within a semantic category. This allows us to synthetically pair smell-only samples with semantically aligned in-the-wild web images, converting a unimodal olfactory dataset into a cross-modal benchmark without costly co-collection. Building on this dataset, we propose See & Sniff, a self-supervised framework that learns joint visuo-olfactory representations via dense local alignment and naturally produces smell saliency maps for spatial grounding of odor sources. We further introduce pixel-level smell localization task and a benchmark for evaluation. Our method surpasses smell-only baselines by 7% in smell classification from smell alone and generalizes to cross-modal retrieval and smell localization, establishing visuo-olfactory learning as a new direction in multimodal perception.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Smell Classification | SmellNet smell-only (test) | Accuracy63.75 | 12 | |
| Smell Localization | SmellNet-V-Source | mIoU0.6456 | 11 | |
| Interactive Localization | SmellNet InteractiveSource V (test) | IIoU35 | 7 | |
| Family-wise Classification | SmellNet V (unseen ingredients) | Accuracy52.97 | 4 | |
| Smell-to-Vision Retrieval | SmellNet-V (test) | R@156.14 | 4 | |
| Vision-to-Smell Retrieval | SmellNet-V (test) | Recall@163.2 | 4 | |
| Smell-to-Smell Retrieval | SmellNet (test) | R@169.44 | 2 |