Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

See & Sniff: Learning Visuo-Olfactory Representations

About

While modern multimodal models integrate vision with language, audio, or touch, olfaction remains largely unexplored due to the lack of paired visuo-olfactory data. We introduce SmellNet-V, a scalable visuo-olfactory dataset built on the insight that odor identity is largely invariant to visual transformations within a semantic category. This allows us to synthetically pair smell-only samples with semantically aligned in-the-wild web images, converting a unimodal olfactory dataset into a cross-modal benchmark without costly co-collection. Building on this dataset, we propose See & Sniff, a self-supervised framework that learns joint visuo-olfactory representations via dense local alignment and naturally produces smell saliency maps for spatial grounding of odor sources. We further introduce pixel-level smell localization task and a benchmark for evaluation. Our method surpasses smell-only baselines by 7% in smell classification from smell alone and generalizes to cross-modal retrieval and smell localization, establishing visuo-olfactory learning as a new direction in multimodal perception.

Seongyu Kim, Seungwoo Lee, Hyeonggon Ryu, Joon Son Chung, Arda Senocak• 2026

Related benchmarks

TaskDatasetResultRank
Smell ClassificationSmellNet smell-only (test)
Accuracy63.75
12
Smell LocalizationSmellNet-V-Source
mIoU0.6456
11
Interactive LocalizationSmellNet InteractiveSource V (test)
IIoU35
7
Family-wise ClassificationSmellNet V (unseen ingredients)
Accuracy52.97
4
Smell-to-Vision RetrievalSmellNet-V (test)
R@156.14
4
Vision-to-Smell RetrievalSmellNet-V (test)
Recall@163.2
4
Smell-to-Smell RetrievalSmellNet (test)
R@169.44
2
Showing 7 of 7 rows

Other info

Follow for update