Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Learning multiple visual domains with residual adapters

About

There is a growing interest in learning data representations that work well for many different types of problems and data. In this paper, we look in particular at the task of learning a single visual representation that can be successfully utilized in the analysis of very different types of images, from dog breeds to stop signs and digits. Inspired by recent work on learning networks that predict the parameters of another, we develop a tunable deep network architecture that, by means of adapter residual modules, can be steered on the fly to diverse visual domains. Our method achieves a high degree of parameter sharing while maintaining or even improving the accuracy of domain-specific representations. We also introduce the Visual Decathlon Challenge, a benchmark that evaluates the ability of representations to capture simultaneously ten very different visual domains and measures their ability to recognize well uniformly.

Sylvestre-Alvise Rebuffi, Hakan Bilen, Andrea Vedaldi• 2017

Related benchmarks

TaskDatasetResultRank
Image ClassificationCIFAR-100 (test)
Accuracy79.31
3518
Image ClassificationVTAB 1K
Overall Mean Accuracy62.05
359
Image ClassificationVTAB 1k (test)
Accuracy (Natural)72.89
145
Visual Task AdaptationVTAB 1K--
95
Image ClassificationOxford Flowers (test)
Accuracy80.65
85
Image ClassificationVisual Decathlon Challenge 1.0 (test)
Mean Accuracy77.17
81
Image ClassificationFGVC
Average Accuracy88.41
78
Binary AnalysisCAPYBARA Dec-Full
ROUGE-L45.2
13
Binary AnalysisCAPYBARA XRep
ROUGE-L28.1
13
Binary AnalysisCAPYBARA Dec-Anon
Exact Match (EM)7.15
13
Showing 10 of 29 rows

Other info

Follow for update