Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

From Sparse to Soft Mixtures of Experts

About

Sparse mixture of expert architectures (MoEs) scale model capacity without significant increases in training or inference costs. Despite their success, MoEs suffer from a number of issues: training instability, token dropping, inability to scale the number of experts, or ineffective finetuning. In this work, we propose Soft MoE, a fully-differentiable sparse Transformer that addresses these challenges, while maintaining the benefits of MoEs. Soft MoE performs an implicit soft assignment by passing different weighted combinations of all input tokens to each expert. As in other MoEs, experts in Soft MoE only process a subset of the (combined) tokens, enabling larger model capacity (and performance) at lower inference cost. In the context of visual recognition, Soft MoE greatly outperforms dense Transformers (ViTs) and popular MoEs (Tokens Choice and Experts Choice). Furthermore, Soft MoE scales well: Soft MoE Huge/14 with 128 experts in 16 MoE layers has over 40x more parameters than ViT Huge/14, with only 2% increased inference time, and substantially better quality.

Joan Puigcerver, Carlos Riquelme, Basil Mustafa, Neil Houlsby• 2023

Related benchmarks

TaskDatasetResultRank
Image ClassificationOxford-IIIT Pets
Accuracy93.32
398
Image ClassificationSVHN
Top-1 Accuracy94.8
209
Image ClassificationImageNet-R (test)--
179
Image ClassificationImageNet-A (test)
Top-1 Acc6.69
177
Mortality PredictionMIMIC IV
F1-score63
154
Image ClassificationCIFAR100
Average Accuracy61.4
150
Image ClassificationCIFAR-10--
122
Semantic segmentationCamVid
mIoU71.54
110
Eye Imaging ClassificationFPRM
F1 Score87
48
Image ClassificationDomainBed
PACS Accuracy86.6
37
Showing 10 of 24 rows

Other info

Follow for update