Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Scaling Language-Image Pre-training via Masking

About

We present Fast Language-Image Pre-training (FLIP), a simple and more efficient method for training CLIP. Our method randomly masks out and removes a large portion of image patches during training. Masking allows us to learn from more image-text pairs given the same wall-clock time and contrast more samples per iteration with similar memory footprint. It leads to a favorable trade-off between accuracy and training time. In our experiments on 400 million image-text pairs, FLIP improves both accuracy and speed over the no-masking baseline. On a large diversity of downstream tasks, FLIP dominantly outperforms the CLIP counterparts trained on the same data. Facilitated by the speedup, we explore the scaling behavior of increasing the model size, data size, or training length, and report encouraging results and comparisons. We hope that our work will foster future research on scaling vision-language learning.

Yanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer, Kaiming He• 2022

Related benchmarks

TaskDatasetResultRank
Object Hallucination EvaluationPOPE
Accuracy86.5
2056
Visual Question AnsweringGQA
Accuracy60.9
1445
Image ClassificationImageNet-1K
Top-1 Acc61.3
1239
Image ClassificationImageNet 1k (test)
Top-1 Accuracy86.9
939
Multimodal EvaluationMME--
902
Multimodal UnderstandingMMBench--
887
Image ClassificationImageNet V2
Top-1 Acc66.8
767
Image ClassificationImageNet A
Top-1 Acc71.9
723
Visual Question AnsweringVQA v2 (test-dev)
Overall Accuracy74.7
721
Image ClassificationStanford Cars
Accuracy90.9
705
Showing 10 of 119 rows
...

Other info

Code

Follow for update