Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

The Challenges of Continuous Self-Supervised Learning

About

Self-supervised learning (SSL) aims to eliminate one of the major bottlenecks in representation learning - the need for human annotations. As a result, SSL holds the promise to learn representations from data in-the-wild, i.e., without the need for finite and static datasets. Instead, true SSL algorithms should be able to exploit the continuous stream of data being generated on the internet or by agents exploring their environments. But do traditional self-supervised learning approaches work in this setup? In this work, we investigate this question by conducting experiments on the continuous self-supervised learning problem. While learning in the wild, we expect to see a continuous (infinite) non-IID data stream that follows a non-stationary distribution of visual concepts. The goal is to learn a representation that can be robust, adaptive yet not forgetful of concepts seen in the past. We show that a direct application of current methods to such continuous setup is 1) inefficient both computationally and in the amount of data required, 2) leads to inferior representations due to temporal correlations (non-IID data) in some sources of streaming data and 3) exhibits signs of catastrophic forgetting when trained on sources with non-stationary data distributions. We propose the use of replay buffers as an approach to alleviate the issues of inefficiency and temporal correlations. We further propose a novel method to enhance the replay buffer by maintaining the least redundant samples. Minimum redundancy (MinRed) buffers allow us to learn effective representations even in the most challenging streaming scenarios composed of sequential visual data obtained from a single embodied agent, and alleviates the problem of catastrophic forgetting when learning from data with non-stationary semantic distributions.

Senthil Purushwalkam, Pedro Morgado, Abhinav Gupta• 2022

Related benchmarks

TaskDatasetResultRank
Class-incremental learningCIFAR-100 20 tasks--
58
Class-incremental learningImageNet-100 20 tasks--
16
Class-incremental classificationCIFAR100 50 tasks
Accuracy43.62
14
Online Continual Self-Supervised LearningCLEAR100 11 experiences (streaming online)
Final Accuracy51.3
9
Online Continual Self-Supervised LearningImageNet100 streaming online 20 experiences
Final Accuracy48
9
Online Continual Self-Supervised LearningCIFAR-100 streaming online 20 experiences
Final Accuracy46.5
9
Class-incremental classificationCIFAR-100 tasks
Class Accuracy (CA)38.3
7
Class-incremental classificationImageNet 100 tasks
CA35.07
7
Class-incremental classificationImageNet-100 50 tasks
Cumulative Accuracy34.83
7
ClassificationCIFAR-100 Split (20 tasks)
Classification Accuracy38.82
6
Showing 10 of 13 rows

Other info

Follow for update