Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Mask-based Predictive Representations for Reinforcement Learning

About

Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information. It is crucial to abstract effective states from high-dimensional image inputs and limited samples for sample-efficient reinforcement learning. To address this challenge, inspired by fields such as natural language processing and computer vision, we propose a self-supervised task based on mask prediction as an auxiliary task for reinforcement learning. This non-reconstruction method uses the sequence information collected by the agent from the environment and the context information in the sequence to predict the masked information, thereby strengthening the agent's understanding of the task and learning effective representations. Combined with transformers, we find that the model reconstructs the masked input sequence in the latent space. By feeding the compressed representations learned by this method into reinforcement learning models, we observe an improvement in the sample efficiency of reinforcement learning. Moreover, the model outperforms state-of-the-art sample-efficient reinforcement learning methods on multiple continuous and discrete control benchmarks.

Kai Zhao• 2026

Related benchmarks

TaskDatasetResultRank
Reinforcement LearningAtari 100k
Alien Score1.19e+3
50
Vision-based Reinforcement LearningDMControl 100k
Finger Spin Score942
8
Vision-based Reinforcement LearningDMControl 500k
Finger Spin Score983
8
Showing 3 of 3 rows

Other info

Follow for update