Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Can We Really Learn One Representation to Optimize All Rewards?

About

As unsupervised pretraining becomes increasingly ubiquitous in reinforcement learning, a more thorough theoretical understanding of these methods becomes of equal importance to their empirical success. We focus on the setting of unsupervised learning via interaction, where the forward-backward (FB) representation learning serves as a prototypical and popular example. In this paper, we shed light on FB by formally contextualizing the method within a broader class of recent methods that use regression to obtain a low-rank approximation of a successor measure ratio. Our analysis clarifies when FB representations can exist and how the low-rank approximation converges in practice. Building upon the theory, we propose a variant of FB that is both more amenable to theoretical understanding and simpler to optimize in practice. Experiments in didactic settings, as well as in $10$ state-based and image-based continuous control domains, demonstrate that our method converges to desired representations with $10^5 \times$ smaller errors than FB, achieving $+24\%$ improved zero-shot performance on average. We also demonstrate that zero-shot policies inferred by our algorithm provide an efficient initialization if the user prefers further fine-tuning on downstream tasks. Our project website is available at https://chongyi-zheng.github.io/onestep-fb.

Chongyi Zheng, Royina Karegoudra Jayanth, Benjamin Eysenbach• 2026

Related benchmarks

TaskDatasetResultRank
Goal Reachingantmaze teleport-navigate v0
Success Rate42
17
Goal ReachingAntMaze Medium navigate-v0 (test-time goals)
Success Rate93
10
Goal ReachingAntMaze Large test-time goals navigate-v0
Success Rate73
10
Goal ReachingAntMaze Giant navigate test-time goals v0
Average Success Rate5
10
Goal ReachingAntMaze Teleport navigate-v0 (test-time goals)
Average Success Rate42
10
Goal-conditioned Reinforcement LearningOGBench scene play (5 tasks) zero-shot
Average Return16
10
Goal Reachingantmaze medium-navigate v0
Success Rate93
8
Goal Reachingantmaze large-navigate v0
Success Rate73
8
Goal Reachingantmaze giant-navigate v0
Success Rate5
8
Unsupervised Reinforcement LearningExORL cheetah (4 tasks) zero-shot
Average Return378
6
Showing 10 of 19 rows

Other info

Follow for update