Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Emergent Agentic Transformer from Chain of Hindsight Experience

About

Large transformer models powered by diverse data and model scale have dominated natural language modeling and computer vision and pushed the frontier of multiple AI areas. In reinforcement learning (RL), despite many efforts into transformer-based policies, a key limitation, however, is that current transformer-based policies cannot learn by directly combining information from multiple sub-optimal trials. In this work, we address this issue using recently proposed chain of hindsight to relabel experience, where we train a transformer on a sequence of trajectory experience ascending sorted according to their total rewards. Our method consists of relabelling target return of each trajectory to the maximum total reward among in sequence of trajectories and training an autoregressive model to predict actions conditioning on past states, actions, rewards, target returns, and task completion tokens, the resulting model, Agentic Transformer (AT), can learn to improve upon itself both at training and test time. As we show on D4RL and ExoRL benchmarks, to the best our knowledge, this is the first time that a simple transformer-based model performs competitively with both temporal-difference and imitation-learning-based approaches, even from sub-optimal data. Our Agentic Transformer also shows a promising scaling trend that bigger models consistently improve results.

Hao Liu, Pieter Abbeel• 2023

Related benchmarks

TaskDatasetResultRank
Continuous ControlD4RL Hopper medium
Normalized Return70.45
24
Continuous ControlD4RL Walker2d medium
Normalized Return88.71
19
Continuous ControlD4RL Halfcheetah medium
Normalized Return45.12
10
Reinforcement LearningD4RL hopper medium-replay
Normalized Return96.85
10
Continuous ControlD4RL Walker Medium-Expert
Normalized Score114.9
5
Continuous ControlD4RL HalfCheetah Medium-Replay
Normalized Score46.86
5
Continuous ControlD4RL Walker Medium-Replay
Normalized Score92.32
5
Continuous ControlD4RL hopper-medium-expert
Normalized Score115.9
5
Continuous ControlD4RL halfcheetah-medium-expert
Normalized Score95.81
5
Showing 9 of 9 rows

Other info

Follow for update