Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Reinforcement Learning on HalfCheetah v4
Loading...
10,554
Max Return
ReLU
1,611.04
3,932.77
6,254.5
8,576.23
Jan 29, 2026
Feb 22, 2026
Mar 18, 2026
Apr 11, 2026
May 5, 2026
May 29, 2026
Jun 22, 2026
Max Return
Updated 1mo ago
Evaluation Results
Method
Method
Links
Max Return
ReLU
Frame=ANN
2026.01
10,554
C-DSAC
Number of runs=100, Se...
2026.04
10,023
S-PLIF
Frame=Vanilla
2026.01
9,828
S-PLIF
Frame=PT
2026.01
9,405
PLIF
Frame=Vanilla
2026.01
9,252
PLIF
Frame=PT
2026.01
9,219
PDA
Environment steps=1M,...
2026.03
5,174.6
TRPO
Environment steps=1M,...
2026.03
4,496.8
PPO
Environment steps=1M,...
2026.03
4,067.5
NPG
Environment steps=1M,...
2026.03
3,556.4
Adaptive β
Number of seeds=5, Eva...
2026.06
2,207
Fixed β
Number of seeds=5, Eva...
2026.06
2,111
PPO-Clip
Number of seeds=5, Eva...
2026.06
1,955
per-sample PPO-KL
Number of seeds=5, Eva...
2026.06
1,955
Feedback
Search any
task
Search any
task