Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Preference-Aligned Reinforcement Learning on MetaWorld

21,000Label Score

PrefVLM

-3205,21510,75016,285Jul 2, 2026
Updated 23d ago

Evaluation Results

MethodLinks
2026.07
21,0002,000
2026.07
5,00010,000
2026.07
4,9004,900
2026.07
5004,000