Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Diffusion Policies for Risk-Averse Behavior Modeling in Offline Reinforcement Learning

About

Offline reinforcement learning (RL) presents distinct challenges as it relies solely on observational data. A central concern in this context is ensuring the safety of the learned policy by quantifying uncertainties associated with various actions and environmental stochasticity. Traditional approaches primarily emphasize mitigating epistemic uncertainty by learning risk-averse policies, often overlooking environmental stochasticity. In this study, we propose an uncertainty-aware distributional offline RL method to simultaneously address both epistemic uncertainty and environmental stochasticity. We propose a model-free offline RL algorithm capable of learning risk-averse policies and characterizing the entire distribution of discounted cumulative rewards, as opposed to merely maximizing the expected value of accumulated discounted returns. Our method is rigorously evaluated through comprehensive experiments in both risk-sensitive and risk-neutral benchmarks, demonstrating its superior performance.

Xiaocong Chen, Siyu Wang, Tong Yu, Lina Yao• 2024

Related benchmarks

TaskDatasetResultRank
Offline Reinforcement LearningStochastic-D4RL Hopper (medium-expert)
Mean Reward577.7
14
Offline Reinforcement LearningStochastic-D4RL Walker2d medium-expert
Mean Return823.2
14
Offline Reinforcement LearningStochastic-D4RL Hopper-medium-replay
Mean Return25.68
14
Offline Reinforcement LearningStochastic-D4RL HalfCheetah (medium-expert)
Mean Return650.7
14
Offline Reinforcement LearningStochastic-D4RL HalfCheetah medium-replay
Mean Return220
14
Offline Reinforcement LearningStochastic-D4RL Walker2d medium-replay
Mean Return4.62
14
Offline Reinforcement LearningD4RL Half-Cheetah risk-sensitive Medium
CVaR 0.1276
7
Offline Reinforcement LearningD4RL Half-Cheetah risk-sensitive Expert
CVaR (0.1)732
7
Offline Reinforcement Learningrisk-sensitive D4RL Half-Cheetah Mixed
CVaR 0.1275
7
Offline Reinforcement Learningrisk-sensitive D4RL Walker-2D Medium
CVaR 0.11.37e+3
7
Showing 10 of 17 rows

Other info

Follow for update