Failure-Aware RL: Reliable Offline-to-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation

About

Post-training algorithms based on deep reinforcement learning can push the limits of robotic models for specific objectives, such as generalizability, accuracy, and robustness. However, Intervention-requiring Failures (IR Failures) (e.g., a robot spilling water or breaking fragile glass) during real-world exploration happen inevitably, hindering the practical deployment of such a paradigm. To tackle this, we introduce Failure-Aware Offline-to-Online Reinforcement Learning (FARL), a new paradigm minimizing failures during real-world reinforcement learning. We create FailureBench, a benchmark that incorporates common failure scenarios requiring human intervention, and propose an algorithm that integrates a world-model-based safety critic and a recovery policy trained offline to prevent failures during online exploration. Extensive simulation and real-world experiments demonstrate the effectiveness of FARL in significantly reducing IR Failures while improving performance and generalization during online reinforcement learning post-training. FARL reduces IR Failures by 73.1% while elevating performance by 11.3% on average during real-world RL post-training. Videos and code are available at https://failure-aware-rl.github.io.

Huanyu Li, Kun Lei, Sheng Zang, Kaizhe Hu, Yongyuan Liang, Bo An, Xiaoli Li, Huazhe Xu• 2026

Related benchmarks

Task	Dataset	Result
Robotic Manipulation	Average All Real-world Manipulation Tasks	Success Rate (SR)77	6
Reinforcement Learning	FailureBench Bounded Push	Average Return4.59e+3	4
Reinforcement Learning	FailureBench Bounded Soccer	Avg Return2.28e+3	4
Reinforcement Learning	FailureBench Fragile Push Wall	Avg Return3.22e+3	4
Reinforcement Learning	FailureBench Obstructed Push	Average Return1.23e+3	4
Contact-rich Robotic Manipulation	Real-world Tube Insertion	Success Rate60	4
Contact-rich Robotic Manipulation	Real-world RAM Insertion	Success Rate (SR)75	4
Contact-rich Robotic Manipulation	Real-world Wipe Whiteboard	Success Rate (SR)80	4
Non-rigid Robotic Manipulation	Real-world Fold Towel	Success Rate (SR)85	4
Multi-object Robotic Manipulation	Real-world Pick Eggplant	Success Rate (SR)85	4

Showing 10 of 10 rows

Other info

Follow for update

@wizwand_team Discord