Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Difference-Aware Retrieval Policies for Imitation Learning

About

Parametric imitation learning via behavior cloning can suffer from poor generalization to out-of-distribution states due to compounding errors during deployment. We show that reusing the training data during inference via a semi-parametric retrieval-based imitation learning approach can alleviate this challenge. We present Difference-Aware Retrieval Policies for Imitation Learning (DARP), a semi-parametric retrieval-based imitation learning approach that addresses this limitation by reparameterizing the imitation learning problem in terms of local neighborhood structure rather than direct state-to-action mappings. Instead of learning a global policy, DARP trains a model to predict actions based on $k$-nearest neighbors from expert demonstrations, their corresponding actions, and the relative distance vectors between neighbor states and query states. DARP requires no additional assumptions beyond those made for standard behavior cloning -- it does not require additional data collection, online expert feedback, or task-specific knowledge. We demonstrate consistent performance improvements of 15-46% over standard behavior cloning across diverse domains, including continuous control and robotic manipulation, and across different representations, including high-dimensional visual features. Code and demos are available at https://weirdlabuw.github.io/darp-site/.

Quinn Pfeifer, Ethan Pronovost, Paarth Shah, Khimya Khetarpal, Siddhartha Srinivasa, Abhishek Gupta• 2026

Related benchmarks

TaskDatasetResultRank
Robotic ManipulationRoboCasa--
68
LocomotionMuJoCo Hopper Low-Dimensional State v1
Score3.55e+3
7
LocomotionMuJoCo Ant Low-Dimensional State v1
Locomotion Score4.38e+3
7
LocomotionMuJoCo Walker Low-Dimensional State v1
Total Score4.89e+3
7
LocomotionMuJoCo HalfCheetah Low-Dimensional State v1
Score5.52e+3
7
Robotic ManipulationRoboSuite
Stack Success Rate75
4
Action ModelingPush T
Score70
2
Reinforcement LearningMuJoCo Hopper
Improvement over BC53.2
2
Reinforcement LearningMuJoCo Ant
Improvement over BC84.5
2
Reinforcement LearningMuJoCo HalfCheetah
Improvement over BC418.7
2
Showing 10 of 12 rows

Other info

Follow for update