Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Grounding Generative Policies in Physics: Optimization-Guided Diffusion for Robot Control

About

Diffusion models sample effectively from high-dimensional, multimodal distributions, but their outputs may violate deployment constraints. For task-space robot policies, generated grasps, waypoints, or trajectories can be distributionally valid yet infeasible, violating reachability, collision-avoidance, or closed-loop executability requirements. This embodiment gap limits zero-shot deployment across robots, even when the task-space behavior itself is transferable. We propose an inference-time optimization framework that couples the behavior generation to physical feasibility by formulating diffusion guidance as a constrained optimization problem. Our key insight is to replace the sampling perturbation in the backward process with an optimized correction, allowing hard constraints or soft penalties to be imposed during sampling without the need to retrain the diffusion model, while keeping samples close to the learned prior. We evaluate the method on dexterous grasp synthesis with reachability and collision-avoidance constraints, and dynamic manipulation with controller-level trackability constraints. Across settings and robot embodiments, optimization-guided denoising matches the feasibility of projection- and gradient-guidance baselines while better preserving grasp quality, and improving controller-level executability and task success, with task success improving by up to 20pp. on dexterous grasping and 23pp. on visuomotor manipulation over the best baseline.

Sabrina Bodmer, Ren\'e Zurbr\"ugg, Tifanny Portela, Hao Ma, Alexandre Didier, Marco Hutter, Colin Jones, Melanie Zeilinger• 2026

Related benchmarks

TaskDatasetResultRank
Dexterous Grasp Synthesis30 objects
Success Rate (SR(1))71
11
Grasp Prediction with Collision AvoidanceFloor environment
Success Rate (Total)65.62
10
Grasp Prediction with Collision AvoidanceWalls environment
SRTot60.62
10
Grasp Prediction with Collision AvoidanceClutter environment
SRTot55.63
10
Grasp Prediction with Collision AvoidanceTunnels environment
SRTot61.88
10
Tabletop pick-and-placeFranka
Task Success Rate (SRtask)67
5
Drawer OpeningDynaarm
Task Success Rate81.8
5
Drawer OpeningFranka
Task Success Rate60
5
Tabletop pick-and-placeDynaarm
SR (Task)41.8
5
Showing 9 of 9 rows

Other info

Follow for update