BayesFP: Posterior Estimation for Flow-Based Policies via Feynman-Kac Sampling
About
Robots must generate trajectories that remain faithful to learned expert behavior while satisfying safety constraints and task-specific objectives specified only at inference time. We formulate constrained trajectory generation for pretrained diffusion and flow-matching policies as Bayesian posterior sampling, with the learned demonstration distribution as a prior and an inference-time, cost-derived likelihood tilting it toward feasible, optimal trajectories. To sample from this posterior without any retraining of the base policy, we leverage the Feynman--Kac corrector framework, originally formulated for diffusion models, and extend it to deterministic flow-matching policies. The result is a unified, inference-time, retraining-free sampler for diffusion and flow policies. We validate the approach on pretrained Diffusion Policy, GR00T-N1.6, and $\pi_{0.5}$ checkpoints across simulated and real-world manipulation tasks, including planning around non-convex obstacles introduced at inference time, and show improvements over the base $\pi_{0.5}$ on zero-shot tasks.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| BBQSauceInBin collision avoidance | π0.5 Simulation tasks v1 (test) | Collision Rate17.9 | 5 | |
| HammersInLeftBin collision avoidance | π0.5 Simulation tasks v1 (test) | Collision Rate2.8 | 5 | |
| LIBERO-Obj. + V-shape obstacle collision avoidance | LIBERO v1 (test) | Collision Rate14 | 5 | |
| LIBERO-Obj. + cylinder obstacle collision avoidance | LIBERO v1 (test) | Collision Rate0.00e+0 | 5 | |
| FoodPacking2Cans collision avoidance | π0.5 Simulation v1 (test) | Collision Rate3 | 5 | |
| TakeMugsOffOfShelf collision avoidance | π0.5 Simulation tasks v1 (test) | Collision Rate9.9 | 5 | |
| Can + cylinder obstacle collision avoidance | Simulation v1 (test) | Collision Rate0.00e+0 | 4 | |
| Transport + cylinder obstacle collision avoidance | Simulation v1 (test) | Collision Rate4 | 4 | |
| BimodalMugSelection | Real-world SO101 arm (inference) | Correct-mug Rate100 | 2 | |
| PickAndPlace | Real-world SO101 arm (inference) | Collision Rate6 | 2 |