Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Robot Critics that Sweat the Small Stuff

About

Large vision-language models contain several priors about the world and object interactions, making them useful critics during inference to steer robot policies towards success. However, closed-loop robot manipulation requires judging small visual differences between success and failure, which remains a challenge for current VLMs. We introduce a method to fine-tune critics by constructing pairwise progress supervision using success and failure rollouts obtained from a policy. Our fine-tuned critic excels at fine-grained progress reasoning and subtle failure detection, outperforming prior progress reasoning baselines. Additionally, we use an action-conditioned video model to predict the visual effect of several candidate actions sampled from a policy, and show that our critic can correctly identify successful candidates to execute, improving the average policy success rate by 11% across real-world tasks and 5.9% across simulation tasks.

Sruthi Sudhakar, Junbang Liang, Sreehari Rammohan, Pavel Tokmakov, Richard Zemel, Carl Vondrick• 2026

Related benchmarks

TaskDatasetResultRank
StackingReal-world
Success Rate32
11
Critic AccuracyRoboCasa Overall v1
Cab Countr Accuracy96.6
7
Critic AccuracyLBM dataset
BimanBike Accuracy98
4
Critic AccuracyRoboCasa OOD Only v1
Countr Cab Accuracy97
3
PickPlace Lego To BowlReal-world
Success Rate8
2
Pickup LegoReal-world
Success Rate16
2
Push-BowlReal-world
mIoU32.1
2
Robotic ManipulationRoboCasa365 simulation atomic tasks
Navigate Kitchen Success Rate4
2
Showing 8 of 8 rows

Other info

Follow for update