Choose What to Observe: Task-Aware Semantic-Geometric Representations for Visuomotor Policy

About

Visuomotor policies learned from demonstrations often overfit to nuisance visual factors in raw RGB observations, resulting in brittle behavior under appearance shifts such as background changes and object recoloring. We propose a task-aware observation interface that canonicalizes visual input into a shared representation, improving robustness to out-of-distribution (OOD) appearance changes without modifying or fine-tuning the policy. Given an RGB image and an open-vocabulary specification of task-relevant entities, we use SAM3 to segment the target object and robot/gripper. We construct an L0 observation by repainting segmented entities with predefined semantic colors on a constant background. For tasks requiring stronger geometric cues, we further inject monocular depth from Depth Anything 3 into the segmented regions via depth-guided overwrite, yielding a unified semantic--geometric observation (L1) that remains a standard 3-channel, image-like input. We evaluate on RoboMimic (Lift), ManiSkill YCB grasping under clutter, four RLBench tasks under controlled appearance shifts, and two real-world Franka tasks (ReachX and CloseCabinet). Across benchmarks and policy backbones (Flow Matching Policy and SmolVLA), our interface preserves in-distribution performance while substantially improving robustness under OOD visual shifts.

Haoran Ding, Liang Ma, Yaxun Yang, Wen Yang, Tianyu Liu, Anqing Duan, Xiaodan Liang, Dezhen Song, Ivan Laptev, Yoshihiko Nakamura• 2026

Related benchmarks

Task	Dataset	Result
open box	RLBench	Success Rate62	13
CloseGrill	RLBench OOD1	Success Rate81.3	6
CloseGrill	RLBench OOD2	Success Rate82	6
CloseMicrowave	RLBench ID	Success Rate96.7	6
CloseMicrowave	RLBench OOD1	Success Rate92	6
CloseMicrowave	RLBench OOD2	Success Rate91.3	6
CloseGrill	RLBench ID	Success Rate82.7	6
CloseGrill	RLBench OOD3	Success Rate82	3
CloseMicrowave	RLBench OOD3	Success Rate89.3	3
OpenBox	RLBench OOD1	Success Rate62	3

Showing 10 of 30 rows

Other info

Follow for update

@wizwand_team Discord