Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LaST-HD: Learning Latent Physical Reasoning from Scalable Human Data for Robot Manipulation

About

Human-hand demonstrations provide a direct and scalable source of physical interaction data for robot learning. While manual retargeting is indispensable for establishing kinematic action correspondence across different morphologies, robust transfer requires going beyond geometry to address the underlying alignment of physical dynamics between human and robot manipulation. To address this, we introduce LaST-HD, a novel human-to-robot action learning paradigm that extends reasoning-before-acting VLA by aligning human-hand and robot demonstrations in a shared latent reasoning space. Rather than mimicking human kinematics, LaST-HD trains an auxiliary action-conditioned world model on unpaired human-hand and robot trajectories to synthesize unified latent targets. After aligning cross-embodiment representations in this shared forward-dynamics space, these targets supervise LaST-HD's latent reasoning process, enabling it to internalize shared physical dynamics and drive efficient human-hand action learning. Moreover, we develop Out-of-Lab (OOL) Glove, a low-cost motion-capture glove tailored to LaST-HD for human-hand data collection. The captured human data provide precise keypoints and serve as universal action supervision across grippers and dexterous hands. Armed with the aligned latent space and high-fidelity human-hand data, we develop a progressive mixed-to-human training recipe comprising mixed human-robot co-training and human-hand online correction post-training. Through mixed co-training, LaST-HD improves generalization to novel objects, scenes, and positions using only human-hand demonstrations. With online correction, LaST-HD further adapts to novel environments and achieves over 90\% accuracy using only 20 minutes of OOL glove data.

Jiaming Liu, Yinxi Wang, Chenyang Gu, Siyuan Qian, Xiangju Mi, Hao Chen, Jiawei Chen, Qingpo Wuwu, Xiaoqi Li, Nuowei Han, Yiming Zhang, Xuheng Zhang, Yang Yue, Yeqing Yang, Lei Wang, Peng Jia, Hao Tang, Shanghang Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Robotic ManipulationReal-robot manipulation tasks Aggregate
Average Success Rate (Avg SR)73
19
Robotic ManipulationUnSeen-Object
Unscrew Cap Success Rate75
6
Robotic ManipulationUnseen Position
Unscrew Cap Success Rate30
6
Robotic ManipulationUnseen Background
Unscrew Cap Success Rate65
6
Pour WaterTianji Marvin + WUJI In-domain
Success Rate60
5
Put Items to Bag and ZipTianji Marvin In-domain
Success Rate80
5
Sort FruitsTianji Marvin In-domain
Success Rate95
5
Unscrew bottle capGalaxea In-domain R1 Lite
Success Rate85
5
Grasp with a ClampTianji Marvin + WUJI In-domain
Success Rate45
5
Organize BoxGalaxea R1 Lite In-domain
Success Rate70
5
Showing 10 of 10 rows

Other info

Follow for update