Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

About

We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointing, within a single architecture toward general physical intelligence. Leveraging three automated data construction pipelines to significantly expand the data coverage of critical capabilities, we build a large-scale data system of over 15B tokens, and design a multi-task balanced RL recipe to alleviate heterogeneous task conflicts. We further introduce a Planner-Grounder-Corrector (PGC) closed-loop framework that enables a single model to autonomously execute and self-correct over long-horizon tasks. With only 8B parameters, Embodied-R1.5 achieves SOTA on 16 out of 24 embodied VLM benchmarks, surpassing leading models like Gemini-Robotics-ER-1.5 and GPT-5.4. Benefiting from the internalized embodied capabilities, Embodied-R1.5 can be fine-tuned into a VLA with only a small amount of data, outperforming leading VLA models like $\pi_{0.5}$ across 4 popular manipulation benchmark suites. We further conduct extensive zero-shot real-robot experiments, validating performance in instruction following, affordance grounding, articulated object manipulation, and long-horizon complex tasks, demonstrating strong generalization to the physical world. We open-source model weights, datasets, training code, and EmbodiedEvalKit, an evaluation framework tailored for embodied tasks, to facilitate future research in EFMs.

Yifu Yuan, Yaoting Huang, Xianze Yao, Yutong Li, Shuoheng Zhang, Linqi Han, Pengyi Li, Jiangeng Sun, Wenting Jia, Zhao Zhang, Yuhao Liu, Ruihao Liao, Yucheng Hu, Qiyu Wu, Yuxiao Li, Zibin Dong, Fei Ni, Yan Zheng, Shuyang Gu, Yi Ma, Hongyao Tang, Han Hu, Jianye Hao• 2026

Related benchmarks

TaskDatasetResultRank
Robotic ManipulationLIBERO-Plus
Language Understanding Score88
414
3D Visual GroundingScanRefer--
172
3D Dense CaptioningScan2Cap--
127
Robotic ManipulationLIBERO
Long Success Rate93.9
108
Spatial ReasoningSPAR-Bench
Overall Score40.3
59
Embodied Reasoning and Question AnsweringERQA
Score46
53
Robot ManipulationSimplerEnv WidowX Visual Matching
Average Success Rate74
52
3D Question AnsweringScanQA--
48
Visual ReasoningCV-Bench
Accuracy86.9
47
Embodied Visual Question AnsweringERQA
Accuracy46
39
Showing 10 of 57 rows

Other info

GitHub

Follow for update