EO-1: An Open Unified Embodied Foundation Model for General Robot Control
About
The human ability to seamlessly perform multimodal reasoning and physical interaction in the open world is a core goal for general purpose embodied intelligent systems. Recent vision-language-action (VLA) models, which are co-trained on large-scale robot and visual-text data, have demonstrated notable progress in general robot control. However, they still fail to achieve human-level flexibility in interleaved reasoning and interaction. In this work, we introduce EO-Robotics, consists of EO-1 model and EO-Data1.5M dataset. EO-1 is a unified embodied foundation model that achieves superior performance in multimodal embodied reasoning and robot control through interleaved vision-text-action pre-training. The development of EO-1 is based on two key pillars: (i) a unified architecture that processes multimodal inputs indiscriminately (image, text, video, and action), and (ii) a massive, high-quality multimodal embodied reasoning dataset, EO-Data1.5M, which contains over 1.5 million samples with emphasis on interleaved vision-text-action comprehension. EO-1 is trained through synergies between auto-regressive decoding and flow matching denoising on EO-Data1.5M, enabling seamless robot action generation and multimodal embodied reasoning. Extensive experiments demonstrate the effectiveness of interleaved vision-text-action learning for open-world understanding and generalization, validated through a variety of long-horizon, dexterous manipulation tasks across multiple embodiments. This paper details the architecture of EO-1, the data construction strategy of EO-Data1.5M, and the training methodology, offering valuable insights for developing advanced embodied foundation models. Project Page: https://eo-robotics.ai/eo-1.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Robot Manipulation | LIBERO | Object Achievement99.8 | 1025 | |
| Robotic Manipulation | LIBERO | Spatial Success Rate99.7 | 570 | |
| Robotic Manipulation | LIBERO | Long-horizon Success Rate94.8 | 165 | |
| Robot Manipulation | SimplerEnv WidowX | Overall Success Rate72.7 | 123 | |
| Robot Manipulation | SimplerEnv Google Robot Visual Matching | Pick Coke Can98 | 79 | |
| Robot Manipulation | SimplerEnv WidowX Visual Matching | Average Success Rate72.7 | 52 | |
| Robotic Manipulation | SimplerEnv Google Robot - Visual Aggregation | Pick Coke Can91.6 | 28 | |
| Robotic Manipulation | RLBench 10 tasks | Take Umbrella Success Rate76 | 20 | |
| Embodied Pointing and Visual Question Answering | Pointing and VQA Benchmarks Suite | -- | 14 | |
| Trajectory Prediction | Trajectory RoboInter Gripper, RoboInter Traj., ShareRobot Bench-T, VA Bench-V | Error (RoboInter Gripper)0.148 | 12 |