Vesta: A Generalist Embodied Reasoning Model
About
Robots operating in open-world environments must seamlessly integrate localization, spatial reasoning, navigation, and long-horizon planning. While specialist models excel at individual tasks, deploying a multi-model stack is computationally expensive and prone to cascading errors. We present Vesta, a unified embodied generalist that consolidates these capabilities into a single foundation model. Our approach combines a diverse and massive curated corpus designed to induce spatial grounding and a simple multimodal memory harness that enables reasoning over extended time horizons. Across diverse benchmarks, Vesta on average beats individual SOTA baselines by >$20\%$ and beats an ensemble of per-category-best baselines by $>10\%$ -- thus demonstrating that a generalist model can match or exceed specialists. On real-world robotic tasks requiring memory and reasoning, Vesta improves task success by >35\%. Our work thus demonstrates that a single generalist is a feasible, scalable, and arguably preferable alternative to combining specialists.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Vision-Language Navigation | R2R-CE (val-unseen) | Success Rate (SR)55.5 | 779 | |
| Embodied Reasoning and Question Answering | ERQA | Score44.9 | 53 | |
| Embodied Cognition | Open-X VQA | Score89.3 | 4 | |
| Embodied Cognition | SAT | Score81.3 | 4 | |
| Embodied Cognition | MMSI-Bench | Score40.8 | 4 | |
| Embodied Cognition | MindCube (tiny) | Score80.9 | 4 | |
| Embodied Cognition | CV-Bench | Score88.1 | 4 | |
| Embodied Cognition | PAI-U | Score57.9 | 4 | |
| Embodied Localization | CrossPoint | Score76 | 4 | |
| Embodied Localization | EmbSpatial | Score81.9 | 4 |