VIMA: General Robot Manipulation with Multimodal Prompts

About

Prompt-based learning has emerged as a successful paradigm in natural language processing, where a single general-purpose language model can be instructed to perform any task specified by input prompts. Yet task specification in robotics comes in various forms, such as imitating one-shot demonstrations, following language instructions, and reaching visual goals. They are often considered different tasks and tackled by specialized models. We show that a wide spectrum of robot manipulation tasks can be expressed with multimodal prompts, interleaving textual and visual tokens. Accordingly, we develop a new simulation benchmark that consists of thousands of procedurally-generated tabletop tasks with multimodal prompts, 600K+ expert trajectories for imitation learning, and a four-level evaluation protocol for systematic generalization. We design a transformer-based robot agent, VIMA, that processes these prompts and outputs motor actions autoregressively. VIMA features a recipe that achieves strong model scalability and data efficiency. It outperforms alternative designs in the hardest zero-shot generalization setting by up to $2.9\times$ task success rate given the same training data. With $10\times$ less training data, VIMA still performs $2.7\times$ better than the best competing variant. Code and video demos are available at https://vimalabs.github.io/

Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, Linxi Fan• 2022

Related benchmarks

Task	Dataset	Result
Robotic Manipulation	VIMA-Bench 1.0 (test)	L1 Score81.5	14
Robotic Manipulation	VIMA-Bench	Task 1 Score99.3	13
Robotic task execution	Simulation Tasks Averaged (test)	Success Rate (SR)30	10
Action Generation	VIMA-Bench	Success Rate72.6	5
Robot manipulation generalization	VIMA-Bench	Novel Task48.8	5

Showing 5 of 5 rows

Other info

Follow for update

@wizwand_team Discord