Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of existing vision-language models (VLMs) with simple fine-tuning on robotics data. To this end, we derive a simple and novel vision-language manipulation framework, dubbed RoboFlamingo, built upon the open-source VLMs, OpenFlamingo. Unlike prior works, RoboFlamingo utilizes pre-trained VLMs for single-step vision-language comprehension, models sequential history information with an explicit policy head, and is slightly fine-tuned by imitation learning only on language-conditioned manipulation datasets. Such a decomposition provides RoboFlamingo the flexibility for open-loop control and deployment on low-performance platforms. By exceeding the state-of-the-art performance with a large margin on the tested benchmark, we show RoboFlamingo can be an effective and competitive alternative to adapt VLMs to robot control. Our extensive experimental results also reveal several interesting conclusions regarding the behavior of different pre-trained VLMs on manipulation tasks. We believe RoboFlamingo has the potential to be a cost-effective and easy-to-use solution for robotics manipulation, empowering everyone with the ability to fine-tune their own robotics policy.

通过对开放源代码的视觉-语言模型进行简单微调，RoboFlamingo构建了一个简单而新颖的视觉-语言操控框架，并利用单步视觉-语言理解的预训练模型、显式策略推测历史信息，通过模仿学习在以语言为条件的操纵数据集上微调。通过在基准测试上超过最先进的性能，表明RoboFlamingo能够有效并具有竞争力地将VLM适应到机器人控制中，为机器人操作提供了一种具有潜力的经济高效和易于使用的解决方案。

视觉语言基础模型作为有效的机器人模仿者