RoboFlamingo
AdvancedAn early piece of work that fine-tuned the open-source OpenFlamingo vision-language model directly into a robot manipulation policy.
RoboFlamingo was released in November 2023 by ByteDance Research together with Tsinghua University, Shanghai Jiao Tong University, and the National University of Singapore, one of the earlier efforts to turn an open-source vision-language model (VLM) directly into a manipulation policy. It uses OpenFlamingo as its backbone, which at each step understands the current image and the language instruction, followed by an explicit policy head (such as an LSTM) that aggregates history and outputs the robot arm's actions; it is fine-tuned with imitation learning purely on language-annotated demonstration data. This split between 'understanding' and 'decision-making' let it be trained on a single 8-GPU server. On the CALVIN long-horizon benchmark, it completed an average of 4.09 consecutive tasks, clearly ahead of prior methods, though the paper only validated it in simulation. The same authors later expanded this idea into the systematic study RoboVLMs.
ExampleGiven five consecutive instructions in CALVIN (such as 'open the drawer' and 'push the blue block to the left'), RoboFlamingo could on average complete about 4 of them in a row.
- Also called
- Vision-Language Foundation Models as Effective Robot Imitators
- Related
- Vision-Language Model · Vision-Language-Action Model · CALVIN Benchmark · Imitation Learning · RoboVLMs · Flamingo
- Sources
- arXiv 2311.01378: Vision-Language Foundation Models as Effective Robot Imitators
RoboFlamingo 项目主页 (Chinese)
GitHub: RoboFlamingo - As of
- 2024-02