Embodied AI Glossary中文

RoboFlamingo

Advanced

An early piece of work that fine-tuned the open-source OpenFlamingo vision-language model directly into a robot manipulation policy.

RoboFlamingo was released in November 2023 by ByteDance Research together with Tsinghua University, Shanghai Jiao Tong University, and the National University of Singapore, one of the earlier efforts to turn an open-source vision-language model (VLM) directly into a manipulation policy. It uses OpenFlamingo as its backbone, which at each step understands the current image and the language instruction, followed by an explicit policy head (such as an LSTM) that aggregates history and outputs the robot arm's actions; it is fine-tuned with imitation learning purely on language-annotated demonstration data. This split between 'understanding' and 'decision-making' let it be trained on a single 8-GPU server. On the CALVIN long-horizon benchmark, it completed an average of 4.09 consecutive tasks, clearly ahead of prior methods, though the paper only validated it in simulation. The same authors later expanded this idea into the systematic study RoboVLMs.

ExampleGiven five consecutive instructions in CALVIN (such as 'open the drawer' and 'push the blue block to the left'), RoboFlamingo could on average complete about 4 of them in a row.

Also called
Vision-Language Foundation Models as Effective Robot Imitators
Related
Vision-Language Model · Vision-Language-Action Model · CALVIN Benchmark · Imitation Learning · RoboVLMs · Flamingo
Sources
arXiv 2311.01378: Vision-Language Foundation Models as Effective Robot Imitators
RoboFlamingo 项目主页 (Chinese)
GitHub: RoboFlamingo
As of
2024-02

See it in the full glossary →