Embodied AI Glossary中文

PaLM-E

Common

Google's 2023 embodied multimodal large language model, which feeds images and robot-state estimates directly into PaLM.

PaLM-E is the embodied multimodal language model Google and TU Berlin released in March 2023; its largest version, PaLM-E-562B, has 562 billion parameters and combines the PaLM language model with a ViT vision encoder. It encodes continuous inputs like images and estimated robot state into vectors, interleaves them with text tokens into a single “multimodal sentence,” and feeds that into the language model. Its output is a mid-level plan in the form of text, which is then handed off to low-level skill policies to execute — the model itself doesn't output motor commands directly. The paper found positive transfer from training on web image-text data together with robot data, and the 562B version also achieved state-of-the-art results on OK-VQA visual question answering at the time. It's a landmark example of using a large model for robot task planning, and a predecessor of VLAs like RT-2.

ExampleGiven a kitchen photo and the instruction “bring me the chips from the drawer,” PaLM-E generates a sequence of substeps — go to the drawer, open the drawer, take out the chips — which a low-level policy then executes one by one.

Also called
PaLM-E-562B, PaLM-E: An Embodied Multimodal Language Model
Related
SayCan · RT-2 · Multimodal Large Language Model · LLM-based Task Planning · Vision Transformer · Language Grounding
Sources
PaLM-E 项目主页 (Chinese)
PaLM-E: An Embodied Multimodal Language Model (arXiv:2303.03378)
As of
2023-03

See it in the full glossary →