Embodied AI Glossary中文

Large Language Model

大语言模型LLMEssential

A very large neural network trained on massive text that can understand and generate natural language.

A large language model is a neural network trained on massive amounts of text, generally based on the Transformer architecture, with next-token prediction — predicting the next token, the smallest unit of text a model processes — as its main training objective. At sufficient scale, these models can perform new tasks just from a few examples in the prompt, without changing their weights: OpenAI's 2020 GPT-3, with 175 billion parameters, demonstrated this few-shot ability in its paper, and after ChatGPT launched in late 2022, large language models saw widespread adoption. In robotics, they are mainly used to understand human instructions, break long tasks into steps, as in SayCan, and generate control code, as in code-as-policies. Most VLA backbones are also large language models — OpenVLA, for instance, is built on Llama 2 with a vision encoder attached.

ExampleTold “I spilled my Coke, can you bring me something to clean it up,” SayCan has a large language model score every skill the robot knows, combines that with an estimate of how likely each step is to succeed right now, and picks, in order, find the sponge, pick up the sponge, bring it over, done, which the robot then carries out step by step.

Also called
LLM, Large Model, Language Large Model
Related
Transformer · Token · Next-Token Prediction · Vision-Language Model · LLM-based Task Planning · SayCan
Sources
Language Models are Few-Shot Learners (GPT-3, arXiv:2005.14165)
Large language model (Wikipedia)
SayCan project page

See it in the full glossary →