Embodied AI Glossary中文

11Models & Architectures

What the resulting models look like: Transformers, VLMs, how actions are generated, VLAs, and world models. · 184 terms

  1. 11.1Neural network basics16
  2. 11.2From attention to language models24
  3. 11.3Vision encoders and vision-language models28
  4. 11.4Introduction to generative models10
  5. 11.5Diffusion models and flow matching22
  6. 11.6How actions are represented, output17
  7. 11.7VLA structure and extensions17
  8. 11.8Hierarchical systems and embodied reasoning12
  9. 11.9World models and video generation23
  10. 11.10Inference speed, deployment, reliability15

11.1Neural network basics

Starting with the basic building blocks and classic architectures of neural networks, which every later model is built from.

11.2From attention to language models

The shared backbone of large models: tokens, attention, and the Transformer, leading up to large language models.

11.3Vision encoders and vision-language models

Giving a language model eyes: first how images become tokens, then VLMs, the foundation VLAs are built on.

11.4Introduction to generative models

Earlier models output a single answer; here, models learn a whole data distribution, plus turning vectors into discrete codes.

11.5Diffusion models and flow matching

The leading approach among generative models: adding and removing noise, and flow matching, which power both action and video generation.

11.6How actions are represented, output

With networks and generative methods in hand, see how robot actions are represented, and output by regression, discretization, or generation.

11.7VLA structure and extensions

Attach action output to a VLM and you get a VLA: how it outputs actions, generalizes across robots, and takes in more sensors.

11.8Hierarchical systems and embodied reasoning

The opposite of end-to-end: splitting planning and control across different models, then having a model reason before it acts.

11.9World models and video generation

Using generative models to predict ‘what happens to the world after an action,’ finally merging with action generation itself.

11.10Inference speed, deployment, reliability

Every model here eventually runs on hardware: how to make it faster and cheaper, and how to tell when it’s unsure.

See it in the full glossary →