Embodied AI Glossary中文

12Landmark Models & Projects

Famous models in historical order: from SayCan and RT-2 to the π series, GR00T, and world models. · 332 terms

  1. 12.1Foundation models embodied AI borrows11
  2. 12.2LLMs as the brain27
  3. 12.3End-to-end learning and the RT series16
  4. 12.4Classic imitation-learning policies21
  5. 12.5Open-source generalist policies and VLA32
  6. 12.6The π series and reinforcement learning28
  7. 12.7Global tech giants and star startups35
  8. 12.8Embodied models in China45
  9. 12.9World models and learning from video43
  10. 12.10Legged, dexterous-hand, and agile skills16
  11. 12.11Humanoid whole-body control and teleop37
  12. 12.12Navigation and autonomous driving21

12.1Foundation models embodied AI borrows

An embodied model’s ‘eyes’ and ‘brain’ are often borrowed: meet these foundation models and vision encoders first.

12.2LLMs as the brain

The earliest use of large models was as the brain: breaking down tasks, writing code and rewards, understanding space, then calling existing skills.

12.3End-to-end learning and the RT series

The other path is learning control end to end: from real-robot grasping to generalist models, up to RT-2, which coined VLA.

12.4Classic imitation-learning policies

Smaller imitation-learning models were evolving in parallel: from PerAct to Diffusion Policy and ACT, mastering fine bimanual work.

12.5Open-source generalist policies and VLA

RT and imitation learning converge into open-source VLA: after Octo and OpenVLA, improvements bloomed in every direction.

12.6The π series and reinforcement learning

Physical Intelligence’s π series set the VLA benchmark; then how reinforcement learning makes policies stronger from experience.

12.7Global tech giants and star startups

From technical approaches to companies: the foundation models of NVIDIA, Figure, Google, and other firms outside China.

12.8Embodied models in China

Now China: foundation models from big tech, research institutes, and startups, mostly split between VLA and world-action-model approaches.

12.9World models and learning from video

Back to the world-model thread: training policies inside imagination first, then generating worlds and learning actions from video.

12.10Legged, dexterous-hand, and agile skills

From manipulation to movement: quadrupeds doing parkour, dexterous hands spinning objects, trained with RL in sim, then deployed to hardware.

12.11Humanoid whole-body control and teleop

The same methods carried over to humanoids: from animated-character imitation and bipedal walking to whole-body teleop and motion tracking.

12.12Navigation and autonomous driving

Finally, moving through the world: navigation from modular pipelines to foundation models, and driving from ALVINN to VLA.

See it in the full glossary →