Embodied AI Glossary中文

Latent Action

潜在动作Common

An abstract action code learned automatically from how consecutive video frames change, with no real action labels needed.

A latent action is an action representation a model infers automatically from the change between adjacent video frames, usually a handful of discrete codes or a single low-dimensional vector. It targets a specific problem: there is a huge amount of human and gameplay video on the internet, but none of it carries real action labels like joint angles or button presses, so it can't be used to train a policy directly. A latent action doesn't correspond to any specific robot's joints; it only describes ‘what kind of change happened on screen,’ so the same coding scheme can be shared across human-hand video and video from very different robots. DeepMind's 2024 Genie learned 8 discrete latent actions from about 30,000 hours of platformer game video and used them to control its generated game worlds; methods such as LAPA, UniVLA, and AgiBot's GO-1 instead pretrain a VLA to predict latent actions first, then fine-tune on a small amount of real robot data to learn how to map them to real actions.

ExampleThe 8 latent actions Genie learned are numbered consistently across the games it can generate, roughly corresponding to moves like left, right, and jump; users can steer a character in a scene the model has never seen just by pressing the numbered action.

Related
Latent Action Model · Latent Action Pretraining · Action-free Video · Genie (Original) · LAPA · AgiBot GO-1
Sources
Genie: Generative Interactive Environments (arXiv 2402.15391)
Latent Action Pretraining from Videos (LAPA, arXiv 2410.11758)
AgiBot World Colosseo (GO-1, arXiv 2503.06669)

See it in the full glossary →