Embodied AI Glossary中文

Latent Action Pretraining

潜在动作预训练Advanced

Extracting unlabeled “latent actions” from video, then pretraining a robot model to predict them.

Latent action pretraining first learns “latent actions” — an encoding of what changed between two adjacent frames — from video that has no action labels, then pretrains a VLA to predict these latent actions; a small amount of real-robot data is used at the end to map the latent actions onto real robot actions. The representative work is LAPA (October 2024), from researchers at the University of Washington, KAIST, Microsoft Research, NVIDIA, and others, which learns discrete latent actions with a VQ-VAE (a vector-quantized autoencoder). Its value is that it can exploit huge amounts of internet and human-manipulation video that carries no robot action labels at all. Genie's latent action model, UniVLA, and AgiBot's GO-1 all take a similar approach.

ExampleThe LAPA paper reports positive transfer even when pretraining only on Something-Something V2 human-manipulation videos; on real-robot tasks that require language conditioning and generalization, it beats OpenVLA, which was trained with real action labels, while using roughly one-thirtieth the pretraining compute.

Related
Latent Action · Latent Action Model · LAPA · Action-free Video · Pretraining on Human Videos · Vector-Quantized Variational Autoencoder
Sources
Latent Action Pretraining from Videos (LAPA, arXiv:2410.11758)
LAPA 论文 HTML 版 (Chinese)
As of
2024-10

See it in the full glossary →