Embodied AI Glossary中文

Vector-Quantized Variational Autoencoder

向量量化变分自编码器VQ-VAEAdvanced

An autoencoder whose encoded output is quantized by a codebook into discrete IDs, commonly used to turn images, video, or actions into tokens.

VQ-VAE was proposed by DeepMind's van den Oord, Vinyals, and Kavukcuoglu in 2017 (NeurIPS 2017). It differs from a variational autoencoder (VAE, a generative model that compresses data into a continuous latent variable and reconstructs it) in two ways: the encoder's output first goes through vector quantization, replaced with the nearest codeword in a codebook, so the latent variable is discrete; and the prior isn't a fixed Gaussian but a separately trained autoregressive model (such as PixelCNN) instead. The quantization step isn't differentiable, so training copies the gradient from the decoder side straight to the encoder with a straight-through estimator, plus a commitment loss that pulls the encoder's output toward the codeword. It also eases posterior collapse, a common VAE problem. Many later image and video tokenizers, and latent action models, are built on top of it.

ExampleGoogle DeepMind's Genie uses a VQ-VAE to compress video into discrete tokens, and its latent action model uses a VQ-VAE-style objective too, learning discrete latent actions with just 8 codewords from unlabeled game video with no action labels, letting a person 'control' the generated world frame by frame.

Also called
VQ-VAE, VQ-VAE-2
Related
Variational Autoencoder · Vector Quantization · Video Tokenizer · Latent Action Model · Genie (Original) · LAPA
Sources
Neural Discrete Representation Learning (arXiv:1711.00937)
Genie: Generative Interactive Environments (arXiv:2402.15391)

See it in the full glossary →