Embodied AI Glossary中文

Variational Autoencoder

变分自编码器VAECommon

A model that compresses data into a distribution over latent variables, and can sample from it to reconstruct or generate data.

The variational autoencoder was proposed by Kingma and Welling in 2013 and is a type of generative model. Its encoder compresses the input (such as an image) into a low-dimensional latent variable, but outputs a distribution (a mean and variance) rather than a single fixed point; its decoder samples from that distribution and reconstructs data from it. Training optimizes two things at once: reconstructions should look right, and the latent distribution should stay close to a standard normal distribution, which keeps the latent space continuous and smooth, so decoding a randomly sampled point still gives a plausible result. Thanks to the reparameterization trick, a model with this kind of random sampling can still be trained with gradient descent. Its most common use today is as a compressor in front of a diffusion model: Stable Diffusion and Wan both use a VAE to compress pixels into latent space before denoising; in robotics, ACT uses its conditional variant, the CVAE, to capture the variety in demonstration actions.

ExampleStable Diffusion's VAE compresses a 512×512 color image into a 64×64×4 latent variable; the diffusion model does all its denoising in that much smaller space, and only the decoder turns the result back into an image at the end.

Also called
VAE
Related
Autoencoder · Latent Space · Conditional Variational Autoencoder · Vector-Quantized Variational Autoencoder · Latent Diffusion Model · Video Tokenizer
Sources
Auto-Encoding Variational Bayes (arXiv 1312.6114)
High-Resolution Image Synthesis with Latent Diffusion Models (arXiv 2112.10752)

See it in the full glossary →