Evidence Lower Bound
证据下界ELBOAdvancedA tractable lower bound on the log-likelihood of data; VAEs and diffusion models both train by maximizing it.
A generative model wants to maximize the data's log-likelihood, log p(x) (also called the “evidence”), but this is usually intractable when the model has latent variables. The ELBO introduces an approximate posterior q(z|x), giving log p(x) ≥ E_q[log p(x|z)] − KL(q(z|x) ‖ p(z)): the first term measures reconstruction quality, and the second pulls the encoding distribution toward the prior, with the gap between the two sides equal to the KL divergence between q and the true posterior. Kingma and Welling's 2013 VAE paper used the reparameterization trick to let the ELBO be optimized directly with gradient descent; Ho and colleagues' 2020 DDPM also starts from a variational lower bound, simplifying and reweighting it into the familiar noise-prediction loss. In robotics, ACT is trained as a conditional VAE.
ExampleACT's training loss is the action-chunk reconstruction error plus β times the KL divergence between the encoding distribution and a standard normal distribution — exactly a β-weighted ELBO.
- Also called
- ELBO, Variational Lower Bound, VLB
- Related
- Variational Autoencoder · Conditional Variational Autoencoder · Kullback-Leibler Divergence · Denoising Diffusion Probabilistic Model · Maximum Likelihood Estimation · Action Chunking with Transformers
- Sources
- Auto-Encoding Variational Bayes (arXiv:1312.6114)
Denoising Diffusion Probabilistic Models (arXiv:2006.11239)
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv:2304.13705)