Embodied AI Glossary中文

Generative Model

生成模型Common

A model that learns the probability distribution behind data and can sample new data from it.

A generative model learns the distribution of the data itself, p(x), or a conditional distribution p(x|c), so it can produce new samples. This is in contrast to a discriminative model, which only learns p(y|x) and is used for classification or scoring. Common generative models include autoregressive models (which predict the next token one at a time, like GPT), variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion models, and flow matching. In embodied AI, generative models show up in three main places. First, for generating actions: demonstration data for the same scene often contains several equally valid ways to solve it (action multimodality), and methods like Diffusion Policy and π0 use diffusion or flow matching to model the whole action distribution instead of regressing to a single average. Second, for generating future frames, which is what world models and video-prediction models do. Third, for generating training data itself, such as using a video generation model to synthesize robot demonstrations.

ExampleSuppose half the demonstrations go left around an obstacle and half go right: a model trained with direct regression outputs the average of the two and drives straight into it, while a generative model like Diffusion Policy samples either ‘go left’ or ‘go right.’

Related
Diffusion Model · Flow Matching · Variational Autoencoder · Generative Adversarial Network · Action Multimodality · Diffusion Policy
Sources
Google Machine Learning: Background: What is a Generative Model?
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)

See it in the full glossary →