Embodied AI Glossary中文

Classifier-Free Guidance

无分类器引导CFGAdvanced

Computing both a conditional and an unconditional prediction and extrapolating between them so generation follows the condition more closely.

Classifier-free guidance was proposed by Jonathan Ho and Tim Salimans (a 2021 NeurIPS workshop paper, with a full arXiv version in 2022), for use with diffusion models, flow matching, and other generative models that denoise step by step. During training, the condition (such as a text prompt) is randomly dropped for some examples, so the same network learns to make both conditional and unconditional predictions; during generation, both are computed at every step, and the final direction is 'unconditional result + w × (conditional result − unconditional result),' where w is called the guidance scale. A larger w follows the condition more closely but reduces diversity, and too large a value introduces artifacts. It replaces the earlier 'classifier guidance,' which needed training a separate classifier, and has become a standard setting in text-to-image and text-to-video models, at the cost of one extra forward pass per step. In embodied AI, diffusion-based video world models and action generation models can also use it to control how closely they follow a language instruction.

ExampleGenerating an image from text with Stable Diffusion v1.5 in Diffusers, turning guidance_scale up from 2.5 to 10.5 makes the image follow the prompt more and more closely, though artifacts start to appear once it's too high.

Also called
CFG, Guidance Scale
Related
Diffusion Model · Flow Matching · Denoising Steps · Text-to-Video / Image-to-Video · Generative Model · Video Generation Model
Sources
Classifier-Free Diffusion Guidance (arXiv 2207.12598)
Hugging Face Diffusers: Text-to-image(guidance_scale 说明) (Chinese)

See it in the full glossary →