Embodied AI Glossary中文

Denoising Steps

去噪步数NFECommon

How many times a diffusion or flow-matching model calls its network to generate one result; this sets inference speed.

Diffusion and flow-matching models don't generate a sample in one shot; they start from noise and call the network repeatedly, each call correcting the result a little. The number of network calls is the number of denoising steps, often written precisely as NFE, the number of function evaluations. More steps usually give a more refined result, but time cost grows roughly proportionally, which matters for robots since a policy has to produce actions in real time inside a closed control loop. Values vary widely by method: the original DDPM used 1,000 steps; Diffusion Policy trains with 100 steps but drops to 10 at inference using DDIM; π0 integrates flow matching over 10 steps; GR00T N1 uses just 4. Consistency models and mean-flow methods push toward generating in 1 or 2 steps. Note that the number of diffusion steps used in training and the number of sampling steps used at inference don't have to match — inference steps are generally adjustable at deployment time.

ExampleWhen π0 predicts a 50-step-long action chunk, it integrates 10 steps, each of size δ=0.1, starting from Gaussian noise, meaning it calls the action expert 10 times to produce one action segment.

Also called
Number of Function Evaluations, NFE, Sampling Steps, Inference Steps
Related
Denoising Diffusion Probabilistic Model · Denoising Diffusion Implicit Model · Flow Matching · Consistency Model · Inference Latency · One-step Generation
Sources
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)

See it in the full glossary →