Embodied AI Glossary中文

Consistency Policy

Consistency Policy(一致性策略)Advanced

Distills a diffusion policy into a fast visuomotor policy that generates an action in a single step.

Consistency Policy was proposed in May 2024 by Jeannette Bohg's group at Stanford (first author Aaditya Prasad) with Jimmy Wu and others at Princeton, published at RSS 2024. Diffusion Policy produces high-quality actions, but generating one requires tens to hundreds of denoising steps, too slow for a laptop GPU or onboard computer. Consistency Policy treats a trained diffusion policy as a “teacher” and distills it using the consistency trajectory model (CTM) objective: the student network is trained so that starting from any point along the denoising trajectory, it can jump straight to the same endpoint (“self-consistency”), so at inference it generates an entire action in one step. Tested across 6 simulated tasks (Robomimic, Push-T, Franka Kitchen) and 3 real-robot tasks, it runs an order of magnitude faster than comparable methods with roughly the same success rate. It's a representative approach for speeding up diffusion policies, and later work such as ConRFT uses it as the policy backbone.

ExampleIn real-robot experiments, a robot running Consistency Policy on just a laptop-class GPU completed tasks like plugging in a cord and cleaning up trash, using single-step generation in place of Diffusion Policy's many denoising steps.

Also called
Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation
Related
Diffusion Policy · Consistency Model · One-step Generation · Knowledge Distillation · Inference Latency · ConRFT
Sources
Consistency Policy (arXiv:2405.07503)
Consistency Policy 项目主页(RSS 2024) (Chinese)
As of
2024-06

See it in the full glossary →