Action Chunking with Transformers
ACTEssentialA 2023 Stanford imitation-learning policy that predicts a whole chunk of actions at once, letting low-cost bimanual robots do fine manipulation.
ACT is an imitation-learning algorithm introduced in April 2023 by Tony Zhao, Chelsea Finn, and colleagues at Stanford, with collaborators from UC Berkeley and Meta, published alongside the low-cost bimanual platform ALOHA (roughly $20,000 in hardware) at RSS 2023. Behavior cloning — learning actions by copying human demonstrations — tends to accumulate small errors at every step, which is especially damaging for fine-grained tasks. ACT has a Transformer output an entire chunk of future actions at once, so fewer decisions are made and less error accumulates. During training it wraps a conditional variational autoencoder (CVAE, a generative model that can represent multiple valid ways of doing the same thing) around the policy to absorb the randomness in human demonstrations; at execution time, temporal ensembling averages overlapping action chunks with weights to keep motion smooth. Because it's structurally simple and needs little data, ACT went on to become a standard baseline for projects like Mobile ALOHA and LeRobot.
ExampleOn ALOHA, ACT learned six fine-grained tasks — such as opening a translucent condiment cup and slotting a battery into place — from only about 10 minutes of teleoperated demonstrations per task, reaching 80–90% success rates.
- Also called
- ACT, Action Chunking Transformer
- Related
- Action Chunking · Temporal Ensembling · Conditional Variational Autoencoder · ALOHA · Mobile ALOHA · Behavior Cloning
- Sources
- Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (arXiv 2304.13705)
ALOHA / ACT 项目主页 (Chinese) - As of
- 2023-04