Embodied AI Glossary中文

Autoregressive Decoding

自回归解码AREssential

Generating output one piece at a time, with each step conditioned on everything generated so far.

Autoregressive decoding is how large language models such as GPT generate text: predict only the next token, append it to the input, predict the next one again, and repeat until done. Applying this to robots first requires discretizing continuous actions into tokens. Google DeepMind's 2023 RT-2 wrote actions as a string of number tokens and had the vision-language model output them the same way it outputs text; OpenVLA splits each action dimension into 256 bins, reusing the 256 least-used tokens in Llama's vocabulary. The upside is directly reusing a language model's architecture and training recipe; the downside is that generation has to happen one token at a time, which is slow — OpenVLA runs at only about 6Hz on an RTX 4090. This led to compressive action tokenizers such as FAST, parallel decoding, and alternatives that generate actions with diffusion or flow matching instead.

ExampleOne of RT-2's action outputs is a string of tokens like “1 128 91 241 5 101 127 217,” corresponding in order to whether the episode should end, end-effector translation and rotation, and gripper open/close, which then get converted back into continuous values sent to the robot.

Also called
AR, Autoregressive Model, Token-by-token Generation
Related
Next-Token Prediction · Action Tokenizer · Action Binning · Parallel Decoding · RT-2 · OpenVLA
Sources
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control(项目页) (Chinese)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)

See it in the full glossary →