Embodied AI Glossary中文

Behavior Transformer

BeT / VQ-BeTBeTAdvanced

A Transformer-based imitation-learning method that learns several different valid behaviors at once from multimodal demonstration data.

BeT (Behavior Transformer) is a behavior-cloning method proposed in 2022 by Lerrel Pinto's group at NYU. Human demonstrations often show “action multimodality”: several equally valid ways to act in the same situation, and direct regression averages them into one wrong action. BeT first clusters continuous actions into a number of bins with k-means, has a Transformer predict which bin to pick, and then predicts a continuous offset to correct it into a precise action — preserving multiple behavior modes this way. VQ-BeT, from 2024 (NYU and Seoul National University, ICML 2024), replaces k-means with residual vector quantization, which suits high-dimensional actions and long action sequences better, running inference at roughly 5x the speed of a diffusion policy. BeT and diffusion policies are two representative lines of work from the same period addressing the same multimodality problem.

ExampleIn the Franka Kitchen simulated kitchen, demonstrators complete a set of subtasks in different orders each time, and BeT is used to test whether a model can learn and reproduce these different behavior patterns.

Also called
BeT, VQ-BeT, Vector-Quantized Behavior Transformer
Related
Action Multimodality · Behavior Cloning · Vector Quantization · Diffusion Policy · Action Tokenizer · Franka Kitchen
Sources
Behavior Transformers: Cloning k modes with one stone (arXiv 2206.11251)
Behavior Generation with Latent Actions (VQ-BeT, arXiv 2403.03181)
VQ-BeT 项目主页 (Chinese)
As of
2024-07

See it in the full glossary →