Embodied AI Glossary中文

Humanoid Locomotion as Next Token Prediction

人形行走即下一个 token 预测Advanced

Treats a humanoid robot's walking control as a language-model-style “predict the next token” problem, learned with a causal Transformer.

This was released in February 2024 by Ilija Radosavovic, Koushil Sreenath, Jitendra Malik, and colleagues at UC Berkeley. It frames real humanoid robot control as a problem similar to language modeling: a causal Transformer autoregressively predicts a sensorimotor trajectory made of observations and actions, where each token predicts the next token of the same modality. This lets data lacking action labels — such as action trajectories extracted from human video — also participate in training. The data comes from simulated trajectories generated by existing neural-network policies and model-based controllers, human motion-capture data, and YouTube videos of humans. The model was deployed on Agility Robotics' full-size humanoid Digit, walking zero-shot on the streets of San Francisco; using just 27 hours of walking data is also enough to transfer to the real robot, and it generalizes to a backward-walking command that was never in the training data.

ExampleThere is no backward-walking command in the training data, yet once deployed, the model can still make Digit walk backward on command.

Related
Next-Token Prediction · Autoregressive Decoding · Action-free Video · Bipedal Locomotion · Agility Robotics Digit · Real-World Humanoid Locomotion with Reinforcement Learning
Sources
Humanoid Locomotion as Next Token Prediction (arXiv 2402.19469)
Project page
As of
2024-02

See it in the full glossary →