BFM-Zero
AdvancedA humanoid whole-body control foundation model trained with no task reward at all, switching tasks through a simple “prompt.”
BFM-Zero is a humanoid behavioral foundation model released in November 2025 by Guanya Shi's group at Carnegie Mellon University together with Meta and other institutions. It's trained with unsupervised reinforcement learning: no specific task reward is given during training; instead, a forward-backward representation (a method that maps states and tasks into the same latent space) learns a shared latent space, into which reference motions, target poses, and reward functions can all be encoded as vectors. At deployment, giving the corresponding vector — the “prompt” — lets the same policy perform motion tracking, reaching a target pose, or optimizing a given reward, all zero-shot, and it also supports few-shot adaptation. The team demonstrated push recovery and getting back up after a fall on a real Unitree G1, calling it the first behavioral foundation model that can switch tasks on a real humanoid through prompts.
ExampleThe same BFM-Zero policy deployed on a Unitree G1 does motion tracking when given a reference motion, poses itself to match a given target pose, or acts to optimize a given reward function — all without retraining in between.
- Also called
- BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised RL
- Related
- Behavior Foundation Model · Unsupervised Skill Discovery · Motion Tracking · Unitree G1 · Meta Motivo · Learning-Based Whole-Body Control
- Sources
- BFM-Zero (arXiv 2511.04131)
BFM-Zero 项目主页 (Chinese) - As of
- 2025-11