Embodied AI Glossary中文

LeVERB

Advanced

A dual-system framework that links a vision-language model to a humanoid whole-body controller through a “latent verb.”

LeVERB was released in June 2025, led by UC Berkeley, with collaborators including Xue Bin Peng, Trevor Darrell, and Koushil Sreenath. Existing VLA models mostly assume the low-level controller only accepts manually defined commands like end-effector pose or base velocity, limiting them to quasi-static tasks. LeVERB splits into two layers: the high-level LeVERB-VL (System 2, 10Hz) encodes first- and third-person images and the instruction with SigLIP, learning a “latent verb” space through a conditional variational autoencoder; the low-level LeVERB-A (System 1, 50Hz) is a whole-body controller, first trained with PPO as a teacher that tracks reference motions, then distilled with DAgger into a student conditioned on the latent verb. The authors used IsaacSim to render motion-capture playback and built LeVERB-Bench, covering more than 150 tasks. Deployed zero-shot on a Unitree G1, it reaches 80% success on simple visual navigation and 58.5% overall, 7.8 times that of a naive hierarchical baseline.

ExampleTell the robot “walk over to the red chair and sit down,” and the high-level model looks at the camera feed and outputs a latent verb, which the low-level controller uses to make the G1 walk over, turn, and sit.

Also called
Latent Vision-Language-Encoded Robot Behavior, LeVERB-Bench
Related
Learning-Based Whole-Body Control · Dual-System Architecture (System 1 / System 2) · Latent Action · Conditional Variational Autoencoder · Unitree G1 · DAgger
Sources
LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction (arXiv 2506.13751)
As of
2025-09

See it in the full glossary →