Hierarchical Architecture
分层架构EssentialSplitting robot decision-making into a high-level planner and a low-level executor, run by two different models.
A hierarchical architecture splits “what to do” from “how to do it”: the high level is usually a large language model or vision-language model that understands the instruction and the scene and breaks a long task into steps, running at low frequency; the low level is a skill library, a VLA policy, or a motion controller that turns each step into motor commands, running at high frequency. The two levels can exchange language instructions, keypoints, or a latent vector. An early example is Google's 2022 SayCan, which has a large language model pick the next step from a robot's existing skill library; 2025's Hi Robot, from Physical Intelligence with Stanford and Berkeley, has a vision-language model output short language instructions for π0 to carry out. It is often contrasted with end-to-end approaches, where one model goes directly from input to action, but the two are not mutually exclusive: China's “brain-cerebellum” framing and dual-system architectures are both hierarchical in spirit, and a dual-system model such as Helix, despite having two layers, is trained end-to-end jointly.
ExampleFigure's Helix splits into two layers: a 7-billion-parameter vision-language model understands the scene and the instruction at 7 to 9Hz, and an 80-million-parameter low-level policy turns what it outputs into continuous actions at 200Hz.
- Also called
- Hierarchical Policy, Hierarchical Scheme, High-Level Planning + Low-Level Execution, Hierarchical VLA
- Related
- Dual-System Architecture (System 1 / System 2) · Brain–Cerebellum Architecture · End-to-End · SayCan · Hi Robot · Long-horizon Task
- Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances (SayCan, arXiv:2204.01691)
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models (arXiv:2502.19417)
Helix: A Vision-Language-Action Model for Generalist Humanoid Control (Figure) - As of
- 2025-02