Embodied AI Glossary中文

Hi Robot

Advanced

A 2025 Physical Intelligence hierarchical system where a high-level VLM breaks down complex instructions and low-level π0 executes them.

Hi Robot was released in February 2025 by Physical Intelligence together with researchers at Stanford and UC Berkeley, published at ICML 2025. A single VLA is good at executing simple instructions like “pick up the cup,” but struggles with complex requests that carry conditions, and with a user interrupting mid-execution to correct it. Hi Robot splits the system into two layers: the high level is a vision-language model built on PaliGemma-3B that looks at the current view and what the user says, reasons out the simple instruction to execute next, and can also talk back to the user; the low level is π0, which turns that simple instruction into continuous actions. To train the high level, the team cut teleoperation demonstrations into short skill segments, then had a large VLM work backward to infer “what the user might have said at that moment, and how the robot should respond,” synthesizing conversational training data and skipping manual annotation. The system was tested on a single-arm UR5e, a bimanual ARX, and a mobile bimanual ARX.

ExampleWhen the user says “make me a vegetarian sandwich,” the high level breaks it down into step-by-step sub-instructions for getting each ingredient, skipping any meat; while clearing a table, if the user says “that's not trash,” the robot stops and adjusts what it's doing.

Also called
Hierarchical Interactive Robot, Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models
Related
Hierarchical Architecture · Dual-System Architecture (System 1 / System 2) · π0 · Instruction Following · Language Corrections · Physical Intelligence
Sources
Hi Robot: Open-Ended Instruction Following with Hierarchical Vision-Language-Action Models (arXiv 2502.19417)
Hi Robot 论文 HTML 全文 (Chinese)
As of
2025-07

See it in the full glossary →