Embodied AI Glossary中文

HAMSTER

HAMSTER(分层动作模型)Advanced

A 2025 hierarchical VLA from NVIDIA and others where a high level sketches a 2D path and a low-level 3D policy follows it.

HAMSTER was proposed in February 2025 by researchers at NVIDIA, the University of Washington, and USC, published at ICLR 2025. Ordinary VLA models fine-tune a vision-language model (VLM) directly to output actions, which can only be trained on expensive real-robot data. HAMSTER splits the system into two layers: the high level is a fine-tuned VLM (built on VILA) that looks at an RGB image and a task description and sketches a rough 2D path showing roughly how the end effector should move; the low level is a control policy that handles 3D input (the paper tries RVT-2 and 3D Diffuser Actor), using that path as guidance for precise manipulation. Because the high level only has to output a 2D path, it can be trained on cheap “out-of-domain” data — action-label-free video, hand-drawn sketches, simulation data. The high level doesn't handle fine motor control and the low level doesn't handle task reasoning; each does what it's good at.

ExampleTested on a real robot across 7 generalization axes (such as new objects, new backgrounds, and new instruction semantics), HAMSTER beats OpenVLA by about 20 percentage points in average success rate, a roughly 50% relative improvement.

Also called
Hierarchical Action Models for Open-World Robot Manipulation
Related
Hierarchical Architecture · Intermediate Representation · Vision-Language-Action Model · RVT-2 · 3D Diffuser Actor · OpenVLA
Sources
HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation (arXiv 2502.05485)
HAMSTER 项目页 (Chinese)
As of
2025-05

See it in the full glossary →