Embodied AI Glossary中文

Spatial Generalization

位置泛化(空间泛化)Common

A policy's ability to still succeed when an object is placed at a position or orientation absent from training.

Spatial (or position) generalization means the task and the object stay the same, but the object's, or the robot's own, initial position or orientation shifts to somewhere the training data never covered, and the question is whether the policy can still succeed. OpenVLA's evaluation calls this “motion generalization,” defining it as unseen object positions and orientations. It is a common weak point for visuomotor policies: a policy learned through imitation learning is often only reliable near the region the demonstrations covered, and success rates drop noticeably once an object is moved further away — which is why data collection involves repeatedly placing objects at different spots on the table. DemoGen addresses this by synthesizing large numbers of demonstrations at different positions from a single real human demonstration; LIBERO-Plus found that success rates for some VLA models fall from 95% to under 30% with only a slight perturbation to the robot's initial state or camera viewpoint.

ExampleDuring training, the block is only ever placed on the left half of the table; at test time it is placed on the right half or rotated 90 degrees, to see whether the arm can still pick it up the same way.

Also called
Position Generalization, Motion Generalization
Related
Generalization · Object Generalization · Scene Generalization · Out-of-Distribution · DemoGen · LIBERO-Plus
Sources
OpenVLA: An Open-Source Vision-Language-Action Model
DemoGen: Synthetic Demonstration Generation for Data-Efficient Visuomotor Policy Learning
LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
As of
2025-10

See it in the full glossary →