Embodied AI Glossary中文

Visual Generalization

视觉泛化Advanced

Still completing the task when the scene looks different: new background, lighting, colors, distractors, or camera angle.

Visual generalization is a subcategory of generalization: whether a robot policy still completes the same task under visual conditions it never saw in training, such as a new background or tablecloth, different lighting, an object with a new color or texture, extra distractor objects in the frame, or a moved camera. The task and the actions needed haven't changed, only what the scene looks like, so it is usually evaluated separately from semantic generalization (new objects, new instructions) and position generalization; OpenVLA's real-robot evaluation lists it as its own category. Imitation-learning policies that learn actions directly from pixels tend to also memorize irrelevant details like background and lighting. In 2023, Xie, Finn, and colleagues isolated these factors one by one and found that new backgrounds are the easiest to adapt to and new camera positions the hardest. Common fixes include domain randomization, data augmentation, pretrained vision encoders, and collecting data across more varied scenes.

ExampleOpenVLA's real-robot evaluation of “put the eggplant in the pot” used a pot made of papier-mâché, visually different from the pots in the BridgeData V2 training data, to test whether the policy could still recognize the pot and complete the task.

Also called
Visual Robustness, Appearance Generalization
Related
Generalization · Semantic Generalization · Spatial Generalization · Distractor Objects · Domain Randomization · Data Augmentation
Sources
Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation (Xie et al., 2023)
OpenVLA: An Open-Source Vision-Language-Action Model
What Can RL Bring to VLA Generalization? An Empirical Study
As of
2025-05

See it in the full glossary →