Object Generalization
物体泛化CommonA policy's ability to still complete the same task when the object is swapped for one it never saw in training.
Object generalization is the most common type of generalization tested for robot policies: whether a task can still be completed once the object is swapped for one absent from the training data — a new category, a new size or shape, a new color or material. OpenVLA's evaluation breaks this down further: changes in color and appearance count as visual generalization, changes in size and shape count as physical generalization, and entirely unseen target objects count as semantic generalization. It matters because the variety of objects in real homes and factories is nearly endless — collecting data for every single one is impossible. A 2024 study on scaling imitation-learning data found that a policy's generalization to new objects grows as a power law with the number of distinct object categories in training, and that object diversity matters more than simply piling up more demonstrations of the same objects.
ExampleDemonstrations of 'put the cup on the plate' are collected using only 3 kinds of cups; at test time the policy is tried with an unseen mug, a paper cup, and a glass to see how much the success rate drops.
- Related
- Generalization · Scene Generalization · Spatial Generalization · Semantic Generalization · Visual Generalization · Zero-shot
- Sources
- OpenVLA: An Open-Source Vision-Language-Action Model
Data Scaling Laws in Imitation Learning for Robotic Manipulation