Embodied AI Glossary中文

Generative Data Augmentation

生成式数据增强Common

Using image or video generation models to rewrite existing robot data, producing new-scene training examples from old demonstrations.

Generative data augmentation means using generative models — text-to-image or video generation — to swap in new objects, backgrounds, distractors, or lighting on top of existing robot demonstrations, producing new training examples while reusing the original action labels. Notable examples include Google's 2023 ROSIE, which uses a text-guided diffusion model to inpaint new objects and backgrounds into images, and the University of Washington and collaborators' GenAug. It targets the problem that real-robot data is expensive and each demonstration only covers one tabletop: after augmentation, the same motion can be paired with many different appearances, improving visual generalization and robustness to distractors. Its limitation is that it mainly changes appearance, not the physical process or the action itself. Video world models such as Cosmos Transfer are also commonly used for this kind of augmentation.

ExampleROSIE uses a diffusion model to inpaint new objects and backgrounds onto Google's existing robot data; a policy trained on these images can handle new objects and distractors that never appeared in the original data.

Also called
Semantic Data Augmentation
Related
Data Augmentation · Synthetic Data · Diffusion Model · NVIDIA Cosmos Transfer · Visual Generalization · Cross-Painting
Sources
ROSIE: Scaling Robot Learning with Semantically Imagined Experience (arXiv 2302.11550)
GenAug: Retargeting behaviors to unseen situations via Generative Augmentation (arXiv 2302.06671)

See it in the full glossary →