Embodied AI Glossary中文

MVP

MVP(掩码视觉预训练)Advanced

Uses a masked autoencoder to pretrain a vision encoder on huge amounts of natural images, then freezes it for robot use.

MVP (Masked Visual Pre-training) refers to two papers from Jitendra Malik and Trevor Darrell's groups at UC Berkeley: Masked Visual Pre-training for Motor Control (Tete Xiao and colleagues), from March 2022, validated in simulation, and Real-World Robot Learning with Masked Visual Pre-training (Ilija Radosavovic and colleagues, CoRL 2022), from October 2022, extending it to the real robot. The method first self-supervised-pretrains a ViT vision encoder on web images and first-person video using a masked autoencoder (MAE, which hides most small patches of an image and has the network fill them back in), then freezes the encoder and trains only a small control module on top of it. Results show this representation beats CLIP, ImageNet-supervised pretraining, and training from scratch; a 307-million-parameter ViT trained on 4.5 million images keeps improving further still. Together with R3M and VC-1, it helped drive forward the direction of pretrained visual representations.

ExampleFreeze an MVP-pretrained ViT to serve as the robot's eyes, and training only a small control head on top of it with a handful of demonstrations is enough to learn grasping on a new task.

Also called
Masked Visual Pre-training, Masked Visual Pre-training for Motor Control, Real-World Robot Learning with Masked Visual Pre-training
Related
Pre-trained Visual Representation · Masked Autoencoder · Self-Supervised Learning · R3M · VC-1 · Vision Transformer
Sources
arXiv 2203.06173: Masked Visual Pre-training for Motor Control
arXiv 2210.03109: Real-World Robot Learning with Masked Visual Pre-training
As of
2022-10

See it in the full glossary →