Embodied AI Glossary中文

VC-1

Advanced

Meta's visual encoder for embodied tasks, pretrained with masked autoencoding on more than 4,000 hours of first-person video.

VC-1 is research released by Meta AI (FAIR) in March 2023, and was, at the time, the largest systematic evaluation of pretrained visual representations — off-the-shelf visual encoders meant to serve as a robot's 'eyes.' The authors first built CortexBench, containing 17 tasks spanning locomotion, navigation, dexterous manipulation, and mobile manipulation; they then combined more than 4,000 hours of first-person video from 7 sources with ImageNet and used a masked autoencoder (MAE, which masks out image patches and reconstructs them) to train ViTs of various sizes, the largest being ViT-L, which is VC-1. The conclusion: no single visual representation is best on every task, and scaling up data size and diversity only helps on average; once adapted to a specific task, VC-1 matches or beats the best previously known results across the board. Both the model and code are open-sourced.

ExampleWhen building a robot-arm imitation-learning project, VC-1 can be used directly as a frozen image encoder, turning camera images into feature vectors on top of which a small policy network is trained.

Also called
Visual Cortex 1, Artificial Visual Cortex
Related
Pre-trained Visual Representation · Masked Autoencoder · R3M · VIP · Ego4D · Vision Transformer
Sources
Where are we in the search for an Artificial Visual Cortex (arXiv 2303.18240)
VC-1 project page
As of
2023-03

See it in the full glossary →