VIP
VIP(价值隐式预训练)AdvancedA self-supervised pretraining method on human videos that produces both a visual representation and a dense reward at once.
VIP was released by Jason Ma, Amy Zhang, and colleagues at Meta AI (FAIR) and the University of Pennsylvania in September 2022, published at ICLR 2023 (Spotlight). Robot reinforcement learning is often stuck on two things at once: no good visual features, and no easy-to-write reward function. VIP frames 'learning a representation from human video' as an offline goal-conditioned reinforcement learning problem, and derives a value-function objective that needs no action labels — essentially a form of implicit time-contrastive learning, where frames closer in time to completing the task sit closer to the goal in feature space. After pretraining on Ego4D first-person videos, it is used frozen: the reward is simply the change in feature-space distance between the current frame and a goal image, which supplies a dense reward for a wide range of simulated and real-robot tasks; on real robots, as few as about 20 trajectories are enough for few-shot offline reinforcement learning.
ExampleGiven a goal photo showing 'the drawer already closed,' VIP encodes every camera frame and the goal image into features; as the distance shrinks, the robot gets positive reward, learning to push the drawer shut without anyone writing a reward function by hand.
- Also called
- Value-Implicit Pre-Training, VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
- Related
- R3M · VC-1 · Pre-trained Visual Representation · Time-Contrastive Networks · Goal-Conditioned Reinforcement Learning · Ego4D
- Sources
- VIP (arXiv 2210.00030)
VIP project page - As of
- 2023-03