Video Prediction Model
视频预测模型CommonA model that predicts upcoming frames from recent ones, often given the action that's about to be taken.
A video prediction model takes in the past several frames and outputs future frames; in robotics it's usually also given the action the robot is about to take, answering 'what will the camera see if I do this?' — this variant is called action-conditioned video prediction. In 2016, Finn and Levine proposed Visual Foresight: a robot pushes objects around on its own with no human labeling to collect data, trains a prediction model on that data, and then combines it with model predictive control — trying many candidate actions inside the model at each step and executing whichever one's predicted outcome comes closest to the goal — to complete pushing tasks. The difference from a video generation model is that it focuses on continuing forward from an existing view rather than generating from scratch given text. Recent work has shifted to large-scale video diffusion models instead; for example, Video Prediction Policy (VPP) uses a video model's predictive features to drive an inverse dynamics model that outputs actions.
ExampleTo push a block on a table to a target position, the robot first tries many candidate action sequences inside a video prediction model, compares which one's predicted frames come closest to the goal image, executes just the first step of that sequence, and repeats.
- Also called
- Action-conditioned Video Prediction
- Related
- Video Generation Model · World Model · Model Predictive Control · Forward Dynamics Model · Video Prediction Policy · World Action Model
- Sources
- Deep Visual Foresight for Planning Robot Motion (arXiv 1610.00696)
Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control (arXiv 1812.00568)
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations (arXiv 2412.14803)