Visual Foresight
视觉预见 / 基于学习模型的规划AdvancedLearning to predict what the camera view will look like after an action, then picking whichever imagined action reaches the goal best.
Visual foresight was proposed by Chelsea Finn and Sergey Levine at UC Berkeley (as ‘deep visual foresight’) at ICRA 2017, and assembled into a complete framework by Frederik Ebert, Finn, and colleagues in 2018. It works in two steps. First, an action-conditioned video prediction model is trained on unlabeled data collected by a robot autonomously pushing objects around: given the current image and a sequence of candidate actions, it predicts the next several frames. Second, the system performs visual model predictive control (MPC): it samples a large number of candidate action sequences — often refined iteratively with the cross-entropy method — lets the model ‘imagine’ the outcome of each, scores them by how close they get to the goal, executes only the first step of the best sequence, and then replans. The goal can be ‘move this pixel here,’ a target image, or a classifier. The approach needs no reward function or human labels and works on objects it has never seen. It is the direct predecessor of today's ‘world model plus planning’ approaches — DINO-WM and V-JEPA 2-AC, for example, move the prediction from pixels into a pretrained feature space but still use MPC to choose actions.
ExampleYou point at a pixel on some object in the robot's camera view and specify where it should end up. The robot imagines trying a large number of action sequences, picks whichever one's predicted outcome pushes that pixel closest to the target, executes just the first step, and then replans.
- Also called
- Visual MPC, Planning with Learned World Models, Visual Model Predictive Control
- Related
- World Model · Video Prediction Model · Model Predictive Control · Cross-Entropy Method · DINO-WM · V-JEPA 2
- Sources
- Finn, Levine: Deep Visual Foresight for Planning Robot Motion (arXiv:1610.00696, ICRA 2017)
Ebert et al.: Visual Foresight: Model-Based Deep RL for Vision-Based Robotic Control (arXiv:1812.00568)
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning (arXiv:2411.04983) - As of
- 2025-06