V-JEPA 2
CommonMeta's self-supervised video model that predicts video in feature space, usable as a world model for planning robot actions.
V-JEPA 2 is an open-source video model, about 1.2 billion parameters, released by Meta FAIR (Yann LeCun's team) in June 2025. It follows the JEPA (Joint Embedding Predictive Architecture) approach: instead of generating pixels, it masks part of a video and has the model predict the masked content in an abstract feature space, learning motion and physical regularities this way. The first stage does self-supervised pretraining on more than 1 million hours of web video and 1 million images; the second stage freezes the encoder and trains an action-conditioned predictor, called V-JEPA 2-AC, on fewer than 62 hours of unlabeled robot video from the DROID dataset. In use, given a goal image, the model searches in feature space for the action whose predicted outcome comes closest to the goal (model-predictive control). On Franka arms in two different labs, it grasped and placed unseen objects zero-shot with 65–80% success. In 2026 Meta also released V-JEPA 2.1, with improved dense features.
ExampleIn a new lab where no data was ever collected, given a photo showing an object already placed at its target position, V-JEPA 2-AC imagines the outcome of several candidate actions in feature space at each step, executes whichever one lands closest to the goal image, and moves the object into place step by step.
- Also called
- V-JEPA 2-AC, V-JEPA, V-JEPA 2.1
- Related
- Joint-Embedding Predictive Architecture · World Model · Latent World Model · Self-Supervised Learning · Model Predictive Control · DROID (Distributed Robot Interaction Dataset)
- Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv 2506.09985)
Introducing V-JEPA 2 (Meta AI blog)
V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning (arXiv 2603.14482) - As of
- 2026-06