Joint-Embedding Predictive Architecture
联合嵌入预测架构JEPACommonA self-supervised architecture that predicts masked or future content in an abstract representation space instead of reconstructing raw pixels.
JEPA is an architecture Yann LeCun proposed in his 2022 position paper ‘A Path Towards Autonomous Machine Intelligence’; Meta later built an image version, I-JEPA (2023), and a video version, V-JEPA. It encodes part of the input (the context), then has a predictor guess the representation of another part — a masked region, or a future frame — with the loss computed on the representation rather than on pixels. That means the model doesn't need to reconstruct hard-to-predict details like individual leaf texture, and can instead focus on semantic information such as objects and motion, though extra care is needed to keep the representation from collapsing. V-JEPA 2, released in June 2025, was pretrained on more than 1 million hours of video and then fine-tuned on fewer than 62 hours of robot video to produce the action-conditioned V-JEPA 2-AC, which can plan pick-and-place actions zero-shot on a Franka arm given a goal image. LeCun left Meta in November 2025 and founded AMI Labs to continue working on world models.
ExampleFor pick-and-place, V-JEPA 2-AC is given a goal image showing the cup already placed on the plate; the model plays out several candidate action sequences in representation space and executes whichever one lands closest to the goal image's predicted representation.
- Also called
- JEPA, I-JEPA, V-JEPA
- Related
- V-JEPA 2 · World Model · Self-Supervised Learning · Latent World Model · Masked Autoencoder · Representation Collapse
- Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv 2506.09985)
Meta AI: I-JEPA, the first AI model based on Yann LeCun's vision
Wikipedia: Yann LeCun - As of
- 2025-11