Embodied AI Glossary中文

Self-Supervised Learning

自监督学习SSLCommon

Training a model on “questions and answers” constructed from the data itself, with no human labeling needed.

Self-supervised learning is training where the supervisory signal comes from the data itself rather than human labels: part of the input is hidden, and the model is trained to predict it from the rest. Yann LeCun and colleagues called it the “dark matter” of intelligence and championed it heavily. Common forms include masking out words in text for the model to fill in (BERT; GPT's next-token prediction is also often grouped under this umbrella), masking out 75% of an image's patches and reconstructing them (MAE, masked autoencoders), and pulling representations of different views of the same content closer together (contrastive learning). It solves the problem that human labeling is expensive and caps how much data can be used, and it underlies the pretraining of today's large models. In embodied AI, since robot data with action labels is scarce, researchers often first self-supervise a visual representation or a world model on huge amounts of human video, then fine-tune it with a small amount of robot data.

ExampleMeta's V-JEPA 2 first self-supervises on over 1 million hours of internet video, then trains an action-conditioned world model, V-JEPA 2-AC, on fewer than 62 hours of unlabeled DROID robot video, letting a Franka arm in labs it never trained on do pick-and-place by planning against an image goal, with no task-specific training.

Also called
SSL, Self-Supervised Pretraining
Related
Contrastive Learning · Masked Autoencoder · Pre-training · Representation Learning · Next-Token Prediction · V-JEPA 2
Sources
Self-supervised learning: The dark matter of intelligence (Meta AI, 2021)
Masked Autoencoders Are Scalable Vision Learners (arXiv 2111.06377)
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning (arXiv 2506.09985)

See it in the full glossary →