Embodied AI Glossary中文

Representation Learning

表征学习Common

Letting a model automatically learn useful feature vectors from raw data, instead of hand-designing the features.

Representation learning studies how to turn raw data — images, text, sensor readings — into a set of numbers (a representation, also called a feature or embedding vector) that's easier for downstream tasks to use. Bengio and colleagues' 2013 survey treats this as deep learning's central problem: a good representation should separate out the different underlying factors of variation behind the data. There are many ways to learn one: supervised classification, self-supervised learning (contrastive learning, masked autoencoders), and image-text alignment (as in CLIP). In embodied AI, because robot data is scarce, it's common to first pretrain a visual encoder on large-scale images or human video, then freeze or fine-tune it while learning a policy, so only a small number of demonstrations are needed — R3M, VC-1, and DINOv2 all follow this path. A world model's latent space and its latent actions are also products of representation learning.

ExampleR3M pretrains a visual representation on Ego4D first-person human video using time-contrastive learning and video-language alignment; once frozen and attached to a policy for a Franka arm, it learns manipulation tasks in a real, cluttered apartment from just 20 demonstrations.

Also called
Feature Learning
Related
Self-Supervised Learning · Contrastive Learning · Pre-trained Visual Representation · Embedding · Latent Space · R3M
Sources
Representation Learning: A Review and New Perspectives (Bengio et al.)
R3M: A Universal Visual Representation for Robot Manipulation

See it in the full glossary →