Embodied AI Glossary中文

Representation Alignment

表征对齐REPAAdvanced

Training a model so its intermediate features match those of an existing pretrained encoder.

REPA was proposed by Sihyun Yu, Saining Xie, and colleagues in October 2024 (an ICLR 2025 Oral), targeting the slow training of diffusion Transformers such as DiT and SiT. It adds an auxiliary loss: the network's intermediate hidden state while processing a noisy image, after passing through a small projection layer, is aligned with that same clean image's features from a pretrained visual encoder such as DINOv2. The authors argue that one bottleneck in diffusion-model training is having to learn good visual representations from scratch, and that borrowing an already-good representation saves that effort — SiT's training sped up by more than 17.5x. This “REPA-style” alignment has since been extended to video generation and robotics: Spatial Forcing, for example, aligns a VLA's intermediate visual tokens with features from the 3D foundation model VGGT, letting a VLA that has only ever seen 2D data implicitly pick up spatial awareness.

ExampleSpatial Forcing adds a cosine-similarity alignment loss against VGGT features on top of OpenVLA-OFT and π0, with no extra depth-map or point-cloud input, speeding up training by up to 3.8x and improving data efficiency.

Also called
REPA, REPA Regularization
Related
Representation Learning · Diffusion Transformer · DINOv2 · Auxiliary Loss / Auxiliary Task · VGGT · Pre-trained Visual Representation
Sources
Yu et al. 2024: Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
GitHub: sihyun-yu/REPA
Li et al. 2025: Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
As of
2025-10

See it in the full glossary →