Pseudo Action Labels
伪动作标签AdvancedAction labels inferred by a model from video, standing in for actions that were never actually recorded.
When video itself doesn't record what action a human or robot took, a model can “guess” the action at each step from the surrounding frames, and those guesses can be used as labels to train a policy — this is a pseudo action label. There are two common approaches. One trains an inverse-dynamics model (which infers the action between two adjacent frames) on a small amount of action-labeled data, then uses it to label a huge amount of video in bulk — OpenAI's VPT did exactly this to add keyboard-and-mouse action labels to a large collection of Minecraft videos from the internet. The other uses a latent-action model to learn an abstract action encoding directly from video. NVIDIA's DreamGen uses both approaches to add pseudo actions to video generated by a world model, producing trainable “neural trajectories.” This lets video with no action labels be used for training too, though the labels carry error and usually still need fine-tuning on real data.
ExampleVPT first had people play Minecraft while recording their keyboard and mouse actions, trained an inverse-dynamics model on that data, then used it to label pseudo actions across a large collection of internet videos, and finally ran behavioral cloning on that labeled data.
- Also called
- Pseudo-Action Labeling, Pseudo-Actions
- Related
- Inverse Dynamics Model · Latent Action Model · Action-free Video · Neural Trajectories · DreamGen · VPT
- Sources
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos (arXiv:2206.11795)
DreamGen: Unlocking Generalization in Robot Learning through Video World Models (arXiv:2505.12705)