Human Video Data
人类视频数据EssentialVideo of people doing everyday tasks, collectable without any robot, used to teach robots to understand and copy manipulation.
Human video data means video recording people performing all kinds of manipulation, including egocentric video, third-person (exocentric) video, and instructional videos pulled from the internet. Its biggest advantage is volume, low cost, and scene diversity, since collecting it doesn't depend on an expensive robot. The difficulty is that the video carries no action labels a robot can directly execute, and a human hand differs structurally from a robot hand or gripper — the embodiment gap. Common approaches: use hand-pose estimation to extract wrist and finger trajectories as actions; train a latent-action model to learn an abstract action from consecutive frames; or use the video only to pretrain visual representations and world models, then fine-tune with a small amount of real-robot data. NVIDIA's 2026 EgoScale pretrained a VLA (vision-language-action) model on more than 20,000 hours of action-annotated egocentric video, and found that data volume scales log-linearly with validation loss.
ExampleEgoScale first pretrains on human video, then runs mid-training — a transitional stage between pretraining and task fine-tuning — on a small amount of paired human-robot data; on a 22-degree-of-freedom dexterous hand, its average success rate came out 54% higher than skipping the pretraining step.
- Also called
- Human Video
- Related
- Egocentric Video · Internet Video Data · Action-free Video · Latent Action Pretraining · Embodiment Gap · Pretraining on Human Videos
- Sources
- EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data (NVIDIA GEAR)
EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video (arXiv 2505.11709) - As of
- 2026-02