EgoScale
AdvancedAn NVIDIA 2026 project that pretrains a dexterous-hand VLA on 20,000 hours of first-person human video.
EgoScale is a paper released in February 2026 by NVIDIA's GEAR Lab together with UC Berkeley and the University of Maryland, aimed at answering whether human video can teach a robot to work with a multi-fingered dexterous hand at scale. The method has three steps: first, extract wrist motion from 20,854 hours of action-annotated, first-person human video, and retarget human hand poses into the joint space of a 22-degree-of-freedom Sharpa dexterous hand, to pretrain a flow-matching VLA structurally similar to GR00T N1; next, mid-train on a small amount of aligned data where a human and a robot perform the same action in the same scene; finally, post-train on specific tasks. The paper finds a log-linear scaling relationship between the amount of human data and validation loss, with average success rate improving 54% compared to skipping human pretraining.
ExampleOn a Galaxea R1 Pro fitted with the 22-DOF Sharpa dexterous hand, given just 1 robot demonstration, average success on a shirt-folding task reached as high as 88%; switching to a three-fingered Unitree G1 hand, human pretraining still delivered more than a 30-percentage-point absolute improvement.
- Also called
- Scaling Dexterous Manipulation with Diverse Egocentric Human Data
- Related
- Egocentric Video · Pretraining on Human Videos · Dexterous Manipulation · Scaling Law · Mid-training · NVIDIA Isaac GR00T N1
- Sources
- EgoScale: Scaling Dexterous Manipulation with Diverse Egocentric Human Data (arXiv 2602.16710)
EgoScale 项目页(NVIDIA GEAR) (Chinese) - As of
- 2026-02