Embodied AI Glossary中文

WiLoR

Advanced

Quickly finding every hand in an image and reconstructing each one as a 3D hand mesh.

WiLoR is a 3D hand reconstruction method from a team at Imperial College London and Shanghai Jiao Tong University, published at CVPR 2025. It works in two steps: a real-time, fully convolutional network first detects every hand in the image, and then a Vision Transformer-based reconstruction network regresses, coarse to fine, the MANO parameters (a parametric hand model that describes hand shape and pose with a small number of parameters) and camera parameters for each hand, producing a 3D hand mesh. The authors also curated WHIM, a dataset of more than 2 million in-the-wild hand images. With no temporal module at all, it produces reasonably smooth hand tracking frame by frame from monocular video; code, models, and data are all open-sourced. In embodied AI, it can be used to extract 3D hand poses from human video and then retarget them onto a dexterous hand as imitation-learning data.

ExampleA first-person video of a person folding laundry is fed into WiLoR frame by frame, producing MANO poses and 3D fingertip positions for both hands, which are then retargeted into joint targets for a dexterous hand.

Also called
End-to-End 3D Hand Localization and Reconstruction in-the-wild
Related
MANO · Hand Pose Estimation · HaMeR · Motion Retargeting · Human Video Data · Egocentric Video
Sources
WiLoR: End-to-end 3D Hand Localization and Reconstruction in-the-wild (arXiv)
WiLoR 项目主页 (Chinese)
As of
2025-03

See it in the full glossary →