HaMeR
AdvancedA Transformer model that reconstructs a 3D hand mesh from a single RGB image.
HaMeR was proposed by researchers at UC Berkeley, the University of Michigan, and NYU, published at CVPR 2024. It uses a ViT-H vision Transformer as its backbone, regressing MANO hand model parameters (MANO describes a hand’s 3D mesh with a small number of pose and shape parameters) plus camera parameters from the hand region of an image. The authors combined 10 datasets with 2D or 3D hand annotations into about 2.7 million training samples, and also annotated the HInt evaluation set from videos like Ego4D, specifically to test hands in real-world footage. It is a single-frame method, but its results are reasonably smooth when applied to video too. In embodied AI, it is a commonly used tool for extracting finger poses from human video, with results retargeted to a dexterous hand or gripper as an action source for imitation learning.
ExampleWhen OKAMI teaches a humanoid robot to manipulate objects from a human demonstration video, it reconstructs body motion with SLAHMR while estimating each hand’s pose separately with HaMeR.
- Also called
- Hand Mesh Recovery, Reconstructing Hands in 3D with Transformers
- Related
- MANO · Hand Pose Estimation · WiLoR · Human Mesh Recovery · OKAMI · Human Video Data
- Sources
- arXiv 2312.05251: Reconstructing Hands in 3D with Transformers
HaMeR 项目主页 (Chinese)
OKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video Imitation - As of
- 2024-06