Embodied AI Glossary中文

HOI4D

HOI4D 数据集Advanced

A first-person 4D hand-object interaction dataset with 2.4 million RGB-D frames, from Tsinghua and collaborators.

HOI4D was released jointly by Tsinghua University, Peking University, and the Shanghai Qi Zhi Institute, published at CVPR 2022. Collectors wore a helmet fitted with two RGB-D cameras, a Kinect v2 and a RealSense D455, recording themselves manipulating objects from a first-person viewpoint — 4,000 sequences and 2.4 million frames in total, covering 16 categories and 800 object instances (7 rigid-body categories and 9 categories of articulated objects with moving parts, such as laptops, cabinets, and scissors), across 610 indoor rooms. Frame-by-frame annotations include panoptic segmentation, motion segmentation, 3D hand pose, category-level object pose (estimating pose when only the object's category, not the specific instance, has been seen before), and action labels, along with object meshes and scene point clouds. Its purpose is to let a model learn “how a hand interacts with a category of object” from a human viewpoint, useful for hand-object interaction understanding and pose tracking, and it also serves as material for robots learning manipulation from human data.

ExampleUsing the “open a laptop” sequences in HOI4D to train category-level pose tracking lets a model keep estimating the pose of a laptop it has never seen before, even while the hand partly blocks it from view.

Also called
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction
Related
Hand-Object Interaction · Egocentric Video · Category-Level Pose Estimation · Hand Pose Estimation · Articulated Object · HOT3D
Sources
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction (arXiv)
HOI4D 项目主页 (Chinese)

See it in the full glossary →