Embodied AI Glossary中文

Human-Object Interaction

人-物交互HOICommon

Studying how people grasp, push, open, and use objects — an important source for robots to learn actions from.

Human-object interaction (HOI) studies what actions take place between a person and an object. In computer vision, the classic task is HOI detection: draw boxes around the person and the object in an image and determine the relationship between them, outputting a triple of “human, action, object,” such as “person, ride, bicycle”; a representative dataset is HICO-DET (WACV 2018). Research has since expanded to video, 3D, and 4D, focusing on how the hand contacts the object and how the object moves as a result; HOI4D (CVPR 2022), for example, contains 2.4 million frames of first-person RGB-D video across 800 object instances. For embodied AI, human interaction videos are far more abundant and far cheaper than robot data, so extracting contact locations, hand trajectories, and affordances (how an object can be used) from them is an important way to expand a robot's training data. The abbreviation HOI is sometimes also used specifically for hand-object interaction.

ExampleFrom a large collection of videos of people opening refrigerators, a system identifies the interaction “person, pull open, fridge door” and where the hand grips the handle, then uses this to teach a robot to open doors.

Also called
HOI
Related
Hand-Object Interaction · Affordance · Human Video Data · HOI4D · Egocentric Video · OMOMO
Sources
Learning to Detect Human-Object Interactions (HICO-DET, arXiv 1702.05448)
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction (arXiv 2203.01577)

See it in the full glossary →