VRB
VRB(从人类视频学可供性)AdvancedA method that learns 'where to grasp and which way to move afterward' from human video and hands that affordance directly to a robot.
VRB (Vision-Robotics Bridge) is work by Shikhar Bahl, Russell Mendonca, and colleagues in Deepak Pathak's group at Carnegie Mellon University, together with Meta AI, posted to arXiv in April 2023 and published at CVPR 2023. Affordance refers to what an object 'can be used for' — a drawer handle, for instance, affords pulling. VRB trains a visual affordance model on large amounts of first-person human video from datasets like EPIC-KITCHENS and Ego4D: given a scene image, it outputs two things — a contact heatmap (where a person is most likely to reach) and the wrist's post-contact motion trajectory (which direction the hand moves after making contact). Because this representation does not depend on any particular robot's body, it can plug into many different robot-learning setups: offline imitation learning, exploration, goal-conditioned learning, and as an action parameterization for reinforcement learning. The paper validated it across 4 real-world environments, more than 10 tasks, and 2 robot platforms, an early example of using human video to make up for scarce robot data.
ExampleGiven a kitchen photo, the model marks a high contact probability at the cabinet handle and predicts a trajectory of 'grasp, then pull outward'; the robot reaches for that spot and follows the trajectory to attempt opening the cabinet.
- Also called
- Vision-Robotics Bridge, VRB: Affordances from Human Videos as a Versatile Representation for Robotics
- Related
- Affordance · Human Video Data · Egocentric Video · EPIC-KITCHENS · Imitation from Observation · Affordance Detection
- Sources
- Affordances from Human Videos as a Versatile Representation for Robotics (arXiv 2304.08488)
VRB 项目主页 (Chinese) - As of
- 2023-06