Embodied Perception
具身感知AdvancedPerception in service of action: observing while moving, understanding 3D space, and supporting decisions.
Embodied perception refers to the perception an agent situated in an environment carries out in order to act. Unlike traditional computer vision, which takes a single image and outputs a label, its input is a first-person observation that keeps changing as the body moves, and it has to answer questions like “where am I, what is the 3D structure around me, where should I look next, and how should I move.” The 2024 embodied-AI survey from Sun Yat-sen University and others lists it as one of four research directions, covering tasks such as visual SLAM (simultaneous localization and mapping), 3D scene understanding, active exploration, and vision-and-language navigation. Its key feature is that it is active: the agent can turn its head, move closer, or shift an occluding object out of the way to gather more information, rather than passively receiving data. A robot's depth cameras, lidar, and tactile sensors, along with representations such as point clouds and semantic maps, all serve embodied perception.
ExampleWhen a robot cannot find a remote control on a table, it changes its viewing angle or moves aside a magazine blocking its view, instead of only running detection on the current single frame.
- Related
- Active Perception · Interactive Perception · Simultaneous Localization and Mapping · Scene Understanding · Embodied Interaction · Multimodal Perception
- Sources
- Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI (Liu et al., 2024)