Embodied AI Glossary中文

3D Object Detection

3D目标检测Advanced

Finds objects in 3D space and outputs their category along with an oriented 3D box.

3D object detection extends 2D object detection into three dimensions: the input can be a lidar point cloud, an RGB-D image, or even a single color image, and the output is each object’s category plus a 3D bounding box, usually described by a center coordinate, length, width, height, and orientation angle. A 2D box only says which pixels an object occupies; robots that need to grasp, avoid, or navigate around objects need to know their position, size, and orientation in real space. Autonomous driving has been the main driver of this field — the classic KITTI benchmark has 7,481 training images with matching point clouds, evaluating cars, pedestrians, and cyclists, where a car’s 3D box needs an intersection-over-union (IoU) of 0.7 to count as correct. For indoor and general scenes there’s the Omni3D benchmark (234,000 images, 98 categories) with its companion Cube R-CNN model. 3D object detection is often contrasted with 6D pose estimation, which is more fine-grained and gives an object’s full 3D rotation.

ExampleA warehouse robot runs 3D detection on a lidar point cloud to get each crate’s center position, size, and orientation, then plans where to slide a fork or place a grasp accordingly.

Also called
3D Bounding Box Detection, 3D Detection
Related
Object Detection · Point Cloud · LiDAR · Bird’s-Eye View · Intersection over Union · 6D Object Pose Estimation
Sources
KITTI 3D Object Detection Evaluation 2017
Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild (arXiv 2207.10660)

See it in the full glossary →