Embodied AI Glossary中文

Bounding Box

检测框(边界框)BBoxCommon

The rectangle an object detector draws around an object, given as a few coordinates marking its location in the image.

A bounding box is the most common output of an object detector: a rectangle enclosing an object, usually paired with a class label and a confidence score. Two representations are common — top-left plus bottom-right corners (x1, y1, x2, y2), or center plus width and height (cx, cy, w, h) — and different datasets follow different conventions, so mixing them up is a common bug. How well a predicted box matches the ground truth is judged with intersection over union (IoU), and duplicate boxes are removed with non-maximum suppression. Extended forms include rotated boxes with an added angle and 3D bounding boxes. In embodied AI, a bounding box is often an intermediate result: an open-vocabulary detector finds a box from text first, and it's then handed to SAM for a mask or used to crop a region for grasping. This is a different concept from the ‘bounding box’ used for collision checking in physics engines.

ExampleGiven the instruction ‘pick up the red cup,’ Grounding DINO returns a box with a confidence around 0.6; downstream modules then only segment and search for grasp points within that box.

Also called
BBox, Detection Box
Related
Object Detection · Intersection over Union · Non-Maximum Suppression · 3D Object Detection · Grounding DINO · Mask
Sources
Dive into Deep Learning: Object Detection and Bounding Boxes

See it in the full glossary →