Mask
掩码CommonA pixel-by-pixel map, the same size as the image, marking which pixels belong to a given object.
A mask is an image the same size as the original, where each pixel takes a value of 0 or 1 (or a class ID), marking which pixels belong to a given object. The output of a segmentation model is essentially a mask: semantic segmentation produces one per class, and instance segmentation produces one per object. A mask is more precise than a bounding box, which often mixes in background, since a mask hugs the object's actual outline. For storage, the COCO dataset uses polygon vertices for a single object and run-length encoding (RLE) compression for groups of objects. Meta's SAM was trained on more than a billion annotated masks across 11 million images, and can produce a mask for essentially any object from just a click or a box. Robots commonly use a mask to pull just the target object out of a depth map or point cloud before computing a grasp. Note that the ‘attention mask’ used in transformers is an unrelated concept — don't confuse the two.
ExampleGiven the instruction ‘grasp the red cup,’ Grounded-SAM first produces the cup's mask, overlays it on the aligned depth map, back-projects only the cup's pixels into a point cloud, and hands that to the grasp-detection network.
- Also called
- Segmentation Mask, Binary Mask
- Related
- Instance Segmentation · Semantic Segmentation · Segment Anything Model · Grounded SAM · Attention Mask · Point Cloud Segmentation
- Sources
- Segment Anything (arXiv:2304.02643)
COCO Data Format(segmentation: polygon / RLE)