Semantic Segmentation
语义分割CommonLabels every pixel in an image with a category, like “table,” “cup,” or “floor.”
Semantic segmentation is a core computer vision task: deciding which category every pixel in an image belongs to, and producing a category map the same size as the original image. It only cares about class, not individual identity — two cups both get labeled “cup” without being told apart; separating individual objects is instance segmentation, and combining the two is panoptic segmentation. In 2015, Long and colleagues introduced the fully convolutional network (FCN), which let a network predict pixel labels end to end, and this became the standard approach. Robots use semantic segmentation to find the region of a graspable object, identify walkable floor, or project labels onto a 3D point cloud to build a semantic map. Early models could only recognize the fixed set of categories they were trained on; today, open-vocabulary segmentation is common, letting a model segment any category described in text.
ExampleGiven a photo of a kitchen, the table pixels are labeled “table,” both cups are labeled “cup,” and everything else is “background”; a robot then pulls out the point cloud corresponding to the “cup” region to plan a grasp.
- Also called
- Semantic Seg, Pixel-Level Classification
- Related
- Instance Segmentation · Panoptic Segmentation · Open-Vocabulary Segmentation · Mask · Semantic Map · Segment Anything Model
- Sources
- Fully Convolutional Networks for Semantic Segmentation (Long, Shelhamer, Darrell, CVPR 2015)
Wikipedia: Image segmentation