Embodied AI Glossary中文

Segment Anything Model

分割一切模型SAMCommon

Meta's 2023 promptable segmentation foundation model: give it a point or box and it returns that object's mask.

The Segment Anything Model (SAM) is an image segmentation foundation model released by Meta AI in April 2023, introducing the task of ‘promptable segmentation’: a user gives a point, a box, or a rough mask as a prompt, and the model outputs a mask for the corresponding object — even for objects it has never seen — and can also automatically segment an entire image on its own. A large ViT image encoder runs once per image, and a lightweight decoder then answers multiple prompts quickly. Its training data, SA-1B, contains 11 million images and more than 1.1 billion masks, and the code and weights are open-sourced under Apache 2.0. The original SAM gives only a mask with no class label, so it's often chained with Grounding DINO into a pipeline called Grounded-SAM for text-based segmentation. It was later followed by the video-focused SAM 2 (2024) and SAM 3 (2025), which supports text-phrase prompts.

ExampleClicking on a screwdriver in a workbench image, SAM returns its pixel mask; back-projecting the depth pixels inside that mask gives a point cloud containing only the screwdriver, ready for grasp detection.

Also called
SAM, Segment Anything
Related
SAM 2 · SAM 3 · Grounded SAM · Instance Segmentation · Mask · Foundation Model
Sources
Segment Anything (arXiv 2304.02643)
GitHub: facebookresearch/segment-anything
Meta AI Blog: Segment Anything Model 3
As of
2025-11

See it in the full glossary →