Embodied AI Glossary中文

SAM 2

SAM 2(视频分割一切)Common

Meta's image-and-video segmentation model: mark a target once and it keeps segmenting and tracking it through the whole video.

SAM 2 is Meta FAIR's second-generation ‘segment anything’ model, released in July 2024, extending SAM from single images to video: given a point, box, or mask specifying a target on one frame, the model continues to output that target's mask on all the frames that follow. Its core is a streaming memory: while processing frame by frame, it stores information about the target from previous frames in a memory bank that the current frame can reference, plus an occlusion head that judges whether the target is currently visible. It was released alongside the SA-V dataset (roughly 51,000 videos and more than 600,000 spatio-temporal masks). The paper reports that video segmentation needs three times fewer user interactions than before, and that image segmentation is more accurate and six times faster than the original SAM. Code and weights are open-sourced under Apache 2.0. Robots often combine it with Grounding DINO: find the object by text first, then track its mask throughout.

ExampleIn the Grounded-SAM-2 pipeline, Grounding DINO first detects a box for ‘red cup’ in the first frame, and SAM 2 then tracks the cup's mask throughout the manipulation video, for use in data labeling or feeding a target mask to the policy.

Also called
Segment Anything 2, SAM 2.1
Related
Segment Anything Model · SAM 3 · Video Object Segmentation · Object Tracking · Grounded SAM · Mask
Sources
SAM 2: Segment Anything in Images and Videos (arXiv 2408.00714)
Meta AI Blog: Introducing SAM 2
GitHub: IDEA-Research/Grounded-SAM-2
As of
2024-10

See it in the full glossary →