Embodied AI Glossary中文

SigLIP

Common

Google's image-text model trained with a sigmoid loss instead of softmax; its image encoder is used by many VLAs.

SigLIP is an image-text pretraining method proposed in 2023 by Xiaohua Zhai and colleagues at Google (ICCV 2023). Like CLIP, it trains an image encoder and a text encoder so that matched image-text pairs end up close together as vectors; the difference is the loss function. CLIP uses a softmax normalized across the whole batch, while SigLIP treats each image-text pair independently as a binary 'do they match or not' classification problem with a sigmoid loss, which doesn't depend on batch-wide normalization, so it trains well even with small batches and uses less GPU memory. Its vision encoder (such as the roughly 400-million-parameter So400m) is the image input stage for many VLMs and VLAs. SigLIP 2, released in February 2025, added objectives like captioning and self-distillation during training, improving multilingual ability, localization, and dense features.

Exampleπ0's backbone, PaliGemma, is made of a SigLIP-So400m vision encoder and a Gemma-2B language model; OpenVLA instead concatenates SigLIP features together with DINOv2 features.

Also called
Sigmoid Loss for Language-Image Pre-training, SigLIP 2, SigLIP-So400m
Related
CLIP · Contrastive Learning · Vision Encoder · PaliGemma · DINOv2 · Softmax
Sources
Sigmoid Loss for Language Image Pre-Training (arXiv 2303.15343)
SigLIP 2: Multilingual Vision-Language Encoders (arXiv 2502.14786)
PaliGemma: A versatile 3B VLM for transfer (arXiv 2407.07726)
As of
2025-02

See it in the full glossary →