PaliGemma
CommonGoogle's open-weight vision-language model, combining a SigLIP visual encoder with a Gemma language model.
PaliGemma is an open-weight vision-language model (VLM) Google released in May 2024, combining a SigLIP-So400m visual encoder with a Gemma-2B language model for about 3B parameters total. It isn't meant to be an out-of-the-box chat model; it's positioned as a base model meant to be transferred and fine-tuned, and the paper validates transfer performance on roughly 40 tasks. PaliGemma 2, released in December 2024, switched to the Gemma 2 language model and comes in three sizes (3B, 10B, 28B) and three input resolutions (224, 448, and 896 pixels). It's well known in embodied AI because Physical Intelligence uses it as the VLM backbone for π0: it's small, fast to run, and carries internet-scale vision-language knowledge, which makes it a good base for attaching an action expert to build a VLA.
Exampleπ0 uses the 3B-parameter PaliGemma as its backbone, adds a separate, randomly initialized action expert of about 300M parameters, for 3.3B parameters total, and generates continuous actions using flow matching.
- Also called
- PaliGemma 2
- Related
- Vision-Language Model · SigLIP · π0 · Action Expert · Vision-Language-Action Model · Open-weight Model
- Sources
- PaliGemma: A versatile 3B VLM for transfer (arXiv:2407.07726)
Google Developers Blog: Introducing PaliGemma 2
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164) - As of
- 2024-12