CLIP
CommonAn OpenAI model trained on 400 million image-text pairs that maps pictures and text into one shared vector space.
CLIP is an image-text model OpenAI released in 2021, made of an image encoder and a text encoder trained together with contrastive learning on 400 million image-caption pairs collected from the internet: matching image-text pairs are pulled together in vector space, and mismatched ones are pushed apart. After training, images and text land in the same embedding space and their similarity can be computed directly, so it can do zero-shot classification with no further training at all — just write each category as a sentence and see which sentence the image is closest to. The paper reports this matched the zero-shot accuracy of a ResNet-50 trained on 1.28 million labeled images. In embodied AI, CLIP is often used as an open-vocabulary vision or language encoder: the early CLIPort used it to understand the semantics of objects named in an instruction, and many later VLMs and VLAs' vision encoders build on CLIP or its improved successor, SigLIP.
ExampleCrop a few small patches from a tabletop photo and encode both them and the sentence “a red cup” with CLIP; the patch whose embedding is most similar to that sentence is the object the instruction is pointing to.
- Also called
- Contrastive Language-Image Pre-training
- Related
- Contrastive Learning · SigLIP · Vision Encoder · Open-vocabulary · Embedding · CLIPort
- Sources
- Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv 2103.00020)
CLIPort: What and Where Pathways for Robotic Manipulation (arXiv 2109.12098)