Florence-2
AdvancedMicrosoft's small open-source vision foundation model that switches between captioning, detection, and segmentation via text prompts.
Florence-2 is a vision foundation model Microsoft released in November 2023, open-sourced under the MIT license in two sizes: 0.23B (base) and 0.77B (large). It uses a sequence-to-sequence design — an image encoder plus a text decoder: given an image and a task prompt, such as <OD> for object detection or <CAPTION> for image captioning, the model outputs the result, including box coordinates, entirely as text, so a single model can do captioning, detection, phrase grounding, segmentation, OCR, and more. Its training data, FLD-5B, contains 126 million images and 5.4 billion annotations generated through an iterative automated pipeline. Because it's small and has strong spatial grounding ability, it's often used as the vision-language backbone for lightweight VLAs.
ExampleFLOWER (CoRL 2025) uses only half of Florence-2-L's layers as its backbone; X-VLA-0.9B also uses Florence-Large to encode the main-view image and the language instruction.
- Also called
- Florence-2-base, Florence-2-large
- Related
- Vision Foundation Model · Vision-Language Model · Open-Vocabulary Object Detection · Visual Grounding · X-VLA · Backbone Network
- Sources
- Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks (arXiv 2311.06242)
microsoft/Florence-2-large (Hugging Face model card)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment VLA (arXiv 2510.10274) - As of
- 2025-10