DINOv3
CommonMeta's August 2025 third-generation DINO: 7 billion parameters, self-supervised on 1.7 billion images.
DINOv3 is a self-supervised vision foundation model Meta released in August 2025, the successor to DINOv2. Its largest model has 7 billion parameters, trained on 1.7 billion images, roughly 7 times the model scale and 12 times the data of its predecessor. When large models train for a long time, their dense, per-patch features tend to degrade, so DINOv3 introduces Gram anchoring to keep those features stable. Meta reports that, with no fine-tuning and only a lightweight task head attached, it beats specially trained models on tasks such as detection and semantic segmentation. Besides the 7-billion-parameter flagship, several smaller ViT and ConvNeXt models were distilled from it, along with a version trained on satellite imagery. For robotics, it can serve as a stronger frozen vision encoder than DINOv2, providing high-resolution, dense features.
ExampleFeed footage from a wrist-mounted camera into a frozen DINOv3, pull out each patch's feature vector, and train a lightweight head on top for object segmentation or grasp-point prediction, with no need to train a vision network from scratch.
- Related
- DINOv2 · Vision Foundation Model · Self-Supervised Learning · Vision Encoder · Pre-trained Visual Representation · Semantic Segmentation
- Sources
- DINOv3: Self-supervised learning for vision at unprecedented scale (Meta AI Blog)
DINOv3 (arXiv 2508.10104) - As of
- 2025-08