Embodied AI Glossary中文

DINOv3

Common

Meta's August 2025 third-generation DINO: 7 billion parameters, self-supervised on 1.7 billion images.

DINOv3 is a self-supervised vision foundation model Meta released in August 2025, the successor to DINOv2. Its largest model has 7 billion parameters, trained on 1.7 billion images, roughly 7 times the model scale and 12 times the data of its predecessor. When large models train for a long time, their dense, per-patch features tend to degrade, so DINOv3 introduces Gram anchoring to keep those features stable. Meta reports that, with no fine-tuning and only a lightweight task head attached, it beats specially trained models on tasks such as detection and semantic segmentation. Besides the 7-billion-parameter flagship, several smaller ViT and ConvNeXt models were distilled from it, along with a version trained on satellite imagery. For robotics, it can serve as a stronger frozen vision encoder than DINOv2, providing high-resolution, dense features.

ExampleFeed footage from a wrist-mounted camera into a frozen DINOv3, pull out each patch's feature vector, and train a lightweight head on top for object segmentation or grasp-point prediction, with no need to train a vision network from scratch.

Related
DINOv2 · Vision Foundation Model · Self-Supervised Learning · Vision Encoder · Pre-trained Visual Representation · Semantic Segmentation
Sources
DINOv3: Self-supervised learning for vision at unprecedented scale (Meta AI Blog)
DINOv3 (arXiv 2508.10104)
As of
2025-08

See it in the full glossary →