Distilled Feature Fields
蒸馏特征场DFFAdvancedDistills features from 2D models like CLIP and DINO into a 3D scene, so every point in space carries semantics.
A distilled feature field adds a learned feature vector at every point in space to a neural radiance field (NeRF, a neural representation that reconstructs a 3D scene from multi-view photos) or a Gaussian splat, on top of color and density; the training objective is to make this feature, once rendered from any viewpoint, match the image features extracted by a 2D foundation model like CLIP or DINO — effectively distilling the 2D model’s knowledge into 3D. The name comes from a 2022 NeurIPS paper by Kobayashi, Sitzmann, and colleagues at Preferred Networks and MIT, originally used to select and edit objects in a NeRF by text or by clicking. It combines accurate geometry with the ability to locate objects in 3D space using natural language. In robotics, MIT’s F3RM uses CLIP-distilled feature fields for few-shot, language-guided 6-DoF grasping and placing, winning the CoRL 2023 best paper award; LERF takes a similar approach.
ExampleF3RM first takes a set of multi-view photos of a tabletop and reconstructs a scene carrying CLIP features; with just two demonstrations for a task like “grasp the cup by the rim,” it can then grasp objects of shapes and categories it has never seen, based on a text instruction.
- Also called
- DFF, Feature Fields, Language-Embedded Fields
- Related
- F3RM · LERF · Neural Radiance Fields · CLIP · Knowledge Distillation · 3D Gaussian Splatting
- Sources
- arXiv 2205.15585: Decomposing NeRF for Editing via Feature Field Distillation
arXiv 2308.07931: Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
F3RM project page