Embodied AI Glossary中文

Distilled Feature Fields

蒸馏特征场DFFAdvanced

Distills features from 2D models like CLIP and DINO into a 3D scene, so every point in space carries semantics.

A distilled feature field adds a learned feature vector at every point in space to a neural radiance field (NeRF, a neural representation that reconstructs a 3D scene from multi-view photos) or a Gaussian splat, on top of color and density; the training objective is to make this feature, once rendered from any viewpoint, match the image features extracted by a 2D foundation model like CLIP or DINO — effectively distilling the 2D model’s knowledge into 3D. The name comes from a 2022 NeurIPS paper by Kobayashi, Sitzmann, and colleagues at Preferred Networks and MIT, originally used to select and edit objects in a NeRF by text or by clicking. It combines accurate geometry with the ability to locate objects in 3D space using natural language. In robotics, MIT’s F3RM uses CLIP-distilled feature fields for few-shot, language-guided 6-DoF grasping and placing, winning the CoRL 2023 best paper award; LERF takes a similar approach.

ExampleF3RM first takes a set of multi-view photos of a tabletop and reconstructs a scene carrying CLIP features; with just two demonstrations for a task like “grasp the cup by the rim,” it can then grasp objects of shapes and categories it has never seen, based on a text instruction.

Also called
DFF, Feature Fields, Language-Embedded Fields
Related
F3RM · LERF · Neural Radiance Fields · CLIP · Knowledge Distillation · 3D Gaussian Splatting
Sources
arXiv 2205.15585: Decomposing NeRF for Editing via Feature Field Distillation
arXiv 2308.07931: Distilled Feature Fields Enable Few-Shot Language-Guided Manipulation
F3RM project page

See it in the full glossary →