Embodied AI Glossary中文

Feed-Forward 3D Reconstruction

前馈式三维重建Advanced

A 3D reconstruction approach that outputs camera parameters, depth, and a point cloud directly from images in a single network forward pass.

Traditional 3D reconstruction follows a structure from motion (SfM) plus multi-view stereo (MVS) pipeline: match feature points, estimate camera poses, then iteratively refine everything with bundle adjustment — a process with many steps that takes a long time and easily fails with little texture or few viewpoints. Feed-forward 3D reconstruction instead uses a large network (usually a Transformer) trained on large-scale 3D data; given one or more images, a single forward pass directly outputs a point map (a 3D coordinate for every pixel), depth, and camera intrinsics and extrinsics. Landmark examples include Naver’s DUSt3R (CVPR 2024, needing no prior camera calibration), Meta and Oxford’s VGGT (CVPR 2025 best paper), π³, and MapAnything. For robotics, this lets scene geometry be recovered quickly from ordinary RGB images, useful for mapping, pose estimation, or feeding 3D input to a policy.

ExampleVGGT takes anywhere from one to several hundred photos of the same scene and, according to the paper, can directly predict camera parameters, depth maps, point maps, and 3D point trajectories in under a second, with no post-hoc optimization like bundle adjustment needed.

Also called
3D Reconstruction Foundation Model
Related
DUSt3R · VGGT · π³ (Pi3) · MapAnything · Pointmap · Structure from Motion
Sources
DUSt3R: Geometric 3D Vision Made Easy
VGGT: Visual Geometry Grounded Transformer
MapAnything: Universal Feed-Forward Metric 3D Reconstruction
As of
2025-09

See it in the full glossary →