Embodied AI Glossary中文

4D Reconstruction

4D重建Advanced

Reconstructs a 3D scene that moves — the geometry plus how it changes over time.

4D reconstruction recovers a 3D scene that changes over time from video, monocular or multi-view — the fourth dimension is time: it needs to capture not just what the scene looks like, but the shape and position of people, hands, and objects at every moment. Static-reconstruction methods like NeRF and 3D Gaussian splatting assume the scene doesn’t move, and produce ghosting artifacts on moving objects. One family of approaches optimizes per scene: Dynamic 3D Gaussians lets Gaussian points move and rotate over time, and 4D Gaussian Splatting (CVPR 2024) uses a deformation field to predict each Gaussian’s displacement at every moment, rendering 800×800 frames at 82 FPS on an RTX 3090. Another family is feed-forward — MonST3R (ICLR 2025) extends the static reconstruction model DUSt3R to dynamic video, outputting point maps directly frame by frame. In embodied AI, 4D reconstruction is used to recover the 3D motion of hand-object interaction from human videos and to build replayable digital twins.

ExampleFilming someone pouring water with a phone and running 4D reconstruction recovers the 3D shape of the hand, cup, and pitcher at every frame, which can be replayed from any viewpoint or mined for the cup’s motion trajectory to have a robot imitate.

Also called
Dynamic Scene Reconstruction, 4D Gaussian Splatting
Related
3D Gaussian Splatting · Neural Radiance Fields · Feed-Forward 3D Reconstruction · DUSt3R · 3D Point Tracking · 4D World Model
Sources
4D Gaussian Splatting for Real-Time Dynamic Scene Rendering (arXiv 2310.08528)
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion (arXiv 2410.03825)
Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis (arXiv 2308.09713)

See it in the full glossary →