Embodied AI Glossary中文

Scene Flow

场景流Advanced

The 3D motion vector of every point in a scene between two adjacent frames.

Scene flow is the 3D version of optical flow: optical flow describes the 2D displacement of every pixel in an image between two frames, while scene flow describes the 3D displacement of every point in the real world. The concept was introduced by Vedula, Kanade, and colleagues at CMU in 1999, with a journal version published in TPAMI in 2005. Early methods solved for it from multiple viewpoints or stereo images; once depth cameras and lidar became common, networks that learn directly on two point clouds emerged, such as FlowNet3D in 2018. Scene flow tells a system which things are moving, in what direction, and how fast, and is used for motion segmentation and dynamic-object tracking in self-driving cars. In robot manipulation, predicting the future 3D motion trajectory of points on an object — often called 3D flow — is also used as an embodiment-agnostic intermediate representation, for learning a skill from human video and then transferring it to a robot.

ExampleGeneral Flow (2024) trains a model on human RGB-D video to predict the future 3D trajectory of points on an object given a language instruction, then converts that into robot actions for zero-shot skill transfer.

Also called
3D Optical Flow, 3D Flow
Related
Optical Flow · Tracking Any Point · Point Cloud · Intermediate Representation · RAFT · 4D Reconstruction
Sources
Three-Dimensional Scene Flow (Vedula et al., CMU RI)
arXiv 1806.01411: FlowNet3D
arXiv 1612.02590: Scene Flow Estimation: A Survey

See it in the full glossary →