Embodied AI Glossary中文

RAFT

RAFT 光流Advanced

A classic optical-flow network that computes correlation between every pair of pixels, then refines the flow through repeated recurrent updates.

RAFT was proposed by Zachary Teed and Jia Deng at Princeton in 2020, published at ECCV 2020, where it reportedly won the Best Paper Award. Optical flow describes how far every pixel moves between two adjacent frames. RAFT first computes feature correlation between every pair of pixels across the two frames, producing a 4D correlation volume, and then uses a module based on a GRU (a type of recurrent neural network unit) to repeatedly refine the flow at a single high resolution, replacing the traditional coarse-to-fine pyramid approach. According to the paper, it reduced error by 16% on KITTI and 30% on Sintel compared with prior work, and it remains a widely used baseline today. The same team’s RAFT-Stereo and DROID-SLAM both reuse this iterative-update idea. In robotics research, optical flow is often used to estimate object motion, or to infer motion from video that has no action labels.

ExampleGiven two frames of a robot arm before and after pushing a block, RAFT outputs each pixel’s displacement, revealing which way the block and the arm each moved.

Also called
Recurrent All-Pairs Field Transforms, RAFT Optical Flow
Related
Optical Flow · Scene Flow · Tracking Any Point · Stereo Matching · DROID-SLAM · Recurrent Neural Network
Sources
RAFT: Recurrent All-Pairs Field Transforms for Optical Flow (arXiv 2003.12039)
princeton-vl/RAFT (GitHub)

See it in the full glossary →