Multi-Sensor Fusion
多传感器融合CommonCombining data from a camera, lidar, IMU, and other sensors to get an estimate more accurate and stable than any one alone.
Multi-sensor fusion combines data from a camera, lidar, an IMU (which measures acceleration and angular velocity), joint encoders, tactile sensors, and more to jointly estimate the same quantity, reducing the uncertainty below what any single sensor could give alone. Fusion is commonly grouped by the stage at which it happens: data-level (or early fusion, merging raw data directly), feature-level (extracting features from each sensor separately and merging those), and decision-level (or late fusion, where each sensor gives its own result and the results are voted or weighted together). The classic approach uses probabilistic estimators from the Kalman-filter family, though neural networks that learn how to fuse the data directly are now also common. Fusion matters because every sensor has blind spots: cameras struggle in the dark and with glare, IMUs drift, and lidar carries no color. Robot state estimation, SLAM, and self-driving perception all depend on it, with the prerequisite that the sensors have already been calibrated and time-synchronized.
ExampleA quadruped estimating its own velocity: the IMU gives high-frequency acceleration and angular velocity, joint encoders combined with foot-contact detection give a leg-odometry estimate, and the two are merged with an extended Kalman filter to reduce the drift either source alone would have.
- Also called
- Sensor Fusion, Early Fusion / Late Fusion
- Related
- Multimodal Perception · Kalman Filter · Extended Kalman Filter · Tightly-Coupled vs. Loosely-Coupled Fusion · Visual-Inertial Odometry · Multi-Sensor Time Synchronization
- Sources
- Wikipedia: Sensor fusion
ORB-SLAM3: An Accurate Open-Source Library for Visual, Visual-Inertial and Multi-Map SLAM (arXiv 2007.11898)