06Perception & Sensors
How robots see and feel: cameras, depth, IMUs, force and touch, plus point clouds, calibration, and SLAM. · 307 terms
- 6.1Perception overview and cameras27
- 6.2Depth cameras and lidar40
- 6.3Proprioception and force sensing14
- 6.4Tactile sensing30
- 6.5Calibration and spatiotemporal alignment13
- 6.6Detection, segmentation, and tracking35
- 6.7Recovering 3D from images34
- 6.8Point clouds and 3D representations24
- 6.9Object pose, grasping, and affordance17
- 6.10Human body, hands, and interaction14
- 6.11State estimation and SLAM36
- 6.12Maps, semantics, and spatial intelligence23
6.1Perception overview and cameras
Perception splits into internal and external sensing; start with the everyday RGB camera — placement, how it images, and its specs.
- Proprioception本体感知A robot's sense of its own joint positions, velocities, forces, and posture, as opposed to sensing the outside world.
- Exteroception外部感知A robot’s sensing of its surrounding environment using sensors like cameras and lidar.
- Understanding a scene and task using vision together with touch, force, sound, and other senses at once.
- Computer Vision计算机视觉The field of making computers extract useful information from images and video and understand what a scene contains.
- Machine Vision机器视觉(工业视觉)Using cameras and software on a factory floor to automatically inspect, measure, identify, and guide robots.
- RGB CameraRGB相机An ordinary camera that outputs color images, recording red, green, and blue brightness at every pixel.
- Wrist Camera腕部相机A camera mounted on a robot arm's wrist or gripper that moves with the hand, giving a close-up view of the workspace.
- Third-Person Camera第三视角相机A camera mounted off the robot’s body, fixed in place, viewing the whole workspace from the side.
- Head Camera头部相机A camera mounted on a robot's head or upper body that gives a near-first-person view of the scene.
- Palm Camera手掌相机A camera mounted in a robot’s palm that gives a close-up view of the hand and object during grasping.
- A target being partly or fully blocked from a sensor's view by something else, including the robot's own body.
- Multi-View多视角Observing the same scene from multiple cameras or multiple angles at once.
- Pinhole Camera Model针孔相机模型A mathematical model treating a camera as an ideal tiny hole, projecting 3D points onto an image plane along straight lines.
- A 3D coordinate frame centered on the camera's optical center, with the z axis along the optical axis, used to locate objects relative to the camera.
- The parameters describing how a camera itself forms an image, determining which pixel a 3D point lands on.
- The rotation-plus-translation parameters describing where a camera sits in space and which way it points.
- Projection maps a 3D point onto a pixel; back-projection uses a pixel plus its depth to recover the 3D point.
- Lens Distortion镜头畸变The way a real lens's image deviates from the ideal pinhole model, bending straight lines into curves.
- The angular range a camera or sensor can see at once, usually given as horizontal, vertical, and diagonal figures.
- Fisheye Camera鱼眼相机A camera with an ultra-wide fisheye lens, often around 180° field of view, whose images bow visibly at the edges.
- CMOS Image SensorCMOS 图像传感器The chip inside a camera that turns light into a digital image, with amplification built into every pixel — the dominant sensor today.
- Image Signal ProcessorISP(图像信号处理器)The processing unit that turns an image sensor’s raw output into a normal color image.
- Global Shutter / Rolling Shutter全局快门 / 卷帘快门The two ways a camera can expose an image: all pixels at once, or row by row in sequence.
- MIPI CSI-2MIPI CSI-2 相机接口A high-speed serial interface standard that carries image data from a camera sensor to a processor.
- Gigabit Multimedia Serial LinkGMSL 相机接口An automotive-grade serial link that powers a camera and carries high-speed image data over a single coaxial cable.
- Event Camera事件相机A camera where each pixel independently outputs an “event” only when brightness changes, instead of capturing full frames.
- Thermal Camera热成像相机(红外热像仪)A camera that captures infrared heat radiation instead of visible light, revealing an object’s temperature distribution.
6.2Depth cameras and lidar
Adding distance on top of ordinary imaging: stereo, structured-light, and ToF depth cameras, plus lidar.
- 3D Vision3D视觉The umbrella term for techniques that recover an object's or scene's 3D geometry from images or sensor data.
- Depth Map深度图An image whose pixel values store distance to the camera instead of color.
- Data made of many 3D coordinate points used to represent the shape of an object or scene.
- Depth Camera深度相机A camera that outputs, for every pixel, both a color and the distance from that point to the camera.
- Stereo Camera双目相机Two side-by-side cameras that shoot the same scene at once, computing distance from how much the two images differ.
- The horizontal position difference of the same point between the left and right images — bigger for closer objects.
- The distance between a stereo camera’s two lens centers, which sets how far and how accurately it can measure depth.
- Stereo Matching立体匹配Finding the same point’s position in both the left and right images to compute disparity, then converting that into depth.
- Active Stereo主动双目A stereo camera pair with an infrared projector adding a speckle pattern to help match the left and right images and compute depth.
- RealSense Depth Camera (D435i / D405)RealSense 深度相机(D435i / D405)A widely used stereo depth camera series, formerly an Intel division, common in robotics research and products.
- Orbbec Gemini 330 Series奥比中光 Gemini 330 系列A stereo depth camera series from Orbbec that works indoors and out, commonly mounted on robots.
- Stereolabs ZEDZED 双目相机A passive stereo depth camera line from the French company Stereolabs, common in mobile robots and data collection.
- Luxonis OAK-DLuxonis OAK-D 相机A stereo depth camera from Luxonis with a built-in AI chip, letting neural networks run directly on the camera.
- Projects a known pattern onto a scene and computes depth from how the pattern gets distorted.
- Projecting a random infrared speckle pattern onto a scene and computing depth from how it deforms or matches.
- Laser Triangulation激光三角测量(线激光轮廓扫描)A laser projects a dot or line, a camera views the spot from the side, and triangulation gives distance and profile.
- Time of Flight飞行时间法Measures how long light takes to bounce off an object and return, then converts that time into distance.
- Direct Time of Flight直接飞行时间Emits a light pulse and directly times its echo, computing distance from the round-trip time.
- Indirect Time of Flight间接飞行时间A depth-sensing method that emits modulated light and computes distance from the phase shift of its echo.
- Azure Kinect DKAzure Kinect 深度相机A developer-oriented RGB-D camera from Microsoft that combines a depth sensor, color camera, microphone array, and IMU in one device.
- Orbbec Femto Bolt奥比中光 Femto BoltAn indirect time-of-flight RGB-D camera from Orbbec, positioned as a drop-in replacement for Microsoft’s discontinued Azure Kinect.
- Vertical-Cavity Surface-Emitting LaserVCSEL(垂直腔面发射激光器)A semiconductor laser that emits light perpendicular to the chip surface, a common light source in depth cameras and lidar.
- Single-Photon Avalanche DiodeSPAD(单光子雪崩二极管)A photodetector sensitive enough to register a single photon, at the core of dToF lidar and ranging chips.
- RoboSense AC1速腾聚创 AC1 主动相机A robot vision sensor from RoboSense that combines active depth sensing, a color camera, and an IMU in one module.
- Depth Holes深度空洞Pixels where a depth camera can’t measure distance, usually recorded as zero in the depth map.
- Erroneous depth points at object edges in a depth map, floating between the foreground and background.
- LiDAR激光雷达A sensor that fires lasers and times their return to measure distance, scanning out a 3D point cloud of its surroundings.
- 2D LiDAR2D激光雷达A spinning laser rangefinder that scans one horizontal plane, measuring the distance to obstacles all around.
- LiDAR Channels激光雷达线数The number of laser channels arranged vertically in a multi-line lidar, which sets how dense the point cloud is top to bottom.
- Solid-State LiDAR固态激光雷达Lidar with no bulk rotating part, scanning instead with electronics or a small internal mechanism.
- Frequency-Modulated Continuous-Wave LiDARFMCW 激光雷达(调频连续波激光雷达)A lidar that continuously emits frequency-swept laser light, measuring both distance and velocity at every point.
- Livox Mid-360览沃 Mid-360A small 360° hybrid solid-state lidar from Livox with a built-in IMU, widely used on mobile robots.
- Hesai JT128禾赛 JT128A compact, ultra-wide-field-of-view 128-channel 3D lidar from Hesai Technology aimed at robots.
- RoboSense Airy速腾聚创 AiryA hemispherical-field-of-view digital lidar from RoboSense for robots, covering 360°×90° with a single unit.
- Unitree 4D LiDAR L1 / L2宇树 4D 激光雷达 L1 / L2Unitree’s own low-cost, hemispherical-field-of-view lidar, where every point carries a 3D coordinate plus a grayscale value.
- A radar that transmits millimeter-wavelength electromagnetic waves to measure a target’s distance, speed, and bearing.
- Ultrasonic Sensor超声波传感器A cheap, short-range sensor that measures distance by emitting ultrasound and listening for its echo.
- Proximity Sensor接近觉传感器A sensor that detects whether something is nearby, and how close, without needing to touch it.
- Safety Laser Scanner / Safety Light Curtain安全激光扫描仪 / 安全光幕Certified safety devices that detect a person entering a danger zone and stop the machine before contact.
- Vision-Only Approach纯视觉方案Perceives the environment using only cameras, without lidar, millimeter-wave radar, or other ranging sensors.
6.3Proprioception and force sensing
From sensing the outside world to sensing itself: encoders and IMUs measure motion, force/torque sensors measure force and contact.
- A position sensor mounted on a motor or joint that converts shaft angle into an electrical signal.
- A sensor combining an accelerometer and gyroscope that measures an object's acceleration and rotation rate.
- Attitude Estimation姿态解算(AHRS 航姿参考系统)Fuses gyroscope, accelerometer, and (optionally) magnetometer data to compute roll, pitch, and yaw angles in real time.
- A simple filter that blends the gyroscope for short-term accuracy and the accelerometer for long-term stability to estimate attitude angle.
- Sensor Drift传感器漂移(零漂)A sensor’s reading slowly shifting over time or temperature even when the true input hasn’t changed.
- Allan VarianceAllan 方差(IMU 噪声标定)A statistical method for how sensor noise changes with averaging time, commonly used to calibrate IMU noise parameters.
- A sensor that measures force along three axes and torque around three axes at once, often mounted on an arm's wrist.
- Joint Torque Sensor关节力矩传感器A sensor built into a robot joint that directly measures the torque the joint is actually producing.
- Strain Gauge应变片A small element bonded to metal whose electrical resistance changes as it deforms, the basic building block of force sensors.
- When a multi-axis force sensor loaded along only one direction still shows nonzero readings on the other axes.
- Sensorless Force Estimation无传感器力估计Estimating the external force on a robot from motor current and a dynamics model, without installing a dedicated force sensor.
- Determining whether, when, and where a robot has touched an object or its environment.
- Foot Force Sensor足底力传感器A sensor mounted on a legged robot’s foot or ankle that measures the contact force between foot and ground.
- Determines whether each foot of a legged robot is currently on the ground or in the air.
6.4Tactile sensing
A finer-grained sense of touch than force sensors: tactile arrays, e-skin, vision-based tactile sensors, and what tactile data is used for.
- Tactile Sensor触觉传感器A sensor that lets a robot feel contact, measuring pressure distribution, contact location, shear force, and slip.
- Electronic Skin电子皮肤A soft, bendable, large-area tactile sensing layer applied to a robot's surface, mimicking human skin.
- Tactile Array触觉阵列Many small tactile sensing elements arranged in a grid, measuring the pressure distribution across a contact surface.
- Taxel触觉单元(触元)The smallest individual sensing element in a tactile array — the tactile equivalent of an image pixel.
- Measuring pressure from a material’s change in electrical resistance under load; cheap and easy to build into large sensor arrays.
- Capacitive Tactile Sensing电容式触觉传感A tactile-sensing method that infers contact force from the change in capacitance as pressure alters electrode spacing or a dielectric layer.
- Sensing contact and vibration using the piezoelectric effect, where a material generates an electric charge when it is deformed by force.
- Magnetic Tactile Sensing磁性触觉传感Mixes magnetic particles into a soft gel and uses a magnetometer to sense contact from the magnetic-field change deformation causes.
- A magnetic tactile skin: a soft pad embedded with magnetic particles sits over a magnetometer that reads contact from field changes.
- A magnetic tactile skin from XELA Robotics where every sensing point measures both pressure and shear force at once.
- PaXini PX-6AX帕西尼 PX-6AX 多维触觉传感器A multi-dimensional tactile sensor from PaXini that outputs distributed force, net force, and torque at the point of contact.
- SynTouch’s biomimetic fingertip tactile sensor, which senses contact force, micro-vibration, and temperature.
- A tactile sensor that uses a built-in camera to photograph how a soft surface deforms on contact.
- Keeping the camera fixed and lighting a surface from several directions, then recovering its shape from how the shading changes.
- A vision-based tactile sensor that reads touch from a camera watching an elastic gel deform, also the name of the company that makes it.
- GelSight’s commercial compact vision-based tactile sensor, ready to use over USB.
- DIGITDIGIT 视触觉传感器Meta's open-source, fingertip-sized, low-cost vision-based tactile sensor that reads contact from a camera watching a gel deform.
- Meta’s humanoid fingertip-shaped multimodal tactile sensor, sensing force, vibration, temperature, and more.
- A slim, finger-shaped vision-based tactile sensor from MIT that can be mounted on a gripper for grasping.
- An open-source, compact vision-based tactile sensor from Tsinghua’s Huazhe Xu group that measures contact shape and 6D force.
- Daimon DM-Tac Visuotactile Sensor戴盟 DM-Tac 视触觉传感器A line of vision-based tactile sensors from Daimon Robotics that senses force by imaging how a contact surface deforms.
- A biomimetic optical tactile fingertip from the Bristol Robotics Laboratory that uses a camera to watch internal marker pins move.
- Marker Tracking标记点跟踪Prints a dot pattern on a vision-based tactile sensor’s soft gel and tracks the dots’ displacement to estimate shear force and slip.
- Slip Detection滑移检测Determining whether an object held in a hand is starting to slide, so the grip can be adjusted in time.
- The perception task of inferring where and how an object held in the hand is making contact with the outside environment.
- Tactile Image触觉图像Arranging tactile sensor readings into a 2D image, so ordinary vision models can process them directly.
- A general-purpose vision-based-touch representation model from Meta, pretrained with self-supervision on a large set of tactile images.
- A unified tactile-representation model spanning multiple vision-based tactile sensors, from Renmin University’s Di Hu group, covering both static and dynamic touch.
- Contact Microphone接触式麦克风(音频触觉)A microphone stuck to a gripper or object that picks up contact vibration, used as a cheap tactile sensor.
- Visuo-Tactile Fusion视触觉融合Combining what a camera sees with what a tactile sensor feels into a single, unified signal.
6.5Calibration and spatiotemporal alignment
Now that you know the sensors, align them: camera and hand-eye calibration, multi-sensor extrinsics, and time synchronization.
- Estimating a camera's intrinsics, distortion coefficients, and extrinsics so pixels and 3D coordinates can convert into each other.
- A flat board printed with a pattern of precisely known size, used to give a camera calibration accurate reference points.
- Reprojection Error重投影误差The pixel distance between a 3D point reprojected onto the image and where it was actually observed.
- A printable black-and-white square marker that a camera can detect to compute the marker's ID and 6D pose.
- ArUco MarkerArUco码A square tag with a black border around a black-and-white grid code; a camera that spots it can compute its own pose relative to the tag.
- Perspective-n-PointPnP(透视n点)Recovering a camera’s pose from a set of known 3D points and where each one lands in the image.
- Finding the fixed coordinate transform that relates a camera to a robot arm.
- Eye-in-Hand眼在手上A camera-mounting scheme where the camera is fixed to the robot arm's end-effector and moves along with it.
- Eye-to-Hand眼在手外A camera-mounting scheme where the camera is fixed outside the robot and doesn't move with the arm, plus its matching calibration.
- Depth-to-Color Alignment深度与彩色对齐Reprojects a depth map into the color camera’s viewpoint so the two images’ pixels line up one to one.
- Camera-IMU Calibration相机-IMU联合标定Finds the relative pose and time offset between a camera and an IMU so their data can be aligned and fused.
- Camera-LiDAR Extrinsic Calibration相机-激光雷达联合标定Finds the rotation and translation between a lidar and a camera so a point cloud can be projected accurately onto the image.
- Making sure data from different sensors used together actually comes from the same moment in time, not slightly different ones.
6.6Detection, segmentation, and tracking
Entering vision algorithms: boxing, cutting out, and tracking objects in images, including targets specified by text.
- Object Detection目标检测Finding what objects appear in an image and where, marking each with a bounding box.
- Bounding Box检测框(边界框)The rectangle an object detector draws around an object, given as a few coordinates marking its location in the image.
- A family of real-time object detectors that locate every object in an image with a single pass through the network.
- The overlap area of two regions divided by their combined area, used to judge how accurate a predicted box or mask is.
- Non-Maximum Suppression非极大值抑制A post-processing step that keeps only the highest-scoring box among a cluster of overlapping detection boxes.
- An end-to-end detector that frames object detection directly as “set prediction” using a Transformer.
- Precision / Recall精确率 / 召回率Precision measures how many reported results were correct; recall measures how many of the true targets were found.
- Mean Average Precision平均精度均值The most common accuracy metric for object detection and segmentation: average precision (AP) per class, averaged across classes.
- Labels every pixel in an image with a category, like “table,” “cup,” or “floor.”
- Cutting out every object in an image pixel by pixel, keeping separate objects of the same class distinct from each other.
- A segmentation task that labels every pixel with a category while also telling individual object instances apart.
- Mask掩码A pixel-by-pixel map, the same size as the image, marking which pixels belong to a given object.
- An instance segmentation model from 2017, proposed by Kaiming He and colleagues, that outputs both detection boxes and a per-object mask.
- Segment Anything Model分割一切模型Meta's 2023 promptable segmentation foundation model: give it a point or box and it returns that object's mask.
- Keypoint Detection关键点检测Locating a handful of pre-defined, meaningful points in an image, like a wrist, a cup's handle, or a box's corner.
- Object Tracking目标跟踪Continuously finding the same object across video frames and keeping its identity consistent over time.
- Continuously cutting out a pixel mask for a specified object in every frame of a video.
- SAM 2SAM 2(视频分割一切)Meta's image-and-video segmentation model: mark a target once and it keeps segmenting and tracking it through the whole video.
- How far, and in which direction, every pixel appears to move between two consecutive video frames.
- RAFTRAFT 光流A classic optical-flow network that computes correlation between every pair of pixels, then refines the flow through repeated recurrent updates.
- Scene Flow场景流The 3D motion vector of every point in a scene between two adjacent frames.
- Tracking Any Point任意点跟踪Given any point in a video, outputting its position in every later frame, and whether it becomes occluded.
- A DeepMind point-tracking model that first finds a rough match per frame, then refines the trajectory over time.
- Meta’s open-source video point-tracking model that jointly tracks large numbers of pixels, including ones that get occluded.
- 3D Point Tracking3D 点跟踪Continuously tracks arbitrary pixels through a video and outputs their motion trajectory in 3D space.
- Detecting whatever object a piece of text names, instead of being limited to a fixed list of trained classes.
- An open-vocabulary detector that finds and boxes whatever object a text description names, given an image and text.
- A Google open-vocabulary object detection model that can find objects from a text description or an example image.
- An open-vocabulary detector that detects objects in real time just from a typed category name.
- DINO-XDINO-X(开放世界检测)IDEA Research’s open-world detection model that can box objects from text, examples, or no prompt at all.
- Segmenting exactly the pixels that match any text description of a category, even one the model never saw during training.
- Grounded SAMGrounded-SAMAn open-source pipeline that boxes and finely segments the object a text phrase refers to in an image.
- SAM 3SAM 3(可提示概念分割)Meta's third-generation segmentation model that finds every instance matching a short phrase or an example image.
- Visual Grounding视觉定位(Grounding)Finds the location in an image that a phrase or sentence refers to.
- Given a sentence describing one specific object, precisely segmenting out exactly that object in the image.
6.7Recovering 3D from images
Without a depth sensor: estimating depth from ordinary images, multi-view geometry, and networks that output 3D in one step.
- Depth Estimation深度估计Inferring how far every pixel in an image is from the camera, producing a depth map.
- Predicting how far away every pixel is from just one ordinary color photo.
- Metric Depth / Relative Depth度量深度 / 相对深度Metric depth gives true distance in meters; relative depth only gives near/far ordering, off by an unknown scale and offset.
- A monocular depth estimation foundation model from HKU and TikTok that produces a depth map from a single photo.
- Apple’s open-source monocular metric-depth model that outputs a sharp depth map in meters from a single image.
- A model that estimates metric depth from a single image, resolving scale ambiguity across cameras with a canonical camera-space transform.
- A Microsoft Research model that estimates full 3D geometry from a single image, predicting a 3D point for every pixel.
- NVIDIA’s stereo-matching foundation model that produces depth in a new scene with no fine-tuning needed.
- Using a cheap lidar’s sparse depth readings as a “prompt” so a large depth model outputs accurate 4K metric depth.
- Depth Completion深度补全Fills in missing or sparse pixels in a depth map to produce complete, dense depth.
- LingBot-Depth蚂蚁灵波 LingBot-DepthAn open-source depth-completion model from Ant Group’s Robbyant that uses a color image to repair a depth camera’s holes and noise.
- Camera Depth ModelsCDM 相机深度模型A depth-repair model from ByteDance Seed that cleans up noisy depth-camera output into near-simulation-quality accurate depth.
- Getting a robot to correctly see glass, metal, and other objects that ordinary depth cameras get wrong.
- Depth and segmentation datasets collected specifically for glass, metal, and other objects that depth cameras struggle to measure.
- Corners, blobs, and other points in an image that can be reliably found again, together with a vector describing their surroundings.
- Feature Matching特征匹配Finds pairs of feature points across two images that correspond to the same physical point.
- SuperPoint / SuperGlue / LightGlueSuperPoint / LightGlueA line of neural-network models for detecting and matching feature points across images, replacing hand-crafted methods like SIFT.
- Finds, for every pixel in one image, the matching location in another image.
- Random Sample Consensus随机采样一致性Repeatedly fitting a model to small random subsets of data and keeping the one most data points agree with, to reject outliers.
- The geometric relationship that forces the matching point of a scene point, seen by two cameras, to lie on a specific line.
- Homography单应性矩阵A 3×3 matrix describing how pixels on the same plane correspond between two images.
- Recovering a point’s 3D coordinates from two cameras’ known positions and where that point lands in each image.
- Bundle Adjustment光束法平差Jointly fine-tunes all camera poses and 3D point coordinates to minimize reprojection error.
- Structure from Motion运动恢复结构Computing both camera poses and a scene’s 3D points at once from a set of photos taken from different angles.
- Multi-View Stereo多视图立体A method that recovers a scene’s dense 3D structure from multiple photos whose camera positions are already known.
- A 3D reconstruction approach that outputs camera parameters, depth, and a point cloud directly from images in a single network forward pass.
- Pointmap点图An array the same size as an image, where each pixel stores a 3D coordinate instead of a color.
- A feed-forward reconstruction model that regresses a per-pixel 3D point map directly from two images, with no camera parameters needed.
- Naver’s model that adds a matching head onto DUSt3R, outputting both a 3D point map and dense matching features.
- A 3D vision model that computes camera parameters, depth, and a point cloud from multiple images in a single forward pass.
- π³ (Pi3)π³(Pi3)A feed-forward 3D reconstruction model with no reference viewpoint, giving the same result no matter what order the images come in.
- A 3D reconstruction model with a persistent memory state that reads images one at a time and updates the whole scene online.
- Meta and CMU’s general-purpose feed-forward 3D reconstruction model that directly outputs true-scale scene structure.
- ByteDance Seed’s geometry model that recovers depth and camera poses from any number of images, with poses optional.
6.8Point clouds and 3D representations
Once you have 3D data: downsampling, registering, and running networks on point clouds, then meshes, NeRF, and Gaussian splatting.
- Voxel体素A small cube in 3D space — the volumetric equivalent of a pixel.
- Point Cloud Filtering & Voxel Downsampling点云滤波与降采样(体素降采样 / 离群点去除)Preprocessing steps that crop irrelevant regions, remove noisy points, and shrink a raw point cloud down to a usable size.
- Repeatedly picks the point farthest from the already-selected set, downsampling a point cloud evenly to a fixed number of points.
- Computing which way each point on an object’s surface faces — the vector perpendicular to the surface there.
- Fast Point Feature HistogramsFPFH 点云特征A hand-crafted feature that describes each point’s local shape using a histogram of normal-vector geometric relationships in its neighborhood.
- Finding the rotation and translation that lines up two point clouds within one shared coordinate frame.
- A classic registration algorithm that repeatedly finds nearest-point pairs and solves for rotation and translation to align two point clouds.
- Normal Distributions TransformNDT 配准(正态分布变换)A point-cloud registration method that divides space into a grid, models each cell as a normal distribution, and aligns scans against it.
- Chamfer Distance倒角距离Finds each point’s nearest neighbor in the other point cloud and averages the distances, to measure how close two shapes are.
- PointNet / PointNet++PointNetA pioneering network that classifies and segments point clouds directly, with PointNet++ its hierarchical upgrade.
- A transformer backbone for point clouds that swaps costly neighbor search for a fast serialization trick, letting it see farther and run faster.
- Labeling every point in a point cloud with which object category, or which individual object, it belongs to.
- 3D Object Detection3D目标检测Finds objects in 3D space and outputs their category along with an oriented 3D box.
- Point Cloud Completion / Shape Completion点云补全 / 形状补全Given a partial point cloud, inferring and filling in the parts of an object that were occluded or never scanned.
- Triangle Mesh网格(三角网格)Represents a 3D object’s surface using vertices connected into triangular faces.
- Storing the signed distance to the nearest surface in a voxel grid, used to fuse multiple depth frames into a 3D model.
- Implicit vs. Explicit 3D Representation隐式表示 / 显式表示Whether a 3D scene is stored directly as geometric elements, or as a function you query by coordinate.
- A neural network that memorizes the color and density of every point in a scene, letting it render any new viewpoint.
- 3D Gaussian Splatting3D高斯泼溅Representing a scene as many colored 3D Gaussian ellipsoids that can be rendered from any new viewpoint in real time.
- Novel View Synthesis新视角合成Generating an image of a scene from a camera viewpoint that was never actually photographed, using a handful of existing photos.
- Reconstructs a 3D scene that moves — the geometry plus how it changes over time.
- Given just one photo, having a model fill in an object’s complete 3D shape and texture.
- Hunyuan3D混元3D(Hunyuan3D)Tencent’s open-source image-to-3D asset model series, generating a textured 3D mesh from a picture.
- Meta’s open single-image 3D reconstruction models, released as separate versions for objects and for human bodies.
6.9Object pose, grasping, and affordance
Putting 3D perception to work in manipulation: estimating 6D object pose, detecting grasps, and finding affordances and articulated structure.
- Computing an object's 3D position and 3D orientation relative to the camera — six degrees of freedom in total.
- NVIDIA's general-purpose 6D pose model that estimates and tracks the pose of new objects without retraining.
- A method that estimates 6D pose for a new object given only its CAD model, with no retraining needed.
- A method that uses SAM’s segmentation to estimate the 6D pose of objects it was never trained on, given only a CAD model.
- Pose Tracking位姿跟踪Continuously estimating an object’s 3D position and orientation across a video’s successive frames.
- Estimates 6D pose and size for an unseen object in a known category, without needing that exact object’s CAD model.
- Normalized Object Coordinate SpaceNOCS 归一化物体坐标空间A shared, standardized coordinate frame for all objects in a category, used to estimate the pose and size of objects never seen before.
- Average Distance of Model PointsADD / ADD-S 位姿误差指标Poses an object’s 3D model under both the true and predicted pose and averages the point-to-point distance, to score 6D pose estimation.
- A public benchmark and yearly challenge for 6D object pose estimation, maintained by the Czech Technical University.
- 3D Vision-Guided Robotics3D 视觉引导Uses a 3D camera to find a workpiece’s position and orientation and guides an industrial robot to pick or process it.
- Grasp Pose Detection抓取位姿检测Computing directly from an image or point cloud where and in what orientation a gripper should grasp an object.
- A general-purpose grasp-detection model from Shanghai Jiao Tong University that outputs many usable grasp poses directly from a point cloud.
- NVIDIA’s grasping network that generates 6-DoF grasps directly from a depth point cloud in cluttered scenes.
- NVIDIA’s framework that generates 6-DoF grasp poses with a diffusion model, then scores and filters them.
- Semantic Keypoints语义关键点Representing an object with a handful of meaningful points on it — a mug’s handle, a kettle’s spout — to make planning manipulation easier.
- Affordance Detection可供性检测Finds where on an object you can grip, press, or pour from an image or point cloud, and marks it to a specific region.
- Articulation Estimation铰接结构估计Infers an object’s parts, how they’re joined, which axis each part moves around, and how far it’s currently open.
6.10Human body, hands, and interaction
Shifting from objects to people: body and hand pose, 3D meshes, plus gesture, gaze, and speech.
- Human Pose Estimation人体姿态估计Finding the positions of a person's body joints from an image or video and connecting them into a skeleton.
- Hand Pose Estimation手部姿态估计Estimating the positions of a human hand's joints and how the fingers are bent, from an image or sensor data.
- MediaPipeMediaPipe(手部/人体关键点)Google's open-source on-device toolkit that extracts hand and body keypoints in real time from an ordinary camera feed.
- Recovering 3D human motion straight from ordinary video, with no reflective markers worn on the body.
- SMPLSMPL 人体模型A parametric 3D human body model that generates a full body mesh from a small set of shape and pose parameters.
- Human Mesh Recovery人体网格恢复The task of estimating a person’s complete 3D body mesh — pose plus body shape — from an image or video.
- A method that recovers a person’s 3D motion in world coordinates from monocular video.
- A Transformer model that reconstructs a 3D hand mesh from a single RGB image.
- Quickly finding every hand in an image and reconstructing each one as a 3D hand mesh.
- The study of how human hands contact, grasp, and manipulate objects, including estimating their joint 3D pose and contact.
- Lets a machine recognize a gesture a person makes from images or sensor signals and understand its meaning.
- Eye Tracking / Gaze Estimation眼动追踪 / 注视估计Measures the movement of a person’s eyes and estimates where they are looking right now.
- Microphone Array麦克风阵列Multiple microphones arranged in a fixed geometry working together to pick up sound directionally and judge where it comes from.
- Automatically converts spoken words into text — the first step in a robot understanding a spoken command.
6.11State estimation and SLAM
Answering ‘where am I’: from odometry and multi-sensor fusion to SLAM and localizing within a known map.
- State Estimation状态估计Estimates a robot’s current position, orientation, and velocity from noisy sensor readings.
- Multi-Sensor Fusion多传感器融合Combining data from a camera, lidar, IMU, and other sensors to get an estimate more accurate and stable than any one alone.
- Wheel Odometry轮式里程计Counts how much the wheels have turned to estimate how far a robot has traveled and how much it has turned.
- Leg Odometry腿式里程计Uses joint encoders, the IMU, and contact detection to work out how far and in what direction a legged robot has moved.
- Learned State Estimator学习型状态估计器A neural network that estimates quantities like body velocity from joint and IMU data, often trained jointly with a locomotion policy.
- Visual Odometry视觉里程计Compares consecutive camera frames to estimate, frame by frame, how far and how much a camera has turned.
- A robot building a map of an unfamiliar environment while simultaneously figuring out where it is on that map.
- SLAM Front-end / Back-endSLAM 前端 / 后端The two-part division of labor in a SLAM system: the front end estimates motion from sensor data, the back end globally corrects error.
- Lets a robot recognize that it has returned to a place it has already been, used to remove SLAM’s accumulated drift.
- Visual Place Recognition视觉位置识别Looking at an image and determining which place it is, and whether the system has been there before.
- Figuring out where a robot is and which way it’s facing on an existing map, after tracking is lost or it restarts.
- Visual SLAM视觉SLAMUses only cameras to simultaneously estimate where it is and build a map of the surroundings.
- A classic open-source visual SLAM system from the University of Zaragoza, supporting monocular, stereo, RGB-D, and IMU input.
- LiDAR SLAM激光SLAMUsing a lidar's scanned point clouds to localize a robot and build a map of the environment at the same time.
- LOAM (LiDAR Odometry and Mapping)LOAM 系列激光里程计A classic method that splits lidar SLAM into a high-frequency odometry step and a low-frequency mapping step, plus its lightweight variants.
- Occupancy Grid Map占据栅格地图A map that divides the environment into cells, each storing the probability that it's blocked by an obstacle.
- Particle Filter粒子滤波A recursive estimation method that approximates a probability distribution over states using a large set of weighted random samples.
- Adaptive Monte Carlo LocalizationAMCL 自适应蒙特卡洛定位A ROS localization module that estimates a robot's position on a known map using particle filtering and laser scans.
- GNSS / Real-Time Kinematic PositioningGNSS / RTK 定位Combines satellite navigation with base-station differential correction to achieve centimeter-level positioning outdoors.
- Ultra-Wideband PositioningUWB 定位(超宽带)Using extremely short radio pulses to measure signal flight time, for centimeter-level indoor ranging and positioning.
- Tightly-Coupled vs. Loosely-Coupled Fusion紧耦合 / 松耦合融合Two architectures for multi-sensor fusion: combining raw measurements together, versus computing each sensor’s own result and merging afterward.
- Error-State Kalman Filter误差状态卡尔曼滤波A Kalman filter formulation that estimates the difference between a nominal state and the true state, rather than the state itself.
- Unscented Kalman Filter无迹卡尔曼滤波A Kalman filter for nonlinear systems that propagates a small set of sample points instead of taking derivatives.
- Invariant Extended Kalman Filter不变扩展卡尔曼滤波An extended Kalman filter that defines error on a Lie group, converging more reliably, commonly used for legged-robot state estimation.
- Draws every measurement constraint as a “variable-factor” graph, then finds the most likely state by least squares.
- IMU PreintegrationIMU 预积分A technique that pre-integrates a large batch of IMU readings between two frames into a single relative-motion constraint.
- Visual-Inertial Odometry视觉惯性里程计Fusing a camera and an IMU to continuously estimate a device’s own motion trajectory.
- An open-source visual-inertial localization system from HKUST that fuses a camera and an IMU to estimate pose in real time.
- LiDAR-Inertial Odometry激光惯性里程计Fuses lidar point clouds with IMU readings to estimate a robot’s pose in real time while building a point-cloud map.
- FAST-LIO2FAST-LIO / FAST-LIO2An open-source lidar-plus-IMU odometry system from HKU that's fast and works with many lidar types.
- An open-source tightly coupled lidar-inertial odometry and mapping system built on factor-graph optimization.
- An open-source graph-based SLAM library that supports mapping and localization with RGB-D, stereo, or lidar sensors.
- A deep-learning visual SLAM system from Princeton that iteratively estimates camera pose and dense depth.
- A real-time monocular dense SLAM system built on MASt3R as a prior, able to map and localize from ordinary uncalibrated video.
- Gaussian Splatting SLAM高斯泼溅 SLAMA family of SLAM methods that use 3D Gaussian splatting as the map, building a photorealistically renderable scene while localizing.
- Absolute Trajectory Error / Relative Pose ErrorATE / RPE(绝对 / 相对轨迹误差)Two standard metrics for how far a SLAM or odometry system’s estimated trajectory deviates from ground truth.
6.12Maps, semantics, and spatial intelligence
Beyond localization, understanding the whole scene: geometric and semantic maps, scene graphs, and LLM-based spatial reasoning and active perception.
- Working out what's in an environment, where it is, and how things relate to each other from images, depth, or point clouds.
- OctoMap八叉树地图A 3D probabilistic occupancy map stored hierarchically as an octree, commonly used for robot obstacle avoidance and planning.
- An open-source, GPU-accelerated 3D mapping library from NVIDIA that builds real-time distance-field maps for obstacle avoidance.
- A neural network that predicts, for every location in 3D space, whether it is occupied by an object.
- A 2D top-down grid that fuses information from multiple sensors onto a single ground plane.
- A 2.5D terrain map that divides the ground into a grid and stores one height value per cell, commonly used by legged robots.
- Height Scan高度扫描Samples ground height around a robot in a fixed pattern, used as terrain input for a locomotion policy.
- Judging which parts of the terrain ahead can be crossed, which can’t, and how costly each path would be.
- Topological Map拓扑地图A map made of place nodes and connecting edges, recording only what leads where, not exact coordinates.
- Semantic Map语义地图A map that labels what each place is, on top of its geometry, so it can be queried by object name or natural language.
- Semantic SLAM语义SLAMSLAM that recognizes object categories while localizing and mapping, producing a map with semantic labels attached.
- 3D Scene Graph3D场景图Organizes a 3D scene’s floors, rooms, objects, and their relationships into a layered graph.
- Fuses multi-view images into an open-vocabulary 3D object scene graph for large-model reasoning and robot planning.
- Distills features from 2D models like CLIP and DINO into a 3D scene, so every point in space carries semantics.
- Embeds CLIP language features into a NeRF, so objects in a 3D scene can be found with natural language.
- 3D Visual Grounding3D视觉定位Finds the object in a 3D scene that a sentence describes.
- Given an image and a natural-language question about it, the model answers in words.
- Pointing指向(点预测)A vision-language model answering by marking exact pixel coordinates on the image, following a text instruction.
- Visual Prompting视觉提示Drawing boxes or numbers directly on an image so a multimodal large model can answer by referring to those marks.
- VSI-BenchVSI-Bench 空间智能基准A benchmark that uses indoor videos to test how well multimodal large models understand space.
- A robot deliberately moves its sensors or body to see or feel a target better, rather than just passively observing.
- Next-Best-View Planning下一最佳视角Deciding where a sensor should look next, given what it has already seen, to gain the most useful information.
- A robot actively pushes, pokes, or pulls objects, using the resulting changes to understand its environment.