Point Cloud Encoder
点云编码器AdvancedA module that converts an unordered set of 3D points into feature vectors a neural network can use.
A point cloud encoder turns a point cloud from a depth camera or lidar — an unordered set of 3D points with xyz coordinates, and sometimes color — into features. The difficulty is that points have no fixed order and no fixed count, so the convolutions built for image grids don't directly apply. Stanford's 2016 PointNet extracts per-point features with a shared MLP, then aggregates them with an order-independent operation like max pooling, pioneering the direct-on-point-cloud approach; stronger architectures followed, such as PointNet++ and the Point Transformer series. Any policy that uses 3D input in embodied AI depends on this: 3D Diffusion Policy (DP3) first downsamples the point cloud to 512 or 1024 points with farthest point sampling, then uses a lightweight encoder of three MLP layers plus max pooling to get a 64-dimensional feature, and the paper's ablations show this outperforms more complex encoders like PointNet++.
ExampleDP3's point cloud encoder deliberately skips the color channel and uses only geometric coordinates, which the paper says generalizes better to changes in an object's appearance.
- Also called
- Point Cloud Backbone
- Related
- Point Cloud · PointNet / PointNet++ · Point Transformer V3 · 3D Diffusion Policy · Farthest Point Sampling · Vision Encoder
- Sources
- PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation (arXiv:1612.00593)
3D Diffusion Policy (arXiv:2403.03954)