Embodied AI Glossary中文

Point Transformer V3

PTv3Advanced

A transformer backbone for point clouds that swaps costly neighbor search for a fast serialization trick, letting it see farther and run faster.

Point Transformer V3 (PTv3) was proposed by Xiaoyang Wu, Hengshuang Zhao, and colleagues from institutions including the University of Hong Kong and the Shanghai AI Laboratory; it was an oral paper at CVPR 2024, with code integrated into the open-source point cloud framework Pointcept. Earlier point cloud transformers had to search for each point’s neighbors and compute complex relative position encodings, which was slow and memory-hungry. PTv3 instead first arranges the unordered points into a 1D sequence using a space-filling curve (“serialization”), then computes attention within chunks of that sequence. According to the paper, it is about 3x faster and uses about 10x less memory than PTv2, expands the receptive field from 16 points to 1024 points, and achieved state-of-the-art results on more than 20 indoor and outdoor tasks at the time. It is commonly used as a point cloud encoder, and the self-supervised pretraining model Sonata also uses it as a backbone.

ExampleAn indoor point cloud from an RGB-D camera is fed into PTv3, which outputs a semantic category for every point — wall, floor, chair, and so on.

Also called
PTv3, Point Transformer V3: Simpler, Faster, Stronger
Related
Point Cloud Encoder · Point Cloud Segmentation · PointNet / PointNet++ · Transformer · Backbone Network · Self-Attention
Sources
Point Transformer V3: Simpler, Faster, Stronger (arXiv 2312.10035)
Pointcept/PointTransformerV3 (GitHub)
As of
2024-06

See it in the full glossary →