Embodied AI Glossary中文

Spatial Softmax

空间 SoftmaxAdvanced

Pooling each channel of a convolutional feature map down to the 2D coordinate of its 'brightest point,' preserving location.

Spatial Softmax is a pooling layer that turns a convolutional feature map into coordinates, proposed by Berkeley's Levine, Finn, Darrell, and Abbeel in their 2015 work on end-to-end visuomotor policies. The procedure applies softmax across all pixel positions within each channel, producing a probability distribution for 'where this feature appears,' then takes the expectation over pixel coordinates to get one (x, y) pair — effectively a differentiable argmax. Ordinary classification networks use global average pooling, which erases location information, but robot manipulation specifically needs to know where things are; softmax also suppresses weak false activations, making it more robust to distractor objects. It later became a common choice in imitation-learning vision encoders — Diffusion Policy uses it at the end of its ResNet-18 in place of global average pooling.

ExampleLevine and colleagues' policy network follows three convolutional layers with a spatial softmax; the last layer's 32 channels each output one feature-point coordinate, which is concatenated with the robot's joint state and passed through fully connected layers to output motor torques, with the whole network having only about 92,000 parameters.

Also called
Spatial Soft-Argmax, Soft-Argmax, Keypoint Pooling
Related
Convolutional Neural Network · Vision Encoder · Visuomotor Policy · Diffusion Policy · Keypoint Detection · End-to-End Training of Deep Visuomotor Policies
Sources
End-to-End Training of Deep Visuomotor Policies (arXiv:1504.00702, JMLR 2016)
Deep Spatial Autoencoders for Visuomotor Learning (arXiv:1509.06113)
Diffusion Policy (arXiv:2303.04137, HTML)

See it in the full glossary →