Spatial Softmax
空间 SoftmaxAdvancedPooling each channel of a convolutional feature map down to the 2D coordinate of its 'brightest point,' preserving location.
Spatial Softmax is a pooling layer that turns a convolutional feature map into coordinates, proposed by Berkeley's Levine, Finn, Darrell, and Abbeel in their 2015 work on end-to-end visuomotor policies. The procedure applies softmax across all pixel positions within each channel, producing a probability distribution for 'where this feature appears,' then takes the expectation over pixel coordinates to get one (x, y) pair — effectively a differentiable argmax. Ordinary classification networks use global average pooling, which erases location information, but robot manipulation specifically needs to know where things are; softmax also suppresses weak false activations, making it more robust to distractor objects. It later became a common choice in imitation-learning vision encoders — Diffusion Policy uses it at the end of its ResNet-18 in place of global average pooling.
ExampleLevine and colleagues' policy network follows three convolutional layers with a spatial softmax; the last layer's 32 channels each output one feature-point coordinate, which is concatenated with the robot's joint state and passed through fully connected layers to output motor torques, with the whole network having only about 92,000 parameters.
- Also called
- Spatial Soft-Argmax, Soft-Argmax, Keypoint Pooling
- Related
- Convolutional Neural Network · Vision Encoder · Visuomotor Policy · Diffusion Policy · Keypoint Detection · End-to-End Training of Deep Visuomotor Policies
- Sources
- End-to-End Training of Deep Visuomotor Policies (arXiv:1504.00702, JMLR 2016)
Deep Spatial Autoencoders for Visuomotor Learning (arXiv:1509.06113)
Diffusion Policy (arXiv:2303.04137, HTML)