Convolutional Neural Network
卷积神经网络CNNCommonA neural network that slides small filters across an image to extract local features; long the workhorse of vision.
A convolutional neural network is designed specifically for grid-shaped data such as images. Its core is the convolutional layer: a set of small, learnable filters, such as 3×3, slide across the image, computing a local weighted sum at each position; the same filter's parameters are shared across every position, so the parameter count is far lower than a fully connected network, and it's naturally suited to recognizing edges, textures, and parts wherever they appear in the frame. Pooling layers then progressively shrink the resolution, letting the network build up from local detail to overall meaning. Historically, LeCun used LeNet to recognize handwritten digits in the 1990s, AlexNet won decisively on ImageNet in 2012, and ResNet used residual connections to go much deeper in 2015. Since Vision Transformers appeared, large models' vision encoders have mostly switched to Transformers, but CNNs remain common in smaller robot models — RT-1, for instance, uses EfficientNet, and Diffusion Policy uses ResNet-18 to process camera images.
ExampleDiffusion Policy replaces a standard ResNet-18's global average pooling with a spatial softmax to preserve positional information, and replaces BatchNorm with GroupNorm to stabilize training, using it to encode each camera frame into features.
- Also called
- CNN, ConvNet
- Related
- Residual Network · EfficientNet · Vision Encoder · Vision Transformer · Backbone Network · Spatial Softmax
- Sources
- CS231n: Convolutional Neural Networks for Visual Recognition (Stanford)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)
RT-1: Robotics Transformer (project page)