Embodied AI Glossary中文

Residual Network

残差网络ResNetCommon

A convolutional network with skip connections that let each layer just learn the difference between input and output, enabling very deep networks.

Residual networks were introduced in 2015 by Kaiming He and colleagues at Microsoft Research Asia, and the paper won the CVPR 2016 Best Paper Award. Before this, adding more layers to a convolutional network actually made it harder to train and hurt accuracy. ResNet adds a skip connection inside each block that adds the input directly to the output, so the block only needs to learn the 'residual' between input and target, and gradients can also flow straight back along this shortcut. This design allowed networks as deep as 152 layers and won the ILSVRC 2015 image classification competition. Residual connections later became standard in nearly all deep networks — each Transformer layer has them too. In robot learning, small ResNets such as ResNet-18 are commonly used as image encoders; both ACT and Diffusion Policy use one to turn camera images into features.

ExampleACT uses a ResNet-18 to compress each 480×640 camera image into a 15×20×512 feature map, then flattens it into 300 tokens fed into the Transformer.

Also called
ResNet, ResNet-18, ResNet-50, Residual Connection, Skip Connection
Related
Convolutional Neural Network · Backbone Network · Vision Encoder · Vanishing / Exploding Gradients · Action Chunking with Transformers · Diffusion Policy
Sources
Deep Residual Learning for Image Recognition (arXiv 1512.03385)
Dive into Deep Learning: Residual Networks (ResNet)
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)

See it in the full glossary →