Dropout
随机失活CommonRandomly “turning off” some neurons during training, a regularization technique that keeps a network from memorizing the training data.
Dropout comes from Hinton's group; Srivastava, Hinton, and colleagues published the systematic paper in JMLR in 2014. At every training step, each neuron's output is randomly zeroed out with some set probability, which is equivalent to training a different “sub-network” each time; at test time dropout is turned off, the full network is used, and the weights are rescaled accordingly to approximate the average over all those sub-networks. This keeps neurons from relying too heavily on each other, which reduces overfitting — doing well on training data but poorly on new data. It's commonly combined with weight decay, data augmentation, and early stopping; a dropout probability of 0.1 is common in Transformers. Robot demonstration datasets often have only dozens to a few hundred examples and overfit easily, so dropout is a common setting there.
ExampleThe official ACT (Action Chunking with Transformers) code uses a default Transformer dropout of 0.1. The reinforcement-learning algorithm DroQ adds dropout and layer normalization inside its Q-network, matching the sample efficiency of the larger ensemble method REDQ while using computation close to plain SAC.
- Also called
- Dropout Regularization
- Related
- Regularization · Overfitting · Early Stopping · Data Augmentation · Normalization Layers
- Sources
- Dropout: A Simple Way to Prevent Neural Networks from Overfitting (JMLR 2014)
ACT 官方代码 detr/main.py(--dropout 默认 0.1) (Chinese)
Dropout Q-Functions for Doubly Efficient Reinforcement Learning (DroQ)