Inductive Bias
归纳偏置AdvancedThe built-in assumptions a model or algorithm relies on to decide how it should generalize to unseen data.
Any finite batch of training data can be explained by many different underlying rules; inductive bias is the set of assumptions a learning algorithm relies on to choose among them. Tom Mitchell pointed out in 1980 that without such assumptions, a model has no way to make predictions about inputs it hasn't seen. A common example: convolutional neural networks assume that nearby pixels are related and that features are translation-equivariant (shift an object, and its feature map shifts with it). The ViT paper notes that a Transformer has much less built-in image inductive bias than a CNN, so it scores slightly lower than a similarly sized ResNet when trained on medium-scale data like ImageNet, and needs pretraining on much larger datasets to catch up or surpass it. A strong inductive bias saves data but may cap performance; a weak one relies more on data scale. In robot learning, voxelizing observations into 3D and equivariant policies that build rotation symmetry into the network both trade a geometric inductive bias for better performance with fewer samples.
ExamplePerAct voxelizes RGB-D observations before predicting actions, and the paper reports it outperforms a baseline that predicts actions directly from 2D images by 34x, which the authors attribute to the structural prior that 3D voxels provide.
- Also called
- Learning Bias, Structural Prior
- Related
- Convolutional Neural Network · Vision Transformer · Equivariant Policy / Equivariant Neural Network · Graph Neural Network · Generalization · The Bitter Lesson
- Sources
- Inductive bias(Wikipedia)
An Image is Worth 16x16 Words (ViT, arXiv:2010.11929)
Relational inductive biases, deep learning, and graph networks (arXiv:1806.01261)