Embodied AI Glossary中文

EfficientNet

Advanced

Google's efficient convolutional-network family that scales depth, width, and resolution together by a single fixed ratio.

EfficientNet is a family of convolutional neural networks proposed by Google's Mingxing Tan and Quoc V. Le at ICML 2019. Scaling up a CNN previously usually meant increasing just one of depth, width, or input resolution, and returns quickly leveled off. The paper introduced 'compound scaling': a single coefficient scales depth, width, and resolution together by a fixed ratio; a small baseline network, B0, is found with neural architecture search, and then scaled up step by step to get B1 through B7. The paper reports B7 reaching 84.3% top-1 accuracy on ImageNet while being 8.4x smaller and 6.1x faster at inference than the best convolutional networks at the time. Because it's small and fast, robot policies commonly used it as their vision backbone before ViT became widespread.

ExampleGoogle's RT-1 uses an ImageNet-pretrained EfficientNet to extract image features, has the language instruction modulate them through FiLM layers, and then passes the result to TokenLearner and a Transformer to output discrete action tokens.

Also called
EfficientNet-B0-B7
Related
Convolutional Neural Network · Backbone Network · Residual Network · RT-1 · Feature-wise Linear Modulation · TokenLearner
Sources
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks (arXiv:1905.11946)
RT-1: Robotics Transformer project page

See it in the full glossary →