Embodied AI Glossary中文

Early Exit

早退机制Advanced

Letting an easy input produce its output at a middle layer of the network, skipping the remaining layers' computation.

A deep network normally runs every input through all of its layers, but many easy examples can already be judged correctly from shallow features. Early exit inserts several 'exits' (small classification or output heads) partway through the network; at inference, as soon as one exit is confident enough, the model outputs there and stops, and only hard examples run through the whole network. BranchyNet, from Harvard's Teerapittayanon and colleagues, is an early representative of this idea, later adopted for BERT and large language models. Robot control fits this pattern well, since the action is simple most of the time and only occasionally needs complex reasoning; Tsinghua's Gao Huang and colleagues proposed DeeR-VLA at NeurIPS 2024, turning a multimodal large model into a multi-exit structure that decides when to exit based on a compute, latency, and memory budget.

ExampleDeeR-VLA reports on the CALVIN benchmark that it cuts the language model's compute by 5.2-6.5x and GPU memory by 2-6x, with essentially no drop in task performance.

Also called
Multi-exit Network, Dynamic Inference
Related
Inference Latency · On-device Model · Visual Token Pruning · Pruning · Speculative Decoding · Mixture of Experts
Sources
BranchyNet: Fast Inference via Early Exiting from Deep Neural Networks (arXiv:1709.01686)
DeeR-VLA: Dynamic Inference of Multimodal Large Language Models for Efficient Robot Execution (arXiv:2411.02359)

See it in the full glossary →