TokenLearner
AdvancedA module that adaptively pools a large number of image tokens down into a handful of key tokens, to save compute.
TokenLearner is a module Google's Ryoo and colleagues proposed in 2021 (NeurIPS 2021), with a paper title that literally asks what 8 learned tokens can do. A Vision Transformer usually cuts an image into tens to hundreds of patches, one token each, and attention's compute grows with the square of the token count. TokenLearner instead computes a spatial attention map for each output token based on the input content, uses it to weight and pool the feature map, and compresses a large number of tokens down to around 8; later layers only process these few tokens. The paper reaches competitive results on benchmarks like ImageNet and the Kinetics video-recognition set while using noticeably less compute. Its best-known use in embodied AI is in Google's RT-1, where it helps a large model meet the speed requirements of real-time control.
ExampleRT-1 uses TokenLearner to compress each image's 81 visual tokens down to 8, so 6 frames of history become just 48 tokens fed into the Transformer; the paper reports this step gives about a 2.4x inference speedup.
- Related
- Visual Token · Visual Token Pruning · RT-1 · Perceiver Resampler · Querying Transformer · Vision Transformer
- Sources
- TokenLearner: What Can 8 Learned Tokens Do for Images and Videos? (arXiv:2106.11297)
RT-1: Robotics Transformer for Real-World Control at Scale (arXiv:2212.06817)