Embodied AI Glossary中文

Transformer

Essential

A neural network architecture built entirely around attention, the shared backbone of large language models and VLAs.

The Transformer was introduced by a Google team in the 2017 paper “Attention Is All You Need,” originally for machine translation. It does away with recurrence and convolution entirely, processing sequences purely through attention, which lets every token in a sequence directly weigh and pull in information from every other token by relevance. Compared with a recurrent network, which processes a sequence step by step, it can be computed in parallel, trains faster, and scales up far more easily, which made it the shared architecture behind large language models such as GPT and Llama, and vision models such as ViT. The original design is an encoder-decoder; GPT-style models use only the decoder half. Robotics started adopting it heavily around 2022, with Google's RT-1 (Robotics Transformer) a notable early example; today's VLAs, world models, and many diffusion policies are all built on a Transformer backbone.

ExampleOpenVLA's backbone, Llama 2, is a decoder-only Transformer: image tokens and instruction tokens are arranged into one sequence as input, and the model outputs, one token at a time, the tokens that represent the action.

Also called
Transformer Architecture
Related
Attention Mechanism · Self-Attention · Decoder-only Architecture · Encoder-Decoder · Vision Transformer · RT-1
Sources
Attention Is All You Need (Vaswani et al., arXiv:1706.03762)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)

See it in the full glossary →