Embodied AI Glossary中文

Positional Encoding

位置编码PECommon

Information added to each Transformer token telling the model which position in the sequence it's at.

Self-attention inside a Transformer is order-blind on its own: shuffle the input tokens and the output just gets shuffled the same way. To give the model a sense of what came before what, the 2017 paper ‘Attention Is All You Need’ added a set of sine and cosine values at different frequencies to the input embeddings — this is positional encoding; the authors also tried learned position embeddings and found the results almost identical. Several variants have followed: Rotary Position Embedding (RoPE), proposed by Jianlin Su and colleagues in 2021, encodes position with a rotation matrix so that attention scores naturally reflect relative distance, and it has been adopted by mainstream large models such as Llama; Qwen2-VL further extends it to M-RoPE, which jointly encodes text order, an image's rows and columns, and a video's time dimension. Embodied models have to handle multiple frames, multiple camera views, and action sequences, so positional encoding determines whether the model can tell which frame came first and which patch belongs where.

ExampleViT splits a 224×224 image into 16×16-pixel patches, 196 in total, and adds a learned position embedding to each one so the model knows which patch is top-left and which is bottom-right.

Also called
PE, Position Embedding
Related
Rotary Position Embedding · Transformer · Self-Attention · Vision Transformer · Context Length · Token
Sources
Attention Is All You Need (arXiv:1706.03762)
RoFormer: Enhanced Transformer with Rotary Position Embedding (arXiv:2104.09864)
Qwen2-VL (arXiv:2409.12191)

See it in the full glossary →