Embodied AI Glossary中文

Attention Mask

注意力掩码Advanced

A matrix specifying which tokens each token in a sequence is allowed to 'see.'

By default, self-attention in a Transformer lets every token attend to every other token in the sequence. An attention mask blocks out disallowed positions when computing attention scores — typically by adding negative infinity, which becomes a weight of zero after softmax — controlling how information can flow. Two kinds are most common: a causal mask, where each token can only see itself and earlier tokens, used in GPT-style models that generate one token at a time; and a padding mask, which blocks out empty positions added just to pad the sequence to a fixed length. VLA models commonly use a block-wise causal mask: the input is split into blocks, tokens within a block can all see each other, but a block can only see blocks before it. This both protects the input distribution the pretrained VLM originally saw and lets earlier blocks' KV cache be reused across multiple denoising steps, saving inference time.

Exampleπ0 splits its sequence into three blocks — 'image + language,' 'robot state,' and 'noisy action' — where the first block can't see the inputs added after it, to limit disruption to PaliGemma's pretrained distribution, and the state block can't see the action block, so its KV can be cached during sampling.

Also called
Block-wise Causal Mask, Blockwise Causal Attention Mask
Related
Attention Mechanism · Causal Attention · Self-Attention · Key-Value Cache · Action Expert · π0
Sources
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv 2410.24164)

See it in the full glossary →