11Models & Architectures
What the resulting models look like: Transformers, VLMs, how actions are generated, VLAs, and world models. · 184 terms
- 11.1Neural network basics16
- 11.2From attention to language models24
- 11.3Vision encoders and vision-language models28
- 11.4Introduction to generative models10
- 11.5Diffusion models and flow matching22
- 11.6How actions are represented, output17
- 11.7VLA structure and extensions17
- 11.8Hierarchical systems and embodied reasoning12
- 11.9World models and video generation23
- 11.10Inference speed, deployment, reliability15
11.1Neural network basics
Starting with the basic building blocks and classic architectures of neural networks, which every later model is built from.
- Neural Network神经网络A model made of many connected artificial neurons that learns by adjusting the strength of those connections.
- The total number of trainable numbers, or weights, in a model, usually written with B for billion or M for million.
- The most basic neural network: stacked fully-connected layers with nonlinear activations in between.
- The nonlinear function after each layer's linear transform, letting a network fit complex relationships.
- SoftmaxSoftmax(归一化指数函数)A function that turns a set of arbitrary real numbers into a probability distribution: all positive, summing to 1.
- Normalization Layers归一化层(层归一化 / RMSNorm / 批归一化)Layers that rescale a network's intermediate features back into a stable numeric range, making deep networks train faster and more stably.
- A neural network that slides small filters across an image to extract local features; long the workhorse of vision.
- Residual Network残差网络A convolutional network with skip connections that let each layer just learn the difference between input and output, enabling very deep networks.
- Backbone Network骨干网络The main network that extracts general-purpose features from raw input, with task-specific heads attached after it.
- Google's efficient convolutional-network family that scales depth, width, and resolution together by a single fixed ratio.
- Recurrent Neural Network循环神经网络A neural network that processes a sequence one step at a time, using a hidden state to remember the past.
- Long Short-Term Memory / Gated Recurrent Unit长短期记忆网络 / 门控循环单元Recurrent neural networks with learned 'gates' that let them retain longer histories, widely used for time-series data.
- Encoder-Decoder编码器-解码器A network design that compresses the input into an intermediate representation, then generates the output from that representation.
- Inductive Bias归纳偏置The built-in assumptions a model or algorithm relies on to decide how it should generalize to unseen data.
- A network whose output rotates or shifts the same way as its input, building symmetry directly into the architecture.
- Graph Neural Network图神经网络A neural network built for graph-structured data, where each node repeatedly exchanges information with its neighbors along edges.
11.2From attention to language models
The shared backbone of large models: tokens, attention, and the Transformer, leading up to large language models.
- Tokentoken(词元)The basic unit a model processes data in; text, image patches, and actions can all be broken into tokens.
- Tokenizer分词器The preprocessing step that splits text, images, or actions into tokens and converts them into integer IDs.
- Byte-Pair Encoding字节对编码A tokenization algorithm that builds a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair.
- Embedding嵌入向量Representing a word, image patch, or action as a string of real numbers, where similar things end up as nearby vectors.
- Attention Mechanism注意力机制A way of weighting parts of the input by relevance and combining them, letting a model focus on what matters.
- Self-Attention自注意力A mechanism where every element in a sequence gathers information from every other element, weighted by relevance.
- Cross-Attention交叉注意力Attention where one set of tokens queries information from a different set; a common way to fuse modalities.
- A neural network architecture built entirely around attention, the shared backbone of large language models and VLAs.
- Information added to each Transformer token telling the model which position in the sequence it's at.
- A positional encoding that writes position as a rotation angle applied to a vector, so attention naturally senses relative distance.
- Causal Attention因果注意力Attention where each position can only see itself and earlier positions, never anything that comes after.
- Attention Mask注意力掩码A matrix specifying which tokens each token in a sequence is allowed to 'see.'
- Foundation Model基础模型A large model pretrained on massive data that can later adapt to many different downstream tasks.
- Large Language Model大语言模型A very large neural network trained on massive text that can understand and generate natural language.
- Generating output one piece at a time, with each step conditioned on everything generated so far.
- Decoding Strategies解码策略(贪心 / 温度采样 / Top-k / Top-p)The rule for picking an actual token once the model has given a probability distribution over what comes next.
- A Transformer design that keeps only the decoder, predicting the next token one at a time with causal attention.
- Meta's family of openly released large language models, widely used as the language backbone inside other models.
- Mixture of Experts混合专家模型A model made of many 'expert' subnetworks where only a few are activated per input, trading a little compute for a lot of capacity.
- Prompt / Prompt Engineering提示词 / 提示工程A prompt is the text fed into a model; prompt engineering is designing it to get the output you want.
- Having a model write out a series of intermediate reasoning steps before giving its final answer.
- Context Length上下文长度The maximum number of tokens a model can process at once, which sets how much input it can 'see' at a time.
- State Space Model状态空间模型A network that processes long sequences with a hidden state updated over time, with compute growing linearly with sequence length.
- A sequence model that replaces attention with an input-dependent state space model, scaling linearly with sequence length.
11.3Vision encoders and vision-language models
Giving a language model eyes: first how images become tokens, then VLMs, the foundation VLAs are built on.
- Vision Encoder视觉编码器A network that turns an image into a set of feature vectors, or visual tokens, for downstream models to use.
- Vision Transformer视觉 TransformerA network that cuts an image into small patches, treats them like a sequence of words, and processes them with a Transformer.
- Visual Token视觉 tokenOne of the vectors produced by cutting an image into patches and encoding them — the basic unit a model uses for images.
- Vision Foundation Model视觉基础模型A general-purpose model pretrained on huge amounts of images, usable for many vision tasks with little or no adaptation.
- An OpenAI model trained on 400 million image-text pairs that maps pictures and text into one shared vector space.
- Google's image-text model trained with a sigmoid loss instead of softmax; its image encoder is used by many VLAs.
- Meta's 2023 self-supervised vision model that learns general-purpose image features with no labels at all.
- Meta's August 2025 third-generation DINO: 7 billion parameters, self-supervised on 1.7 billion images.
- Microsoft's small open-source vision foundation model that switches between captioning, detection, and segmentation via text prompts.
- A vision encoder pretrained on large-scale images or video, reused as the 'eyes' for a robot policy.
- Spatial Softmax空间 SoftmaxPooling each channel of a convolutional feature map down to the 2D coordinate of its 'brightest point,' preserving location.
- Text Encoder文本编码器A network that turns a text instruction into a sequence of vectors for the rest of the model to use.
- Multimodal Fusion多模态融合Combining information from different sources — image, language, robot state — into a form a model can use jointly.
- Vision-Language Model视觉语言模型A model that takes in both images and text and answers in text; the base that most VLAs build on.
- Multimodal Large Language Model多模态大语言模型A large language model that can also understand non-text inputs like images, video, and audio.
- A small network that maps features from a vision encoder into the input space of a language model.
- Learnable Query可学习查询A set of vectors trained as model parameters that use attention to gather task-relevant information out of the input.
- Querying TransformerQ-FormerThe small module in BLIP-2 that uses 32 learnable queries to distill image information before handing it to a large language model.
- Perceiver ResamplerPerceiver 重采样器A module that uses a small set of learnable query vectors to compress a variable number of visual features into a fixed number of tokens.
- A module that adaptively pools a large number of image tokens down into a handful of key tokens, to save compute.
- Google's open-weight vision-language model, combining a SigLIP visual encoder with a Gemma language model.
- Qwen-VL通义千问 Qwen-VLAlibaba's Qwen team's open-source vision-language model series, frequently used as a backbone for VLA models.
- Dynamic / Native Resolution动态分辨率(原生分辨率输入)Letting a vision encoder cut an image at its original size, so bigger images produce more visual tokens instead of being rescaled.
- Prismatic VLMsPrismatic VLMA Stanford/Toyota Research Institute study and open model that systematically compares VLM design choices; OpenVLA's base model.
- NVIDIA Eagle VLMEagle(英伟达 VLM)NVIDIA's open-source vision-language model series, used as the vision-language backbone in GR00T N1 through N1.6.
- Native Multimodal原生多模态Training a model on multiple modalities together from the very start of pretraining, instead of bolting a vision module onto a language model.
- Unified Multimodal Model统一多模态模型A single model that can both understand images and video by answering questions about them, and generate or edit images and video.
- ByteDance Seed's open-source unified multimodal model that can both understand images and generate or edit them.
11.4Introduction to generative models
Earlier models output a single answer; here, models learn a whole data distribution, plus turning vectors into discrete codes.
- Generative Model生成模型A model that learns the probability distribution behind data and can sample new data from it.
- Autoencoder自编码器A network that compresses input into a short vector and then reconstructs it, learning what matters most in the data.
- Latent Space潜在空间The internal representation space a model compresses raw data into, where similar things end up close together.
- Variational Autoencoder变分自编码器A model that compresses data into a distribution over latent variables, and can sample from it to reconstruct or generate data.
- A VAE that learns an output distribution given a condition, so the same input can generate several valid outputs.
- A generative model where a generator fakes data and a discriminator tries to catch the fakes, trained against each other.
- Normalizing Flow标准化流A generative model that turns a simple distribution into a complex one through a chain of invertible transforms, with exact probabilities computable.
- Preparing a 'codebook' and replacing a continuous vector with the ID of its nearest codeword, turning it into a discrete token.
- Vector-Quantized Variational Autoencoder向量量化变分自编码器An autoencoder whose encoded output is quantized by a codebook into discrete IDs, commonly used to turn images, video, or actions into tokens.
- Rounding each dimension of a continuous vector directly to a few fixed levels, replacing a VQ codebook for discretization.
11.5Diffusion models and flow matching
The leading approach among generative models: adding and removing noise, and flow matching, which power both action and video generation.
- Diffusion Model扩散模型A generative model that learns to remove noise step by step, then generates new data by denoising from pure noise.
- The 2020 diffusion model that made the approach work well: add noise step by step, then learn to remove it step by step.
- Learning a velocity field that smoothly transports noise into real data, then generating samples by integrating along it.
- The function a flow-matching model learns: it tells each sample which direction to move, and how fast, at every timestep.
- A generative model that moves noise toward data along a straight line — the straighter the path, the fewer sampling steps needed.
- A U-shaped convolutional network that downsamples layer by layer, then upsamples back, with matching layers connected directly.
- Diffusion Transformer扩散 TransformerUsing a Transformer instead of a U-Net as a diffusion model's denoising network architecture.
- Feature-wise Linear ModulationFiLM 特征调制Using conditioning information to compute a per-channel scale and shift that modulates a network's intermediate features.
- Adaptive Layer Normalization自适应层归一化Computing layer normalization's scale and shift dynamically from conditioning information, such as a diffusion timestep, instead of fixing them.
- Classifier-Free Guidance无分类器引导Computing both a conditional and an unconditional prediction and extrapolating between them so generation follows the condition more closely.
- Latent Diffusion Model潜在扩散模型A model that first compresses data into a low-dimensional latent space with an autoencoder, then runs diffusion generation there.
- Noise Schedule噪声调度The timetable specifying how much noise a diffusion model adds at each step, which affects training and generation quality.
- Prediction Target Parameterization预测目标参数化(ε / v / x₀ 预测)The choice of whether a diffusion model's network should output noise, velocity, or the clean data itself.
- Score Function / Score Matching分数函数 / 分数匹配The 'score' is the gradient of log-probability with respect to the input; score matching is how a network learns it.
- Denoising Steps去噪步数How many times a diffusion or flow-matching model calls its network to generate one result; this sets inference speed.
- A diffusion speedup method that keeps DDPM's training but turns sampling into a deterministic process that can skip steps.
- Diffusion / Flow Samplers扩散 / 流采样器(ODE / SDE 求解器)The numerical algorithm that integrates a trained diffusion or flow model from random noise into a sample, step by step.
- Consistency Model一致性模型A generative model that maps noise back to data in one step, used to compress diffusion's many sampling steps into one or two.
- Producing a result directly from noise with just one network forward pass, solving diffusion models' slow sampling.
- MeanFlow平均流A generative method that learns the 'average velocity' over a time interval, letting it produce a sample from noise in a single step.
- A diffusion model that adds and removes noise on discrete symbols, like text tokens, instead of continuous values.
- Diffusion Language Model扩散语言模型A language model that denoises a whole span of text in parallel using diffusion, instead of generating it word by word.
11.6How actions are represented, output
With networks and generative methods in hand, see how robot actions are represented, and output by regression, discretization, or generation.
- Action Head动作头The output module attached after a backbone that turns extracted features into concrete robot actions.
- The quantity, reference frame, and format used to describe what action a robot should take.
- Delta (Relative) Action vs. Absolute Action增量动作 / 绝对动作Writing an action as “move by this much” (delta) versus “go to this coordinate” (absolute).
- Having a network output continuous action values directly, trained with L1 or mean-squared error against demonstrations.
- Action Chunking动作分块A policy predicts a short sequence of upcoming actions at once, instead of outputting just one action per step.
- Action Horizon动作视界How many steps of history a policy looks at, how many steps of action it predicts, and how many it actually executes.
- Predicting an action chunk at every step, then averaging the overlapping chunks' predictions for the same moment before executing.
- Keyframe Action Prediction关键帧动作预测Predicting only a few key end-effector poses for a task, and leaving the path between them to a motion planner.
- Action Binning分箱离散化Dividing each continuous action dimension into equal bins and using the bin number as a discrete token.
- Action Tokenizer动作分词器A module that encodes continuous actions into discrete tokens and decodes them back into actions at inference time.
- A transform that breaks a signal into a weighted sum of cosine waves at different frequencies, widely used for compression.
- Gaussian Policy高斯策略A stochastic policy where the network outputs an action's mean and standard deviation, then samples from that normal distribution.
- Gaussian Mixture Model高斯混合模型A probability model that describes data as a weighted combination of several Gaussian distributions, which can have multiple peaks.
- Mixture Density Network混合密度网络A neural network that outputs the parameters of a Gaussian mixture distribution, instead of a single value directly.
- A model that scores input-output pairs instead of predicting the output directly; the lowest-scoring output is the answer.
- An output module attached after a policy's backbone that generates continuous actions through diffusion denoising.
- Residual Policy残差策略Learning a correction on top of an existing controller or policy's output, and adding the two together as the final action.
11.7VLA structure and extensions
Attach action output to a VLM and you get a VLA: how it outputs actions, generalizes across robots, and takes in more sensors.
- End-to-End端到端Using one model to go straight from raw sensor input to control commands, with no hand-designed intermediate modules.
- A large model pretrained on data spanning many robots and tasks, adaptable to many different robots and jobs.
- Large Behavior Model大行为模型Toyota Research Institute's term for a general-purpose, multi-task robot policy pretrained on large amounts of demonstration data.
- Vision-Language-Action Model视觉-语言-动作模型A large model that looks at images, understands language instructions, and directly outputs robot actions.
- Producing an entire sequence in one forward pass, instead of generating it token by token like autoregressive decoding.
- Action Expert动作专家A separate set of parameters inside a VLA dedicated to producing continuous actions, apart from the vision-and-language backbone.
- Mixture-of-Transformers混合 Transformer 架构Giving each modality its own set of Transformer parameters, while every layer's self-attention still lets them all see each other.
- An architecture where discrete content is generated autoregressively and continuous content is generated by diffusion, in the same model.
- Unified Action Space统一动作空间A single fixed format that maps different robots' actions by physical meaning, so their data can be trained together.
- In a cross-embodiment model, a separate output layer per robot type that turns shared features into that robot's actions.
- A small network that turns a robot's own state — joint angles, gripper opening — into a vector the model can use.
- History Encoder历史编码器A module that compresses a recent window of observations and actions into a vector, letting the policy infer current conditions.
- Point Cloud Encoder点云编码器A module that converts an unordered set of 3D points into feature vectors a neural network can use.
- A VLA that explicitly feeds spatial information — depth, point clouds, 3D position — into the model.
- A VLA that takes force/torque signals as input alongside vision and language.
- Vision-Tactile-Language-Action Model视觉-触觉-语言-动作模型A robot model that adds touch sensing to VLA, built for contact-heavy tasks that need a physical feel for the world.
- Memory-Augmented VLA记忆增强 VLAA VLA with an added history-memory module, so it can act based on what happened earlier, not just the current frame.
11.8Hierarchical systems and embodied reasoning
The opposite of end-to-end: splitting planning and control across different models, then having a model reason before it acts.
- Splitting robot decision-making into a high-level planner and a low-level executor, run by two different models.
- Embodied Reasoning Model具身推理模型A multimodal model specialized in understanding the physical world and planning tasks, often called a robot's ‘brain’.
- A slow, deliberate large model handles understanding and planning, while a fast, lightweight model handles real-time motor control.
- System 0System 0(三层系统架构)A kilohertz-rate whole-body control network added below the fast/slow two-system split, dedicated to balance, contact, and coordination.
- A large-pretrained humanoid whole-body control model that can adapt to many motion tasks zero-shot or with little tuning.
- Human Motion Generation人体动作生成模型(文本生成动作)A model that generates a sequence of 3D human skeletal poses from a condition such as a text description.
- Transitional information a model predicts between an instruction and low-level action, such as a 2D trajectory or keypoints.
- Visual Prompting (Set-of-Mark)视觉提示(Set-of-Mark 标记提示)Drawing numbers, boxes, or dots on an image so a vision-language model can point at a location just by naming its label.
- A method that has a robot model reason step by step — plan, subtask, object locations — before producing an action.
- Having a model first produce visual intermediate steps, like generated images or marked regions, before giving its final answer or action.
- Having a VLA first reason out a coarse trajectory in action space, then generate the fine-grained action from it.
- Latent Reasoning潜在推理Having a model carry out multi-step reasoning inside its internal hidden states, instead of writing out a chain of thought in words.
11.9World models and video generation
Using generative models to predict ‘what happens to the world after an action,’ finally merging with action generation itself.
- World Model世界模型A model that predicts what the world will look like after a given action is taken.
- Forward Dynamics Model正向动力学模型A model that takes the current state and an action and predicts what the next state will look like.
- Inverse Dynamics Model逆动力学模型A model that looks at two frames, before and after, and infers what action happened in between.
- Latent Action潜在动作An abstract action code learned automatically from how consecutive video frames change, with no real action labels needed.
- Latent Action Model潜在动作模型The network that produces latent actions from unlabeled video, trained by having an encoder-decoder pair reconstruct the next frame.
- Vision-Language-Latent-ActionViLLA 架构The hierarchical architecture behind AgiBot's GO-1: it first predicts latent action tokens, then decodes them into real robot actions.
- Video Prediction Model视频预测模型A model that predicts upcoming frames from recent ones, often given the action that's about to be taken.
- Video Generation Model视频生成模型A generative model that produces a new, coherent video from text, an image, or an existing clip.
- Text-to-Video / Image-to-Video文生视频 / 图生视频Generating a video from a text description, or from an image plus text.
- Video Tokenizer视频分词器A codec network that compresses video into a small number of tokens or latent vectors and can reconstruct the footage from them.
- Small cubes cut jointly across time and space from a video, each treated as one token for a Transformer.
- Alibaba's Tongyi-team video generation model; Wan2.1 and Wan2.2 are open-weight and widely used as a base for world models.
- World Foundation Model世界基础模型A general-purpose world model pretrained on huge amounts of real video that can be fine-tuned into specialized world models.
- Generating video forward in time, one frame or chunk at a time, so it can be played as it's produced.
- Exposure Bias曝光偏差(自回归误差累积)The problem where a model sees ground truth during training but its own output at inference, so errors compound.
- A causal diffusion sequence model trained by adding an independently random noise level to each token in the sequence.
- Interactive World Model可交互世界模型A world model that takes an action as input at every step and generates the next frame from it, usable as a neural simulator.
- Latent World Model隐空间世界模型A world model that predicts 'what happens if this action is taken' in a compressed latent state instead of pixels.
- Recurrent State-Space Model循环状态空间模型The core structure of the Dreamer world models, combining a deterministic recurrent state with a stochastic latent to predict the future.
- A self-supervised architecture that predicts masked or future content in an abstract representation space instead of reconstructing raw pixels.
- When a self-supervised encoder outputs the same or nearly the same vector for every input, so the representation carries no information.
- 4D World Model4D 世界模型A world model that predicts how a 3D scene changes over time — 3D space plus time.
- World Action Model世界动作模型A model that predicts future frames and robot actions together, using its prediction about the world to guide the action.
11.10Inference speed, deployment, reliability
Every model here eventually runs on hardware: how to make it faster and cheaper, and how to tell when it’s unsure.
- The time a model takes from receiving input to producing output, which determines how quickly a robot can react.
- The robot keeps executing the current action chunk while the model computes the next one, with no pause between.
- Real-Time Chunking实时动作分块A method that computes the next action chunk while still executing the current one, blended smoothly into what's already running.
- Floating-Point Operations (FLOPs)浮点运算量(FLOPs)The number of floating-point additions and multiplications a computation takes — a measure of how computationally heavy a model is.
- Key-Value CacheKV 缓存Storing already-computed attention keys and values so later tokens can reuse them instead of recomputing.
- Having a small model quickly guess a few tokens, then having the large model verify them all in one parallel pass, with no change in output.
- Converting a trained model's weights and activations directly to low-bit numbers after training, with no retraining needed.
- Simulating low-bit error during training itself, so the model adapts in advance and loses less accuracy after quantization.
- Pruning剪枝Deleting unimportant weights, channels, or layers from a model to make it smaller and faster.
- Visual Token Pruning视觉 token 剪枝Dropping or merging unimportant image tokens at inference time so vision-language and VLA models run faster.
- Early Exit早退机制Letting an easy input produce its output at a middle layer of the network, skipping the remaining layers' computation.
- On-device Model端侧模型A model that runs locally on a device's own chip — a robot, a phone — instead of relying on the cloud.
- Uncertainty Estimation不确定性估计Having a model report how confident it is alongside its prediction, so it's clear when to stop or ask for help.
- Gaussian Process高斯过程A method that puts a probability distribution directly over functions, giving a prediction with uncertainty at every point.
- Interpretability可解释性(机制可解释性)Studying how a neural network's internals produce its output; mechanistic interpretability breaks that computation into human-understandable pieces.