Embodied AI Glossary中文

Fully Sharded Data Parallel (FSDP)

全分片数据并行FSDPAdvanced

A data-parallel training method that splits a model's parameters, gradients, and optimizer state across multiple GPUs.

FSDP is PyTorch's built-in distributed training approach, based on the idea behind Microsoft DeepSpeed's ZeRO-3. In ordinary distributed data parallel (DDP), every GPU stores a full copy of the model parameters, gradients, and optimizer state, so once a model gets large enough it simply doesn't fit on one GPU. FSDP shards all three across the available GPUs; before each layer's forward and backward pass, the full parameters for that layer are temporarily gathered from the other GPUs and released again right after, so each GPU only needs to hold a fraction of the total at any moment, making it possible to train models far larger than a single GPU's capacity — at the cost of more inter-GPU communication. Full-parameter fine-tuning of a multi-billion-parameter VLA or video world model commonly uses FSDP or DeepSpeed; newer PyTorch versions recommend the per-layer FSDP2 (fully_shard) interface.

ExampleFull-parameter fine-tuning a roughly 3-billion-parameter VLA on 8 GPUs with 80GB each runs out of memory under DDP, but switching to FSDP makes it fit.

Also called
FSDP2, fully_shard
Related
Distributed Data Parallel (DDP) · DeepSpeed · Distributed Training · PyTorch · Full Fine-Tuning · Mixed-Precision Training
Sources
PyTorch FSDP 文档 (Chinese)
PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

See it in the full glossary →