Embodied AI Glossary中文

Distributed Data Parallel (DDP)

分布式数据并行DDPAdvanced

A multi-GPU training method where each GPU holds a full copy of the model and processes a slice of the data, then syncs gradients.

Distributed data parallel is the most common way to train across multiple GPUs, implemented in PyTorch as torch.nn.parallel.DistributedDataParallel. The approach: every GPU (every process) holds a complete copy of the model; a batch of data is split up and divided among the GPUs, each running its own forward and backward pass, then gradients are synchronized through all-reduce (a form of collective communication where every GPU's gradients are averaged together), after which each GPU updates its parameters, keeping every copy identical. It lets training speed scale close to linearly with the number of GPUs, but requires the full model to fit on a single GPU; when the model is too large for that, a parameter-sharding approach like FSDP or DeepSpeed's ZeRO is needed instead. It's typically launched with torchrun.

ExampleRunning torchrun --nproc_per_node=8 train.py trains a diffusion policy across 8 GPUs on one machine, with a per-GPU batch size of 32 and an effective total batch size of 256.

Also called
DistributedDataParallel
Related
Fully Sharded Data Parallel (FSDP) · DeepSpeed · Distributed Training · Batch Size · PyTorch · Gradient Accumulation
Sources
PyTorch DDP 文档 (Chinese)

See it in the full glossary →