Distributed Training
分布式训练(数据并行 / 模型并行)CommonSplitting a single training run across multiple GPUs or machines so larger models and more data can be trained.
When a model or dataset is too large for one GPU to hold or process, training gets split across multiple GPUs or machines. There are two main approaches. Data parallelism puts a full copy of the model on each GPU, has each one process a different slice of the data, and after backpropagation averages the gradients across GPUs before applying a synchronized update — PyTorch's DDP (Distributed Data Parallel) works this way. Model parallelism instead splits the model itself across GPUs; NVIDIA's Megatron-LM, for instance, splits each Transformer layer's matrices across multiple GPUs (tensor parallelism), and there's also pipeline parallelism, which assigns different layers to different GPUs. FSDP (Fully Sharded Data Parallel) still splits by data, but also shards the parameters, gradients, and optimizer state across GPUs, substantially cutting memory use per device. Pretraining and full fine-tuning of multi-billion-parameter VLA models depend on these techniques.
ExampleOpenVLA (7B parameters) was pretrained on 64 A100 GPUs over 14 days; the paper's full fine-tuning also needs 8 A100s running 5–15 hours, and its memory-usage tests split the model across 2 GPUs with FSDP.
- Also called
- Data Parallelism, Model Parallelism, Tensor Parallelism, Pipeline Parallelism, Multi-GPU Training
- Related
- Distributed Data Parallel (DDP) · Fully Sharded Data Parallel (FSDP) · DeepSpeed · Mixed-Precision Training · Gradient Accumulation · Batch Size
- Sources
- PyTorch Tutorial: Getting Started with Distributed Data Parallel
PyTorch Tutorial: Getting Started with Fully Sharded Data Parallel (FSDP)
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism