DeepSpeed
AdvancedMicrosoft's open-source library for accelerating large-model distributed training, best known for its ZeRO memory optimizer.
DeepSpeed is an open-source deep-learning training and inference optimization library from Microsoft, built on top of PyTorch. It's best known for ZeRO (Zero Redundancy Optimizer): instead of every GPU storing a full copy of the optimizer state, gradients, and even the model parameters, ZeRO shards them across multiple GPUs, making it possible to train larger models within limited memory. It also offers mixed precision, gradient accumulation, CPU/NVMe offloading, and pipeline parallelism. When training a VLA or large language model with parameters in the billions, a single GPU often can't hold it all, and DeepSpeed — or PyTorch's own FSDP, which follows a similar idea — is a common solution. Frameworks such as Hugging Face Accelerate can call into it directly.
ExampleWhen fine-tuning a 7B-parameter VLA, turning on DeepSpeed's ZeRO-2 or ZeRO-3 configuration in the training script lets 8 GPUs split the optimizer state and gradients between them.
- Related
- Fully Sharded Data Parallel (FSDP) · Distributed Data Parallel (DDP) · Mixed-Precision Training · Distributed Training · Hugging Face Accelerate · PyTorch
- Sources
- DeepSpeed 官网 (Chinese)
DeepSpeed GitHub