Batch Size
批大小CommonThe number of samples fed through the model together for each single parameter update.
Training doesn't use the whole dataset at once; instead the data is split into small mini-batches, and each one produces one average loss, one gradient, and one parameter update. The number of samples in each of those batches is the batch size. A larger batch size gives a more stable gradient estimate and better GPU utilization, but uses more memory; too small a batch makes the gradient noisier and training slower. Batch size is closely tied to the learning rate, so changing one usually means retuning the other. When memory is too limited for a large batch, two common workarounds are gradient accumulation (summing gradients over several small batches before updating) and multi-GPU data parallelism. It's worth distinguishing batch size from a training epoch, which is one full pass through the entire dataset and consists of many batches.
ExampleOpenVLA trained for 14 days on 64 A100 GPUs, using a global batch size of 2048 and a fixed learning rate of 2e-5, passing over the training set 27 times in total.
- Also called
- Mini-batch Size
- Related
- Learning Rate · Epoch · Gradient Accumulation · Distributed Training · Hyperparameter · Gradient Descent
- Sources
- Google Machine Learning Glossary
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246) - As of
- 2024-06