Embodied AI Glossary中文

Mixed-Precision Training

混合精度训练Advanced

Running most computation in 16-bit floating point and keeping only the sensitive parts in 32-bit, to save memory and gain speed.

Deep learning defaults to 32-bit floating point (FP32) for storing parameters and doing computation. Mixed-precision training switches operations like matrix multiplication and convolution to a 16-bit format (FP16 or BF16), while reductions, normalization, and other numerically sensitive operations, plus the master copy of the parameters, stay in FP32. Researchers at NVIDIA and Baidu systematically proposed this approach in 2017 and introduced “loss scaling” to stop small FP16 gradients from underflowing to zero; memory use drops by nearly half, and GPU tensor cores have higher throughput for 16-bit math, so training runs faster too. BF16 has the same numeric range as FP32 and usually needs no loss scaling, so most large-model and VLA training today uses BF16. PyTorch's autocast handles most of the precision switching automatically.

Exampleopenpi trains the π0 family in JAX with mixed precision by default: weights and gradients are kept in FP32, while most activations and computation run in BF16; training in full FP32 is just a matter of changing the config's dtype to float32.

Also called
bf16 Training, Half-Precision Training, Automatic Mixed Precision (AMP)
Related
Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4) · GPU Memory (VRAM) · Distributed Training · Gradient Checkpointing (Activation Recomputation) · Quantization-Aware Training · openpi (Physical Intelligence)
Sources
Micikevicius et al. 2017: Mixed Precision Training
NVIDIA Docs: Train With Mixed Precision
GitHub: Physical-Intelligence/openpi (Precision Settings)

See it in the full glossary →