Quantization-Aware Training
量化感知训练QATAdvancedSimulating low-bit error during training itself, so the model adapts in advance and loses less accuracy after quantization.
Quantization-aware training inserts 'fake quantization' operations during training or fine-tuning: the forward pass rounds weights and activations to low-bit values (such as INT8 or INT4) before using them in computation, while backpropagation treats the rounding as an identity function (a straight-through estimator) to work around its non-differentiability, letting the model learn to work under quantization error. Google's Jacob and colleagues' 2018 paper on integer-only inference is a representative example of this approach. Compared with post-training quantization, QAT needs training data and extra compute, but keeps accuracy noticeably better at 4 bits and below; a common recipe is to train normally first, then do a short QAT fine-tuning pass. In embodied AI, BitVLA uses a 'quantize then distill' form of quantization-aware training, guided by a full-precision teacher model, to compress its vision encoder to 1.58 bits, and the paper reports 11x less memory than OpenVLA-OFT at comparable performance.
ExampleGoogle's April 2025 QAT release of Gemma 3 runs about 5,000 steps of QAT before quantizing to int4, cutting the 27B model's memory from 54GB at BF16 down to 14.1GB, with Google reporting 54% less perplexity loss from quantization than quantizing directly.
- Also called
- QAT, Fake-Quantization Training
- Related
- Post-Training Quantization · Pruning · Knowledge Distillation · On-Device / Edge Deployment · Mixed-Precision Training · Numerical Precision Formats (FP32 / FP16 / BF16 / FP8 / INT8 / INT4)
- Sources
- Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference (arXiv 1712.05877)
Gemma 3 QAT Models: Bringing state-of-the-Art AI to consumer GPUs (Google Developers Blog)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation (arXiv 2506.07530) - As of
- 2025-06