Neural Network Quantization is the model compression technique that reduces the numerical precision of network weights and activations from 32-bit floating-point (FP32) to lower bit-widths (FP16, INT8, INT4, or even binary) — shrinking model size by 2-8x, reducing memory bandwidth requirements proportionally, and enabling execution on integer arithmetic units that are 2-4x more power-efficient than floating-point units, all while maintaining acceptable accuracy degradation.
Why Quantization Matters for LLMs
A 70B parameter model in FP16 requires 140 GB of GPU memory — exceeding single-GPU capacity. INT4 quantization reduces this to ~35 GB, fitting on a single 48 GB GPU. Since LLM inference is memory-bandwidth bound (loading weights dominates compute time), 4x smaller weights directly translates to ~4x faster token generation.
Quantization Approaches
- Post-Training Quantization (PTQ): Quantize a pretrained FP16 model without retraining. A small calibration dataset (128-512 samples) determines the quantization parameters (scale and zero-point). Fast (minutes to hours) but may lose accuracy at low bit-widths.
- Quantization-Aware Training (QAT): Insert fake quantization operators during training that simulate low-precision arithmetic while maintaining FP32 gradients. The model learns to be robust to quantization noise. Higher accuracy than PTQ at the same bit-width, but requires the full training pipeline.
LLM-Specific PTQ Methods
- GPTQ: Layer-wise quantization using optimal brain quantization (OBQ) with Hessian-based error correction. Quantizes weights to INT4/INT3 while compensating for quantization error by adjusting remaining weights. The standard for INT4 weight-only quantization.
- AWQ (Activation-Aware Weight Quantization): Identifies salient weight channels (those multiplied by large activation magnitudes) and scales them up before quantization, protecting important weights from quantization error. Simpler than GPTQ with comparable accuracy.
- SqueezeLLM: Sensitivity-based non-uniform quantization that allocates more bits to sensitive weight clusters and fewer to insensitive ones.
- QuIP/QuIP#: Uses random orthogonal transformations to decorrelate weights before quantization, enabling sub-4-bit precision with incoherence processing.
Quantization Formats
| Format | Bits | Memory Saving | Accuracy Impact | Hardware |
|---|---|---|---|---|
| FP16/BF16 | 16 | 2x vs FP32 | Negligible | All modern GPUs |
| INT8 | 8 | 4x vs FP32 | Minimal | GPU Tensor Cores, CPUs |
| INT4 (weight-only) | 4 | 8x vs FP32 | Small (~1-2% task degradation) | GPU with dequant kernels |
| NF4 (QLoRA) | 4 | 8x vs FP32 | Optimized for normal distribution | GPU software |
| INT2-3 | 2-3 | 10-16x vs FP32 | Moderate-significant | Research |
Neural Network Quantization is the practical engineering that makes large language models deployable on real hardware — converting academic-scale models into production-ready systems that serve millions of users at acceptable latency and cost.
Related Topics
Explore 500+ Semiconductor & AI Topics
From EUV lithography to CUDA optimization — search the full knowledge base or chat with our AI assistant.