Quantization
Quantization is a model compression technique that reduces the precision of neural network weights and activations from high-precision floating-point numbers to lower-precision representations, significantly decreasing memory usage and computational requirements. This process converts 32-bit or 16-bit floating-point values to 8-bit integers or even lower precision formats while maintaining acceptable model performance through careful calibration and optimization strategies.
Quantization enables deployment of large language models on resource-constrained devices, reduces inference latency, and lowers energy consumption without requiring architectural changes. The technique encompasses various approaches including post-training quantization, quantization-aware training, and dynamic quantization that balance compression ratio with accuracy preservation. Advanced quantization methods include mixed-precision quantization, outlier-aware approaches such as LLM.int8() and SmoothQuant, and calibrated post-training methods such as GPTQ and AWQ that limit accuracy loss. Pruning is a separate compression technique — it removes parameters instead of reducing their precision — and the two are often applied together. Quantization serves as a critical enabler for edge AI deployment, mobile applications, and cost-effective large-scale inference scenarios.
Related terms
Related services: Agentic AI consulting.
Ready to put agentic AI to work?
Book a free 45-minute consultation. We'll map one real process worth automating with production-grade AI.