Date Notice: Today is September 25, 2026. This article reflects knowledge verified through January 2026. Specific benchmark numbers, chip specs, and framework version details should be checked against current vendor documentation before you cite them anywhere serious.
A Pixel phone running a 7B parameter model on-device sounds impossible until you check the math twice. Full precision, that model needs 28GB of memory just to hold the weights — more RAM than the phone has, period. Ship the same model at 4-bit precision and it drops to roughly 3.5GB. Suddenly it fits, with room left for the KV cache and the rest of the OS. Nothing about the model’s architecture changed. The numbers representing its weights got smaller. That’s quantization, and it’s the reason on-device AI stopped being a demo and started being a product category.
Two Different Knives
Quantization and pruning both shrink a model, but they cut in different places, and conflating them causes real design mistakes.
Quantization reduces the numerical precision used to store and compute weights and activations — typically FP32 or BF16 down to INT8, INT4, or even lower. A weight of 0.7328491 stored in FP32 takes 32 bits. Quantized to INT8 with a scale factor, it might round to a value representable in 8 bits, reconstructed at inference as scale × int_value. You lose precision, not parameters. The model keeps every neuron, every connection — they just get expressed more coarsely.
Pruning removes parameters entirely. Unstructured pruning zeroes out individual weights below some magnitude threshold, producing a sparse matrix. Structured pruning removes whole channels, attention heads, or layers, which is the version that actually speeds things up on standard hardware, because unstructured sparsity needs specialized kernels (NVIDIA’s 2:4 sparsity on Ampere+ GPUs being the main mainstream exception) to translate into real latency wins. Without that hardware support, an unstructured-pruned model still runs the same dense matmuls, just with a lot of the multiplications wasted on zeros.
The two compose. A common pipeline: prune a model structurally to cut parameter count 30-40%, fine-tune to recover accuracy, then quantize the survivors to INT8 or INT4 for the final size and latency win. Doing it in the other order — quantize then prune — usually produces worse results, because pruning decisions made on a lower-precision model are less reliable; small weight magnitudes get harder to distinguish from quantization noise.


