Quantization, Explained: How a Huge Model Fits on a Phone

If you have ever downloaded a model to run locally, you have run into filenames like llama-3-8b-Q4_K_M.gguf and wondered what the middle part means. It is the quantization level, and it is the single most important thing determining whether the model fits on your machine.

The underlying idea is simple, and once you see it the filenames stop being cryptic.

What a model actually is #

A language model is a very large pile of numbers called weights or parameters. An “8B” model has eight billion of them. These numbers were learned during training, and running the model means doing an enormous amount of arithmetic with them.

To run the model, those numbers have to be in memory. So the memory requirement is roughly the number of parameters times the bytes per parameter.

During training, each weight is typically stored as a 16-bit or 32-bit floating point number. At 16 bits, which is 2 bytes, an 8-billion-parameter model needs about 16 GB of memory just to hold the weights, before you account for anything else.

That is more RAM than most laptops have available and far more than any phone. Which is the problem.

The trick #

Quantization stores each weight using fewer bits.

Instead of 16 bits per weight, use 8, or 4, or in aggressive cases fewer. The model gets proportionally smaller:

PrecisionBytes per weight8B model size
FP16 (16-bit float)2~16 GB
8-bit1~8 GB
5-bit0.625~5 GB
4-bit0.5~4 GB
3-bit0.375~3 GB

An 8-billion-parameter model at 4-bit is about 4 GB, which fits comfortably on a laptop and is plausible on a high-end phone. That is a factor of four, and it is the difference between “needs a server” and “runs on your hardware.”

Why it does not destroy the model #

The intuitive objection is that throwing away three quarters of the precision should ruin everything. It mostly does not, for two reasons.

The weights do not need that much precision. A weight of 0.4172839 and a weight of 0.42 produce nearly the same result. Neural networks are trained with noise, dropout, and stochastic gradients, and they turn out to be robust to small perturbations in individual weights. The precision was there because that is how floating point math works, not because the model needed all of it.

Errors also partly cancel. Any single weight is rounded up or down, and across billions of weights and many operations, those rounding errors are somewhat random and partly average out rather than compounding in one direction.

This is an empirical finding rather than a theoretical guarantee. People tried it, measured the quality, and discovered the loss was small. The theory came afterward.

How it actually works, one level down #

The naive version would be to round every number to the nearest of 16 possible values, which for 4-bit is all you have. That would be bad, because weights in different parts of the model have very different ranges.

The real version works in blocks. Take a group of weights, say 32 of them. Find the minimum and maximum in that block. Then map that specific range onto the available values, and store a scale factor alongside the block so you can reconstruct approximately the original numbers.

The result is that each block gets its own scale, so a block of tiny weights and a block of large weights are each represented accurately within their own range. The scale factors add a small amount of overhead, which is why a “4-bit” model averages slightly more than 4 bits per weight in practice.

More sophisticated schemes go further. They use different precision for different layers, because some layers matter more than others. They use calibration data to figure out which weights are most sensitive. They keep certain critical tensors, like embeddings and attention outputs, at higher precision while quantizing the bulk aggressively.

That is what the letters in the filename are about.

Decoding the filenames #

The common naming convention comes from the GGUF format used by llama.cpp and the tools built on it. Q4_K_M breaks down as three parts. Q4 means roughly 4 bits per weight. K means the k-quant method, a more sophisticated scheme that varies precision across the model rather than applying one uniform level. M means medium within that family, and you will also see S for small and L for large, which trade size against quality.

So Q4_K_M is 4-bit k-quant, medium. Q5_K_S is 5-bit k-quant, small. Q8_0 is straightforward 8-bit. Q2_K is aggressive 2-bit and is usually where things start to visibly break.

You will also encounter I-quants (IQ4_XS, IQ3_M and similar), a newer family that squeezes more quality out of very low bit rates using importance information, at the cost of somewhat slower inference on some hardware.

What you actually give up #

Measured on benchmarks, roughly:

At 8-bit the model is essentially indistinguishable from full precision. If you have the memory, this is free. At 6-bit and 5-bit it is very close, with differences showing up only on careful measurement. At 4-bit there is small but real degradation that most people cannot detect in casual use, which is why it is the sweet spot and what most people run. At 3-bit it becomes noticeable, with more repetition, weaker reasoning, and more mistakes on tasks requiring precision. At 2-bit the loss is substantial, sometimes still usable for simple tasks and often not.

The degradation is not uniform across tasks. Casual conversation and summarization hold up well at low precision. Mathematical reasoning, code generation, and long multi-step tasks degrade faster, because they need the model to be right at each step rather than approximately right on average.

The rule that actually matters #

The practical guidance is the opposite of what people’s instincts suggest.

A larger model at lower precision usually beats a smaller model at higher precision, at the same file size.

A 13B model at Q4 is around 7 GB. A 7B model at Q8 is around 7 GB. The quantized 13B is generally better on most tasks.

So the selection procedure is:

  1. Figure out how much memory you can actually spare. Leave headroom for the operating system, your other applications, and the model’s context, which also consumes memory and grows with conversation length.
  2. Pick the largest model that fits in that budget at Q4_K_M.
  3. If you have room left over, move up in precision before you move up in model size.

The exception is when you need precision specifically. For code generation or math, going from Q4 to Q5 or Q6 on the same model is often worth more than a larger model at Q4.

Quantization on phones #

On a phone the constraint is much tighter, and quantization stops being an optimization and becomes the only reason any of this is possible.

A phone realistically has 3 to 6 GB available for a model after the OS takes its share. At 4-bit, that means models in the 1 to 8 billion parameter range depending on the device. At full precision, none of them would fit.

When I built Personal LLM, matching quantized model sizes to device memory tiers was most of the engineering work. A model that loads fine on a flagship gets killed by the OS on a three-year-old phone, and there is no graceful failure mode: the app just disappears. The full account of what that took is in how I built an offline AI chat app, and the broader context is in on-device AI explained.

Other things quantization touches #

Speed. Quantized models are usually faster, not just smaller, because inference on consumer hardware is typically limited by memory bandwidth rather than raw compute. Moving fewer bytes means more tokens per second, so you get the speedup and the memory savings together.

The KV cache. The context of your conversation is also stored in memory and also grows. Long conversations consume real memory on top of the model. Some tools let you quantize the cache too.

Training versus inference. Everything here is about inference. Training generally needs higher precision, which is why quantization is applied afterward rather than during.

The summary #

Quantization stores each model weight in fewer bits, cutting memory use by a factor of four at 4-bit with modest quality loss. It works because neural network weights do not need much precision, and it works in blocks with per-block scale factors so different parts of the model keep their own range. Pick Q4_K_M unless you have a reason not to, and prefer a bigger model at lower precision over a smaller model at higher precision.

It is the single technique most responsible for AI running on hardware you own rather than hardware you rent.

Next: running AI models locally for the practical setup, and what is a token if you want to understand the other number that governs everything.