Quantization explained: Q4 vs Q8 and what you lose

By Tried AI Tools · Published

Short answer

Quantization stores a model's weights with fewer bits so it uses less memory and runs faster. 8-bit is nearly indistinguishable from the full model, 4-bit (Q4_K_M or MLX 4-bit) is the usual sweet spot with a small quality loss, and below 4-bit quality drops noticeably. A bigger model at 4-bit usually beats a smaller model at 8-bit in the same memory.

Language models are made of billions of numbers called weights. Full-size models store each weight in 16 bits. Quantization rounds those weights to fewer bits, for example 8 or 4, which shrinks the file and the memory it needs.

Why it matters on a Mac

Memory is the limit on what you can run, and speed is limited by how many bytes are read per word generated. A 4-bit model takes about a quarter of the memory of the 16-bit original and runs up to several times faster.

Common formats

Name you will see Bits per weight (approx.) Size vs full model Quality
F16 / BF16 16 100% Original
Q8_0, MLX 8-bit 8.5 ~53% Practically identical
Q6_K 6.6 ~41% Very close
Q5_K_M 5.7 ~36% Very close
Q4_K_M, MLX 4-bit 4.5 to 5 ~30% Small loss, the usual default
Q3_K_M 3.9 ~24% Noticeable loss
Q2_K, IQ2 2.5 to 3 ~17% Large loss, last resort

GGUF names (Q4_K_M and so on) are used by llama.cpp, Ollama and LM Studio. MLX models are usually labelled simply 4-bit, 6-bit or 8-bit.

Bigger model, fewer bits

With a fixed amount of memory, you usually get better answers from a larger model at 4-bit than a smaller one at 8-bit. A 14B model at 4-bit (about 8.5 GB) generally outperforms an 8B model at 8-bit (about 8.5 GB). This stops being true at very low bit counts, where quality falls off quickly.

What to pick

For the memory each choice needs, see how much memory you need to run an LLM.

Frequently asked questions

What does Q4_K_M mean?

It is a llama.cpp (GGUF) quantization type. Q4 means about 4 bits per weight, K means the k-quant method that groups weights in blocks with shared scales, and M means the medium variant, which keeps some sensitive layers at higher precision.

Which quantization should I download?

Start with Q4_K_M for GGUF or 4-bit for MLX. If you have memory to spare, try Q5_K_M, Q6_K or 8-bit. Only drop to 3-bit or lower when it is the only way to fit a much larger model.