How Quantization Shrinks AI Models
In the previous article we saw that a model is mostly a huge pile of weights. Every weight is a number, and every number takes up memory. Many models today store each weight with 16 bits, which is 2 bytes. So a model with 7 billion weights needs about 14 GB of memory just to hold them.
Quantization is the trick of storing each weight with fewer bits: 8, 4 or even fewer. Fewer bits per weight means a smaller model. It is often faster too, because writing text is often limited by how fast the weights can be read from memory. Pick a model size and a number of bits.
- 8 GBMany gaming graphics cardsToo big
- 24 GBHigh-end gaming graphics cardsFits
- 80 GBLarge data-center GPUsFits
GB here means a billion bytes. Running a model needs extra memory on top of the weights, so "fits" is only about the weights. The memory sizes are typical examples.
Going from 16 to 4 bits makes the weights four times smaller. That can be the difference between needing a large data-center GPU and running the model on a single graphics card, or even on a laptop.
But there is a catch. With fewer bits, there are fewer different numbers you can write down. With 4 bits there are only 16. So every weight has to be rounded to one of a few allowed steps, a bit like rounding prices to whole dollars. The model then stores only the number of the step. Pick the number of bits and watch the weights snap to the nearest step: the rounded weights.
- Allowed values
- 16
- Average rounding error per weight
- 0.029
- Neuron output, original weights
- 0.246
- Neuron output, rounded weights
- 0.237
Example weights and inputs, chosen by hand. The steps are spread evenly from the smallest to the largest weight. Real formats vary; some use small floating-point numbers instead of evenly spaced steps.
With 8 bits the steps are so close together that you can hardly see the rounding. With 4 bits you can, but each weight still moves only a little. With 2 bits there are only four allowed values, and the weights get pushed around a lot. Notice the neuron output too: some weights round up and others round down, so their errors can partly cancel out.
This is why 8-bit models usually behave almost exactly like the originals, and 4-bit versions of large models often work surprisingly well. Below that, the quality usually drops more noticeably. How much depends on the model and on how carefully it was quantized.
One thing makes rounding much worse: a few unusually large weights, which real models often have. The steps are spread evenly between the smallest and the largest weight, so a single big weight stretches the range, and all the small weights have to share just a few steps. Add a big weight and watch the others.
- Group 1
- 0.007
- Group 2
- 0.009
Example weights, rounded to 4 bits (16 steps from the smallest to the largest value in the range). Real tools often use groups of 32 to 128 weights; we use 8 so you can see them.
The common fix is to split the weights into small groups and give every group its own range. A big weight then only spoils its own group. Each group has to store its own range too, which costs a little extra memory. That is one reason a "4-bit" model is usually a bit bigger than exactly 4 bits per weight.
Real tools add more tricks. Some keep the most sensitive parts of the model at higher precision. Some choose the rounding carefully, using sample text to see which errors matter most. But the core idea stays the same: fewer bits, a few allowed values and careful rounding.
So quantization is a trade: a much smaller and often faster model, in exchange for a little precision. Most of the time the trade is worth it, which is why many of the models people run on their own computers are quantized.
One thing quantization does not do is make a model smarter. It changes how the weights are stored, not what the model has learned, so the goal is only to stay as close to the original as possible. When a model is quantized too hard, it can start making mistakes the original would not make.
If this content helped you, you can buy me a coffee.
You can join the newsletter to be notified of awesome interactive articles and courses about AI, software and design. You will receive at most a few emails per month.
Join 800+ curious readers.