we_are_coded.by CODE · The world, decoded
БГ
Concept

Quantization: how a model shrinks to fit

The BasicsUpdated on 25 August 2026we are coded

How a model made of a pile of numbers shrinks to fit a smaller machine and run faster, and what exactly you lose when you do it. Worth knowing, because it decides whether a local model on your computer is practical at all.

Checked on25 August 2026
In short: inside, a model is a pile of numbers, billions of them. Quantization records them more coarsely, with fewer digits per number, so they take up less memory and compute faster. You gain space and speed. You lose a bit of accuracy, sometimes in places you don't expect.

Same picture, fewer colors

Picture a photo saved with millions of shades of color, and then the same photo saved with only a few hundred. You still recognize everything in it, the face, the background, the dog in the corner, but the detail is coarser and the edges between shadows are sharper. That's exactly what quantizing a model does. The numbers it computes with are usually stored at 16 bits of precision, a measure of how finely a number is recorded. Quantization brings that down to 8 bits, or even 4. The same knowledge, recorded more coarsely.

The gain is direct. A model at 16 bits weighs twice as much as the same model at 8 bits, and four times as much as the 4-bit version. Less weight means less memory used, and less memory used means faster work, because the machine doesn't have to dig through as much data for every word of the answer.

The loss is trickier. Accuracy drops only a little on average, but sometimes right at the edge, on a rare word or a hard calculation, the gap shows up. That's why shrunk models get tested, not taken on faith.

Quantization doesn't make the model dumber. It makes it coarser on the detail you rarely need.

The number behind the claim

JetBrains shrank to 4 bits exactly the model that runs their local Junie, Qwen 3.6-27B. By their numbers, generation came out about twice as fast compared to the 8-bit version, on the same machine. That's why the shrunk version became the one they actually shipped: on a local machine, speed decides whether you use it at all.

The visual is generated code art. No third-party images.
Official primary sources
→JetBrains: Qwen for Junie