XavierFok
← all posts

Quantization levels explained: what Q4, Q5, and Q8 actually trade away

2026-08-15 · by Xavier Fok

# Quantization levels explained: what Q4, Q5, and Q8 actually trade away

Every guide to running AI at home tells you to grab the quantized version of the model. Almost none of them explain what that word means or what you give up for it. So you end up staring at a download page full of files labelled Q4 and Q5 and Q8, picking one on vibes, with no idea whether you just made a sensible choice or quietly wrecked the model's quality.

I run quantized models every day on a seven year old gaming card with eleven gigabytes of memory, so this trade is one I live inside. Here is what the labels mean, what each level costs, and how I pick.

The idea in one paragraph

A language model is a giant pile of numbers called weights. Normally each weight is stored at high precision, sixteen bits per number, sometimes thirty two. All that precision eats memory. Quantization squeezes each number down to fewer bits: eight, five, four. You give up some exactness in every individual weight, and in exchange the model takes far less memory and moves faster. That is the entire concept. Everything after this is detail about how hard you squeeze.

Why the squeeze decides what you can run

The binding limit on a home GPU is VRAM, the memory on the card itself. Mine has eleven gigabytes. A seven billion parameter model at full precision wants around fourteen gigabytes for the weights alone, so it does not fit on my card at all. The same model quantized to four bits per weight shrinks to roughly four gigabytes, which fits with room left over for the actual work the model has to do. Same model, same card, opposite outcome.

That framing matters. Quantization gets described as an optimization, and for home hardware it is closer to a gate. It moves a model from impossible on your card to comfortable on your card.

Reading the label

The number after the Q is roughly bits per weight. Q4 means about four bits, Q5 about five, Q8 about eight. Higher number: more of the original precision kept, bigger file, more VRAM needed. Lower number: smaller file, less VRAM, more of the original information thrown away.

The letters that follow, things like K with an S, M, or L, describe the method used to do the squeezing. The K variants are the ones worth knowing about. Instead of compressing every part of the model equally, they keep the most important parts at higher precision and squeeze the less important parts harder. The S, M, and L are small, medium, and large versions of that trade, each spending a little more size for a little more quality. The practical takeaway is short. Given a K file and a plain file at the same bit count, take the K file. It spends its precision where the model feels it most.

What each level costs

From real use rather than theory:

At Q8, the loss is small enough that normal use will not surface it. The model behaves almost exactly like the full precision version. The catch is that Q8 is the largest of the quantized options, so it saves the least memory.

Q5 is the middle ground. It is a good one. The loss is small, and you only notice it on hard tasks if you go looking, while the memory savings are meaningful enough to change what else fits on the card alongside it.

Q4 saves the most memory and the loss becomes real. It shows up where you would expect: small factual slips get a little more frequent, and long, complicated instructions get followed slightly less reliably. Careful step by step reasoning, math especially, stumbles more often than it does at higher precision. Meanwhile casual chat, summarizing an article, rewriting a paragraph, answering a plain question, all of that is territory where telling Q4 from Q8 is a struggle. The damage concentrates in the hard cases and spares the easy ones, which means the right question is never which quant is best in general. It is which quant is best for the work you actually do.

Measurements of degradation back this shape up. Dropping from full precision to eight bits costs a fraction of a percent on most measures. Five bits costs low single digits on hard tasks. Four bits costs a bit more, and below four bits the loss starts climbing steeply. That cliff is why four bits became the popular floor: most of the memory is already squeezed out by then, while the quality is still clearly acceptable. The two and three bit files exist. Treat them as a last resort.

Bigger model at a lower quant usually wins

This is the pattern that surprised me, and it changed how I choose. Suppose a thirteen billion parameter model at four bits and a seven billion parameter model at eight bits take roughly the same VRAM. The bigger model at the harsher quant is usually the smarter pick. Counterintuitive, but it holds. The extra parameters carry more capability than the extra precision does.

So resist the instinct to max out the quant level on a small model. Stepping up to a larger model at a more aggressive quant is often where the real quality lives for the same memory.

One honest exception, because blanket rules mislead: very small models take quantization badly. A tiny model has little redundancy to spare, so squeezing it hurts proportionally more, and you may want to hold it at five or even eight bits to keep it usable. A very large model has so much capacity that heavy quantization barely dents it, so you can go harsher than you would expect. The four bit floor is a rule for mid sized models, and the model's size shifts it in both directions.

The memory arithmetic

Rough numbers, so you can plan a download instead of gambling on one. A seven billion parameter model at four bits lands around four to five gigabytes. The same model at eight bits sits closer to eight. A thirteen billion model at four bits is roughly eight gigabytes, and at eight bits it pushes past thirteen. The models with tens of billions of parameters only become possible on consumer cards at all because of aggressive quantization.

To use these numbers: take your card's VRAM, subtract a couple of gigabytes for overhead and working memory, and whatever remains is your budget for weights. Any model and quant combination inside that budget is a candidate.

Speed is part of the trade too

People treat quantization as purely a quality question, and it also touches speed, in both directions. Fewer bits means less data moving around, which can make a model faster. But some quantization formats need extra work at runtime to unpack the squeezed numbers into something the GPU can compute with, and that unpacking costs time. On my card the well supported four and five bit methods run fast and clean, while some of the more exotic formats run slower than their file size would suggest.

So when you try a new quant, check more than whether it loads. Watch the tokens per second you actually get. Two files of the same size can run at very different speeds.

Where the files come from

Most people meet quantized models through a tool like Ollama, which picks a quant and handles the download behind a simple model name. That is a fine place to start, and it is where I would start anyone. Once you want control, you go to where the community publishes quantized files directly, and you find a whole menu for a single model: every quant level laid out with its file size. That file size is your VRAM preview. If the file is bigger than the free memory on your card, it will not run well, and the fix is to step down a level.

Testing a quant without fooling yourself

Quantization loss is statistical. It arrives as a slightly higher rate of small mistakes spread across many runs, never as one obviously broken response. A single good answer does not prove a quant is fine, and a single bad answer does not prove it is broken. If you want to know whether a level is good enough for your work, push your real tasks through it a couple of dozen times and watch for the pattern of small slips. That is the only test that tells the truth.

I also give every fresh download a two minute acceptance check covering quality and speed. I have had files that loaded fine and ran strangely slow because my tools supported that format poorly. I have had others that loaded and produced subtly worse output because I grabbed a harsher quant than I meant to. Two minutes catches both before anything gets built on top of the mistake.

The rule I actually use

Stripped down: if the model fits comfortably at eight bits with room left for the work, I run Q8 and stop thinking about it. If Q8 is tight, I drop to Q5, my default workhorse, since the loss is small and the savings are real. Q4 is reserved for fitting a bigger model than would otherwise be possible, with eyes open about the hard tasks getting a little weaker. And for anything where correctness carries real weight, careful reasoning or code or precise instructions, I lean toward higher precision or a bigger model rather than the deepest squeeze.

One last reframe, since it took me a while to accept. I used to treat the quantized file as the budget version I settled for because of my hardware. The people running these models on data center cards quantize too, because the memory and speed wins are worth it even when better hardware is affordable. It is a good engineering trade. Understanding the levels just lets you pick your point on that trade deliberately, instead of guessing at a cryptic label and hoping the guess was cheap.

Get new guides and videos first — join the Telegram channel.