VRAM math: how to know what fits before you download
# VRAM math: how to know what fits before you download
The number one reason a local AI model runs painfully slow is embarrassingly simple: it does not fit on the graphics card, so part of it spills over to the CPU, and everything crawls. Almost nobody checks for this before downloading. They grab a model, watch it produce a few words per second, and conclude their hardware is too weak or that local AI is overhyped. Usually neither is true. They picked a model that does not fit.
I size every model against my card before downloading a single gigabyte, and the arithmetic is short enough to do in your head. Here it is, end to end.
VRAM is the ceiling
Everything hangs on one resource: VRAM, the memory built onto the graphics card itself. It is separate from your regular system memory, it sits right next to the silicon doing the math, it is fast, and there is never much of it. A typical gaming card carries eight, eleven, or sixteen gigabytes.
That number is a hard ceiling. When a model fits inside it, the model flies. When it does not, the model stalls. There is no in between that feels okay, which is why the whole game is predicting fit before you commit.
Three things share the card
Most guides only count one of them. What actually has to live in VRAM, all at the same time:
The weights, the giant pile of numbers that is the model itself. This is the big chunk and the one everybody budgets for.
The context, the working memory the model builds up from your conversation and whatever text it is currently processing. This one grows as you use the model.
The overhead, the scratch space the system needs to actually run the computation.
Budget only for the weights and you will load a model that technically fits, then watch it spill the first time you hand it a long prompt.
The weights formula
Take the parameter count in billions and multiply by bytes per parameter. At full precision that is two bytes each, so a seven billion parameter model wants about fourteen gigabytes for weights alone, which rules out most consumer cards immediately. That is exactly why quantization exists. At four bits per weight you are at roughly half a byte per parameter, and the same model drops to around four gigabytes.
The mental shortcut I use daily: a four bit model costs roughly half a gigabyte of VRAM per billion parameters. Seven billion lands around three and a half to four gigabytes. Thirteen billion lands around seven. That one rule of thumb does ninety percent of the work.
Context is the part people forget
As the model processes text, it accumulates a working memory of everything in the current conversation, and that memory grows with the amount of text involved. A short question costs almost nothing. Paste in a long document, or carry on a very long back and forth, and the context can balloon to a gigabyte or several on its own.
This explains a failure that confuses everyone the first time: a model that ran fine on short questions suddenly chokes when you feed it something big. The model never changed. The context grew and pushed the total past what the card holds. For a typical mid sized model, a short conversation might use a few hundred megabytes, barely a rounding error, while a document heavy session climbs into the gigabytes.
So when you plan, picture the model working on your largest realistic input rather than the model sitting idle. The worst case is what has to fit.
Leave a buffer
The practical version of all that: after estimating the weights, leave headroom on top. A couple of gigabytes minimum, more if long documents are part of your normal use. On my eleven gigabyte card I do not shop for eleven gigabytes of weights. I shop for around seven or eight, keeping three or four free for context and overhead.
That buffer separates a setup that stays fast all day from one that mysteriously drags whenever it gets real work. People who skip the buffer are heavily represented among people posting that local AI is unreliable. Nothing is wrong with their hardware. The card is simply full.
Two worked examples
An eight gigabyte card, considering a seven billion parameter model at four bits: about four gigabytes of weights, plus a couple for context and overhead, totals around six. Comfortable fit, will run fast, good choice.
Same card, eyeing a thirteen billion model at four bits: about seven gigabytes of weights, plus buffer, reaches nine or ten. Over the ceiling. That one spills and crawls. On eight gigabytes, the seven billion class is home and thirteen billion is a stretch to avoid unless you compress it much harder and accept what that costs.
Move to a sixteen gigabyte card and the picture flips. The thirteen billion model at four bits, seven gigabytes plus buffer, fits comfortably and runs fast, with room to reach larger models or run the thirteen billion at higher precision for better quality.
Spilling is a cliff, not a slope
When a model does not fit, the system does not refuse to run it. It quietly moves part of the model into regular system memory and runs that part on the CPU. The CPU and system RAM are dramatically slower at this work than the GPU, so even a small spilled fraction becomes the bottleneck for every single token. Speed can drop by ten or twenty times.
Internalize the shape of that: being slightly over the limit costs you almost as much as being way over it. There is no gentle degradation to hide in. Fit is close to binary, which is why doing the arithmetic beforehand pays off so well.
The file size trap
The download size you see listed is a decent first estimate of the weight memory. It is also only the weights, as they sit on disk. Once loaded, you still pay for context and overhead on top, so a four gigabyte file does not mean four gigabytes of VRAM. Treat the file size as the floor of the true cost. If the file alone already crowds your card's capacity, that is your signal to step down a size or compress harder, because nothing would be left for the actual work.
Two cards, and the system RAM misconception
Two questions come up constantly, so here are the honest answers.
Can you run two graphics cards? Yes. A model can be split across them and their memory adds up, so two eight gigabyte cards can hold a model that needs around sixteen. The catch is that the cards must talk to each other constantly while running, and that communication adds overhead, so two cards run somewhat slower than one card with the same total memory would. For fitting a model no single card you own can hold, splitting is a real and useful option. Just do not expect speed to scale with the added memory.
Does lots of system RAM help? Less than people assume. VRAM and system RAM are separate pools, and only VRAM is fast for this work. Plenty of system memory simply lets the slow overflow happen instead of failing outright. People see sixty four gigabytes of RAM and assume big models are in reach, and technically they are, at speeds you will not enjoy. The fast capacity is the VRAM, and that is the number governing your experience.
Buying with this arithmetic
This math is exactly what should drive a hardware purchase for local AI. Prioritize the amount of VRAM ahead of raw speed, because a faster card with less memory cannot load models that a slower card with more memory runs every day. A used card from a previous generation with a generous memory pool is very often the smarter buy than a newer card with less. And each step from eight to twelve to sixteen gigabytes unlocks a meaningfully larger class of models. When your current card starts feeling tight, the upgrade that helps is the one that adds memory, ahead of the one with bigger benchmark numbers.
The ladder, and one free knob
The day to day payoff is a rough ladder you memorize once. On eight gigabytes, you live comfortably in the seven to eight billion parameter range at four bit compression, with room for normal context. On twelve, mid sized models become comfortable and the smaller ones can run at higher quality. On sixteen, the thirteen billion class is comfortable with real headroom for long context, and above that the larger models start opening up. Know where your card sits and you can judge any model you come across at a glance, without redoing the full arithmetic.
One more lever is worth knowing. Many tools let you cap the context length, and a smaller cap reserves less memory for context, which frees VRAM. When a model is almost fitting, dialing the context limit down is often what makes it run smoothly instead of spilling. The trade is that the model handles less text at once, and for plenty of everyday tasks you never needed a huge context anyway. It is the first knob I reach for when something nearly fits.
The whole method in order
Find the model's parameter count and precision. Estimate the weights, roughly half a gigabyte per billion parameters at four bits. Add a buffer of a couple of gigabytes for context and overhead, more for document heavy work. Compare the total to your card's VRAM. Fits with room to spare: it will run fast. Over the line: step down a size, compress harder, or cap the context, then check again.
Do that before you download and you will almost never be surprised by a model that crawls. It is boring arithmetic, and it is precisely the kind of boring that separates a local setup that just works from one that fights you constantly.
Get new guides and videos first — join the Telegram channel.