Three years ago, running a serious language model on your own hardware was a hobby for people with water-cooled workstations. Today a laptop bought for ordinary office work can serve a capable open model, and the difference between a smooth experience and a frustrating one comes down to a handful of numbers. You do not need to memorise benchmarks. You need to understand how model size, quantisation, memory and bandwidth interact โ because once you do, you can look at any model card and know in under a minute whether your machine can run it.
The only number that matters: memory
Weights are the bulk of any model: billions of floating-point numbers, each occupying a fixed amount of space. At 16-bit precision, a 7-billion-parameter model needs roughly 14 GB just to hold its weights. Quantise that same model to 4 bits and the requirement collapses to about 4 GB. Inference adds a little on top โ the KV cache that tracks your conversation plus activation buffers โ so budget an extra 10โ20 percent over the raw weights.
The rule of thumb fits on an index card: parameter count ร bytes per weight + 20 percent = memory needed. If that result fits in your GPU's VRAM, the model runs at full speed. If it only fits in system RAM, the model still runs, but tokens arrive at a crawl, because system memory was never designed to feed a compute engine this hungry.
- 8 GB VRAM: 7โ8B models at 4-bit โ genuinely useful for chat, summaries and drafting.
- 12โ16 GB VRAM: 13โ14B models at 4-bit, or an 8B model at 8-bit โ the sweet spot for daily work.
- 24 GB VRAM: 30B-class quantised models, where a local assistant starts feeling like a hosted one.
- 48 GB and up: 70B-class quantised models โ workstation territory, but no longer exotic.
VRAM versus RAM: capacity fits, bandwidth speeds
Capacity and speed are two different problems. VRAM and fast system RAM can both hold a model, but they feed the processor at very different rates. A modern desktop GPU moves hundreds of gigabytes per second; a dual-channel desktop CPU setup manages a fraction of that. Because generating every token requires streaming the weights through the compute units, memory bandwidth is the true ceiling on tokens per second.
Unified-memory machines such as Apple Silicon laptops blur the categories in a good way: a 36 GB MacBook Air will happily load a 30B quantised model, though output arrives at reading speed rather than typing speed. That is fine for overnight batch jobs and painful for interactive chat. Decide honestly which one you are buying hardware for. And before you spend anything, run the free test: download a quantised model at the size you think you need, run it, and watch the tokens per second โ ten minutes of measurement beats hours of forum reading.
Compute: what the GPU actually does during inference
Here is the counterintuitive part: text generation barely stresses a GPU's arithmetic. Inference repeats the same matrix multiplications over and over, and moving weights into the compute units โ not the math itself โ is almost always the bottleneck. That is why a five-year-old GPU with abundant VRAM often matches a newer card for local inference, and why consumer cards with 16โ24 GB of memory dominate recommendations despite modest tensor performance.
Where compute does matter is prompt processing. Feeding a 20,000-token document into the context window is a compute-heavy burst, and a stronger GPU chews through it several times faster. If your use case is long-document question answering, spend on compute. If it is conversation, spend on memory instead.
Storage, power and cooling: the boring constraints
Model files are enormous. The full-precision download for a 70B model can exceed 130 GB; the 4-bit build of the same weights typically lands between 35 and 45 GB. A mechanical hard drive turns every launch into a coffee break, so an NVMe SSD is non-negotiable. Plan on roughly 200 GB of free space if you want to keep two or three models installed side by side.
Power and cooling come last on most shopping lists and first in most regrets. A loaded desktop GPU pulls 300โ450 watts, which means a real power supply, real airflow, and a room you can hear it in. Laptops throttle under sustained load; desktops sustain. Choose accordingly.
A realistic hardware ladder
- Try it free first: any machine with 16 GB of RAM runs 7โ8B quantised models, slowly. Perfect for learning the tooling.
- Entry desktop build: a GPU with 16 GB of VRAM, 32 GB of RAM and a 1 TB NVMe drive runs 13โ14B models comfortably.
- Serious local AI: 24 GB of VRAM and 64 GB of RAM for 30B-class models at interactive speed.
- Workstation: 48 GB or more of VRAM, or a multi-GPU box, for 70B-class models and long-context work.
Whichever tier you land on, start with the smallest model that can do the job, measure tokens per second on your real workloads, and only then spend more. Specific hardware advice ages quickly; the sizing method above does not.



