Blog ยท Local AI

Running a Large Model Locally: What Hardware You Actually Need

2026-10-03ยท4 min readยทMD ABU SAYEED
In short: A model runs locally when its quantised weights plus about 20 percent overhead fit in your memory budget โ€” VRAM for speed, system RAM for patience โ€” so size the parameter count against your memory before anything else.
Network cables and ports in a server rack
Network cables and ports in a server rack

Three years ago, running a serious language model on your own hardware was a hobby for people with water-cooled workstations. Today a laptop bought for ordinary office work can serve a capable open model, and the difference between a smooth experience and a frustrating one comes down to a handful of numbers. You do not need to memorise benchmarks. You need to understand how model size, quantisation, memory and bandwidth interact โ€” because once you do, you can look at any model card and know in under a minute whether your machine can run it.

The only number that matters: memory

Weights are the bulk of any model: billions of floating-point numbers, each occupying a fixed amount of space. At 16-bit precision, a 7-billion-parameter model needs roughly 14 GB just to hold its weights. Quantise that same model to 4 bits and the requirement collapses to about 4 GB. Inference adds a little on top โ€” the KV cache that tracks your conversation plus activation buffers โ€” so budget an extra 10โ€“20 percent over the raw weights.

The rule of thumb fits on an index card: parameter count ร— bytes per weight + 20 percent = memory needed. If that result fits in your GPU's VRAM, the model runs at full speed. If it only fits in system RAM, the model still runs, but tokens arrive at a crawl, because system memory was never designed to feed a compute engine this hungry.

VRAM versus RAM: capacity fits, bandwidth speeds

Capacity and speed are two different problems. VRAM and fast system RAM can both hold a model, but they feed the processor at very different rates. A modern desktop GPU moves hundreds of gigabytes per second; a dual-channel desktop CPU setup manages a fraction of that. Because generating every token requires streaming the weights through the compute units, memory bandwidth is the true ceiling on tokens per second.

Unified-memory machines such as Apple Silicon laptops blur the categories in a good way: a 36 GB MacBook Air will happily load a 30B quantised model, though output arrives at reading speed rather than typing speed. That is fine for overnight batch jobs and painful for interactive chat. Decide honestly which one you are buying hardware for. And before you spend anything, run the free test: download a quantised model at the size you think you need, run it, and watch the tokens per second โ€” ten minutes of measurement beats hours of forum reading.

Compute: what the GPU actually does during inference

Here is the counterintuitive part: text generation barely stresses a GPU's arithmetic. Inference repeats the same matrix multiplications over and over, and moving weights into the compute units โ€” not the math itself โ€” is almost always the bottleneck. That is why a five-year-old GPU with abundant VRAM often matches a newer card for local inference, and why consumer cards with 16โ€“24 GB of memory dominate recommendations despite modest tensor performance.

Where compute does matter is prompt processing. Feeding a 20,000-token document into the context window is a compute-heavy burst, and a stronger GPU chews through it several times faster. If your use case is long-document question answering, spend on compute. If it is conversation, spend on memory instead.

Storage, power and cooling: the boring constraints

Model files are enormous. The full-precision download for a 70B model can exceed 130 GB; the 4-bit build of the same weights typically lands between 35 and 45 GB. A mechanical hard drive turns every launch into a coffee break, so an NVMe SSD is non-negotiable. Plan on roughly 200 GB of free space if you want to keep two or three models installed side by side.

Power and cooling come last on most shopping lists and first in most regrets. A loaded desktop GPU pulls 300โ€“450 watts, which means a real power supply, real airflow, and a room you can hear it in. Laptops throttle under sustained load; desktops sustain. Choose accordingly.

A realistic hardware ladder

Whichever tier you land on, start with the smallest model that can do the job, measure tokens per second on your real workloads, and only then spend more. Specific hardware advice ages quickly; the sizing method above does not.

Frequently asked questions

Can you run a 70B model on a laptop?
Yes, if the laptop has enough unified memory โ€” a 64 GB Apple Silicon machine can run a 70B model at 4-bit quantisation. Expect reading-speed output of a few tokens per second rather than instant chat responses, and keep the charger plugged in.
Do you need a GPU to run a large language model?
No. CPU inference works and the tooling has improved enormously, but memory bandwidth limits a typical desktop to a few tokens per second. A GPU matters for speed; it is optional for capability.
How much does context length change the memory budget?
Substantially at the high end. The KV cache grows with context, and at very long sessions on a 30B-class model it can add several gigabytes on its own. If you plan long conversations or big documents, reserve 20โ€“30 percent of VRAM beyond the weights.
Does quantisation ruin model quality?
Modern 4-bit quantisation costs a small amount of quality โ€” usually a point or two on benchmarks โ€” in exchange for a quarter of the memory footprint. For most assistant tasks the difference is hard to notice, and 8-bit is nearly lossless if you have the room.

Want a model trained on your own data?

Tell us what you want to build and we will send a scoped plan within 1โ€“2 business days.

Request a fit assessment

More from the blog

LoRA vs Full Fine-Tuning: How to Choose

LoRA vs Full Fine-Tuning: How to Choose

LoRA or full fine-tuning? Compare cost, quality, hardware and the signals that tell you which method fits when adapting an open model to your own domain.
2026-09-29
How to Keep a Trained Model From Going Stale

How to Keep a Trained Model From Going Stale

A trained model starts drifting the day you ship it. A practical routine for spotting staleness, refreshing data and retraining without breaking what works.
2026-09-28
Choosing a Base Model for a Domain-Specific Assistant

Choosing a Base Model for a Domain-Specific Assistant

How to pick the right base model for a domain-specific AI assistant: size, licence, context, language coverage and evaluation before you fine-tune.
2026-09-28