Two ways to change a model's behaviour
Every team that adapts an open model faces the same fork early: freeze the base network and train a small set of new weights, or let the whole model move. LoRA โ low-rank adaptation โ takes the first path. It adds small trainable matrices to existing layers and leaves the original weights untouched. Full fine-tuning takes the second: gradients flow through every parameter, and you end up with a complete new set of weights.
The difference is not academic. A LoRA run on a 7B model can finish on a single consumer GPU in an hour or two. Full fine-tuning of the same model usually wants several high-memory GPUs, careful checkpointing and a real budget. That gap shapes most decisions before quality even enters the room.
The cost and hardware gap
LoRA's advantage is arithmetic. When one to five percent of parameters are trainable, optimiser state shrinks with them, cutting memory needs by roughly an order of magnitude. Full fine-tuning stores gradients and optimiser statistics for every weight in the network, so a 7B model can need well over 100 GB of GPU memory before it even starts learning.
- Hardware โ LoRA fits on a single 16 to 24 GB GPU for 7B to 13B models; full fine-tuning usually wants a multi-GPU node or aggressive sharding.
- Storage โ adapters weigh a few megabytes; each full checkpoint is tens of gigabytes.
- Serving โ adapters can be swapped per request at serving time; full models each need their own deployment.
- Iteration โ you can train a dozen LoRA variants in the time it takes to run one full fine-tune.
When LoRA is the right call
Most teams should start with LoRA, and most never need to leave it. It is strongest when the goal is to teach a style, a format, a tone or a bounded skill: structured output that always matches your field names, a support voice that sounds like your brand, consistent classification into your own categories. Because the base weights stay frozen, LoRA is also much less prone to catastrophic forgetting โ the model keeps its general abilities while picking up your specifics.
LoRA is also the pragmatic choice when data is limited. Adapter training is gentler, so a few thousand high-quality examples can shift behaviour meaningfully without the model collapsing into your dataset's quirks. And because the base model is shared, you can serve many customers or use cases from one deployment by swapping adapters.
When full fine-tuning earns its cost
Full fine-tuning starts to pay off when you need deep change. A model that must reason in a genuinely new syntax, absorb a large corpus as durable knowledge, or move far from its pretraining distribution often needs more than a thin adapter can store. Domain shifts โ turning a general model into a medical, legal or code specialist โ are the classic case. Work on full-rank updates keeps finding that when new knowledge must be memorised rather than merely elicited, low-rank adapters can hit a ceiling.
If you plan to distil the model, continue pretraining on billions of in-domain tokens, or ship it as a standalone artefact that others will fine-tune further, full fine-tuning is usually the honest choice.
Does one actually beat the other on quality?
On many benchmarks, a well-tuned LoRA lands within a point or two of full fine-tuning. The gaps show up in specific places: memorisation of new facts, very long training runs, and tasks where every fraction of a percent matters. If your evaluation set is stable and the difference is small, you have your answer โ take the cheaper method. Run both on a slice before committing to the expensive one, because a full run is a costly way to discover you did not need it.
Middle ground and hybrid approaches
The choice is not binary. You can fine-tune only the upper blocks of a network, train bias terms alone, or use prompt-based methods. Some teams run one longer full fine-tune, then maintain customer-specific variants as small adapters on top. Others merge several adapters into a single deployment. These hybrids often capture most of the benefit of full training while keeping the cost and flexibility of adapters for everyday iteration.
A practical decision checklist
- Start with LoRA unless you have evidence it is not enough.
- Match the method to the goal: style and format suit adapters; deep domain absorption points to full fine-tuning.
- Run a small bake-off with the same data and the same evaluation set for both methods.
- Judge on your own task, not on public leaderboards.
- Factor in serving: adapters swap cheaply, full models do not.
The bottom line
LoRA has become the default for good reason: it is fast, cheap, portable and surprisingly strong. Full fine-tuning remains the heavier tool for heavier jobs โ new domains, durable knowledge, maximum quality. The mature move is not to argue for one camp but to run a cheap experiment and let your own evaluation data pick the winner.



