Featured image of post LoRA: Teach a Model Without Rewriting the Whole Thing

LoRA: Teach a Model Without Rewriting the Whole Thing

LoRA freezes pretrained weights and learns a low-rank update. This article explains the matrix factorization, the real training ledger, adapter merging, QLoRA, and the limits of the low-rank assumption.

The obvious way to fine-tune one model for ten tasks is to save ten complete weight sets. Yet the tasks did not create ten foundations. Most pretrained capability remains shared; only the change required by each dataset is different. As models grow, copying the whole object merely to preserve that difference becomes the expensive part.

LoRA—Low-Rank Adaptation—changes the object being optimized. Instead of directly learning a new weight \(W'\), it freezes the pretrained weight \(W\) and learns only an update \(\Delta W\). It then assumes that this update does not need the entire high-dimensional space and can be represented by two narrow matrices:

\[ W' = W + \Delta W, \qquad \Delta W = \frac{\alpha}{r}BA. \]

The important move is not simply “attach a small network.” It is to separate the foundation from the task-specific delta. The base still supplies almost all computation and capability; the adapter records the directions in which this task needs to move it. In its GPT-3 175B setup, the LoRA paper reported up to 10,000 times fewer trainable parameters and roughly one-third the GPU memory of full Adam fine-tuning. Those are results for particular configurations, not universal multipliers.

A Full Update Becomes Two Narrow Projections

Consider a linear weight \(W\in\mathbb{R}^{d_{out}\times d_{in}}\). Full fine-tuning can change all \(d_{out}d_{in}\) entries. LoRA uses

\[ A\in\mathbb{R}^{r\times d_{in}},\qquad B\in\mathbb{R}^{d_{out}\times r},\qquad r\ll\min(d_{in},d_{out}), \]

and computes

\[ y=Wx+\frac{\alpha}{r}B(Ax). \]

The base path \(Wx\) stays frozen. The side path compresses the input to \(r\) dimensions and projects it back to the output space. It contains only \(r(d_{in}+d_{out})\) parameters. For a \(4096\times4096\) matrix with \(r=8\), the full matrix has 16,777,216 entries while the two LoRA matrices contain 65,536—exactly \(1/256\) as many. That is the geometry for one adapted matrix; the whole-model ratio also depends on which modules receive adapters.

Figure 1: LoRA leaves W frozen and represents its update with a down-projection followed by an up-projection

The original implementation initializes \(A\) randomly and \(B\) to zero. Training therefore begins with \(BA=0\), so the model initially behaves exactly like the base. The factor \(\alpha/r\) controls update magnitude and makes scale easier to manage when rank changes. Gradients and optimizer states are maintained only for \(A\) and \(B\); \(W\) still participates in forward and backward computation but is not updated.

That distinction matters. LoRA reduces trainable parameters, parameter gradients, and optimizer states. It does not delete the large model’s compute. Frozen weights still need to be stored or streamed, and activations still need to be retained for backpropagation. Sequence length, batch size, activation checkpointing, and kernel efficiency still affect throughput and peak memory. Training 0.x% of the parameters does not mean using 0.x% of the memory or arithmetic.

Low Rank Is a Hypothesis About Task Change

Any matrix can be represented if the rank is high enough. LoRA’s real bet is that a downstream task needs a \(\Delta W\) with low “intrinsic rank.” A pretrained model already contains broad linguistic and world patterns. Adaptation may only need to amplify, suppress, or recombine a small set of directions rather than relearn those patterns.

In some GPT-3 experiments, the original paper found that rank one remained competitive when adapting query and value projections together. Updates learned at different ranks also shared a few leading singular directions. The authors explicitly cautioned that a tiny rank should not be expected to work for every task or dataset. The evidence supports low-rank approximations for many adaptations; it does not prove that every new domain, fact set, or layer is naturally low-rank.

Rank is therefore neither a contest to choose the smallest number nor a guarantee that larger is safer. Too little rank can bottleneck expression. More rank consumes memory and communication bandwidth and may simply saturate. Target modules matter just as much. The original work focused on self-attention projections and found that, under the same parameter budget, covering query and value could be better than giving a larger rank to one projection. Modern libraries can target key, output, and MLP linear layers as well, but the choice should be validated for the task rather than inherited as folklore.

The Largest Savings Are in Training State and Versioning

Full fine-tuning produces a complete model per task. Adaptive optimizers also maintain gradients plus first- and second-moment states for every trainable weight. LoRA confines these task-growing objects to small matrices. One base can support many adapters, and each checkpoint can contain only LoRA parameters; Microsoft’s reference library exposes a LoRA-only state_dict for exactly this purpose.

Figure 2: full fine-tuning, LoRA, and QLoRA reduce different parts of the ledger; freezing weights does not remove base computation or activations

This changes model delivery. “One task, one model” previously meant copying, transferring, and deploying every weight. With LoRA, the base becomes a shared runtime and task differences become compact patches. Experiments, rollbacks, and variant storage become cheaper, and the delta learned from private data can be governed separately.

It does not make every cost small. Base-weight matrix multiplication remains. Activations can dominate training memory. Distributed execution still has to place the frozen weights and communicate adapter gradients. LoRA most directly cuts the state that must be updated and persisted; end-to-end speed and peak memory still require workload-level measurement.

Merging Removes the Extra Path, but Multi-Tenant Serving Is Not Free

Because the update is linear, deployment can precompute

\[ W_{merged}=W+\frac{\alpha}{r}BA. \]

Inference then performs an ordinary \(W_{merged}x\). A single merged adapter does not increase network depth the way a serial adapter layer inside each Transformer block would. This is the precise scope of the paper’s “no additional inference latency” claim: a merged single-adapter path has the same matrix-multiplication shape as a fully fine-tuned weight.

A service hosting many adapters on one base has a different problem. Mutually different updates cannot all be permanently merged into the same \(W\). The system must select adapters per request, execute the side path dynamically, or keep multiple merged copies. Small checkpoints make storage and switching easier, but batching, residency, cache behavior, and kernel fusion remain systems concerns. “Can be merged” does not mean “unlimited concurrent adapters are free.”

Figure 3: one base can share many small adapters; one adapter can be merged, while concurrent adapters still require routing and scheduling

The publishable artifact is consequently no longer just “the model.” It is the base version, adapter version, and exact configuration. Rank, alpha, target modules, bias handling, and a base-model identifier all matter. An adapter applied to the wrong base may have compatible dimensions while carrying the wrong semantics.

QLoRA Compresses the Frozen Base to 4 Bits

Ordinary LoRA still has to fit the base model. QLoRA stores the frozen pretrained weights in 4-bit form, dequantizes them to a compute type as needed, and lets gradients flow through that path into higher-precision LoRA adapters. The quantized base itself is not updated. This is not “train every parameter in 4 bits”; it is “hold the frozen base in 4 bits and train a low-rank delta at a suitable precision.”

The QLoRA paper introduced NF4 for approximately normally distributed weights, double quantization to compress the quantization constants, and paged optimizers to manage memory spikes. It demonstrated fine-tuning a 65B-parameter model on one 48 GB GPU and matched full 16-bit fine-tuning in the evaluated settings. That result shows that base storage and trainable state can be compressed together. It does not imply that every 65B model, sequence length, and batch fits in 48 GB; activations, target modules, optimizer choice, and implementation all change the peak.

Quantization does not make multiplication free. Training and inference still read, dequantize, and compute with the base weights, and hardware support for 4-bit kernels varies. QLoRA’s durable contribution is to make base precision and update precision separate design axes, not to guarantee that it is faster than ordinary LoRA in every environment.

When Another Method Is the Better Answer

LoRA fits tasks where the base already has sufficient capability, the goal is behavioral or domain adaptation, and many variants must be stored. It does not guarantee that a small rank can reliably absorb a large body of knowledge. Nor does it correct poor data, mislabeled examples, or evaluation leakage. A small adapter only means a small parameter ledger, not behavior that is easy to audit.

If LoRA at a reasonable rank consistently trails full fine-tuning on a task far from the base distribution, the low-rank assumption should be treated as a failed experimental hypothesis—not defended by adding more tricks. If a few examples in context solve the problem, prompting or retrieval may be simpler than training. If knowledge must remain current and attributable, RAG may be preferable to embedding facts in weights. The right method depends on whether the desired change is behavior, knowledge, or foundational capability.

LoRA’s most useful intuition is to price the model separately from its task-specific change. An effective change to a huge system may occupy only a few directions in a high-dimensional space. Saving those directions can remove most gradient, optimizer, checkpoint, and deployment duplication. Base computation, activation memory, and data risk remain. LoRA is not free fine-tuning; it is a more accurate ledger.

References

  1. Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, ICLR 2022.
  2. Microsoft, Official LoRA / loralib implementation.
  3. Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs, NeurIPS 2023.
  4. QLoRA authors, Official QLoRA repository.
  5. Hugging Face, Official PEFT LoRA documentation.