# LoRA and Parameter-Efficient Fine-Tuning — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/h-lora

> Train a tiny low-rank update instead of the whole weight matrix.

## The needed change is often low-rank

Updating every weight is expensive and yields a full model copy per task. **Parameter-efficient fine-tuning (PEFT)** trains only a few added parameters. **LoRA** freezes each chosen weight matrix `W` and learns `ΔW = A·B`, where `A` is `d×r` and `B` is `r×d` with a small **rank** `r` (for example 8). That is `r·2d` trainable numbers instead of `d²`: for `d = 4096`, `r = 8`, about 65,000 instead of 16.8 million per matrix (0.39%). The bet is that the needed change is **low-rank**. The adapter is megabytes, can be **swapped per task** on one base model or **merged** for serving; **QLoRA** also stores the frozen base in 4-bit form.

## Recovering a rank-2 change with a rank-2 adapter, run

I ran this with Python, numpy 2.5.3 and scikit-learn 1.9.1, with fixed random seeds. It trains small classical models, not a language model: the mechanics (gradient descent, learning rate, overfitting, forgetting, low-rank updates) are the same ideas that apply to fine-tuning an LLM, but the numbers are not LLM results. A frozen 32x32 matrix needs a rank-2 change. With no update the error is 3.29; a rank-1 adapter (64 trainable numbers, 6%) cuts it to 1.64; rank 2 (128 numbers, 12%) reaches essentially zero, and rank 4 also reaches zero.

```python
import numpy as np
rng = np.random.default_rng(1)
d, true_rank = 32, 2

W0 = rng.normal(size=(d, d)) / np.sqrt(d)                      # frozen "pretrained" weight
delta = (rng.normal(size=(d, true_rank)) @ rng.normal(size=(true_rank, d))) * 0.3   # the change the new task needs: rank 2
W_target = W0 + delta

X = rng.normal(size=(400, d)); Y = X @ W_target.T               # task data produced by the target weights

def train_lora(r, steps=3000, lr=0.02):
    A = rng.normal(size=(d, r)) * 0.1; B = np.zeros((r, d))     # W = W0 + A @ B ; only A and B are trained
    for _ in range(steps):
        W = W0 + A @ B
        err = X @ W.T - Y                                       # (n, d)
        G = err.T @ X / len(X)                                   # gradient wrt W  (d, d)
        gA, gB = G @ B.T, A.T @ G
        A -= lr * gA; B -= lr * gB
    return float(np.mean((X @ (W0 + A @ B).T - Y) ** 2)), A.size + B.size

print("frozen W0 only: loss", round(float(np.mean((X @ W0.T - Y) ** 2)), 4), "| full matrix has", d * d, "parameters")
for r in (1, 2, 4):
    loss, params = train_lora(r)
    print(f"LoRA rank {r}: loss {loss:.5f} with {params} trainable parameters ({params / (d * d):.0%} of the full matrix)")

```

Output:

```
frozen W0 only: loss 3.2905 | full matrix has 1024 parameters
LoRA rank 1: loss 1.63657 with 64 trainable parameters (6% of the full matrix)
LoRA rank 2: loss 0.00000 with 128 trainable parameters (12% of the full matrix)
LoRA rank 4: loss 0.00000 with 256 trainable parameters (25% of the full matrix)
```

## Training memory: full fine-tune versus LoRA, run

I ran this with plain Python 3 (standard library only), using example numbers. With the rough estimate of 16 bytes per trained parameter, a 7B model needs about 112 GB to fully fine-tune but about 14 GB when only a 0.4% adapter trains over a frozen 2-byte base. Rough planning numbers before activations and overhead.

```python
# Rough training memory (mixed precision with Adam): weights 2 B + gradients 2 B + fp32 master weights 4 B + 2 optimizer states 8 B
# = about 16 bytes per TRAINED parameter, plus frozen weights at 2 B per parameter. Activations come on top.
def gb(x): return x / 1e9
def full_ft(params): return params * 16
def lora(params, trainable): return params * 2 + trainable * 16     # frozen base at 16-bit + trained adapters

for name, params in (("7B", 7e9), ("13B", 13e9), ("70B", 70e9)):
    trainable = params * 0.004                                       # ~0.4% trainable with a small-rank adapter (example)
    print(f"{name:4} full fine-tune ~ {gb(full_ft(params)):7.0f} GB   LoRA ~ {gb(lora(params, trainable)):6.0f} GB   (before activations)")

```

Output:

```
7B   full fine-tune ~     112 GB   LoRA ~     14 GB   (before activations)
13B  full fine-tune ~     208 GB   LoRA ~     27 GB   (before activations)
70B  full fine-tune ~    1120 GB   LoRA ~    144 GB   (before activations)
```

**Quiz:** What does LoRA train?

- [x] Small low-rank matrices added to frozen weights
- [ ] Every weight of the model
- [ ] Only the tokenizer
- [ ] The user interface

*Answer:* Small low-rank matrices added to frozen weights. The adapter learns a compact update while the base stays frozen.
