# Learning Rate: The Most Sensitive Knob — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/h-lr

> Step size decides whether training converges or explodes.

## Too small is slow, too large diverges

The **learning rate** sets the size of each weight update. Too small and training barely moves; too large and updates overshoot, so loss bounces or **explodes**. Fine-tuning usually needs a **much smaller** rate than training from scratch because you want to refine, not overwrite. Typical full fine-tuning rates are small (around 1e-5 to 5e-5) and LoRA adapters often use larger ones (around 1e-4 to 3e-4), but follow current guidance for your model and tune on a validation set. Use **warm-up** and decay, and always plot training and validation loss.

## Four learning rates on the same problem, run

I ran this with Python, numpy 2.5.3 and scikit-learn 1.9.1, with fixed random seeds. It trains small classical models, not a language model: the mechanics (gradient descent, learning rate, overfitting, forgetting, low-rank updates) are the same ideas that apply to fine-tuning an LLM, but the numbers are not LLM results. The same least-squares problem is solved with four step sizes: 0.001 learns slowly (loss 11.8 to 9.7 in 60 steps), 0.05 and 0.4 reach near 0.01, and 1.2 diverges to about 7e31.

```python
import numpy as np
rng = np.random.default_rng(0)
X = rng.normal(size=(200, 5)); w_true = np.array([1.0, -2.0, 0.5, 0.0, 3.0]); y = X @ w_true + rng.normal(size=200) * 0.1

def train(lr, steps=60):
    w = np.zeros(5); losses = []
    for _ in range(steps):
        err = X @ w - y; losses.append(float(np.mean(err ** 2)))
        w -= lr * (2 / len(X)) * X.T @ err
    return losses
for lr in (0.001, 0.05, 0.4, 1.2):
    l = train(lr)
    note = "diverges (loss explodes)" if l[-1] > l[0] else "learns"
    print(f"learning rate {lr:<6} loss start {l[0]:9.3f} -> end {l[-1]:12.4g}   {note}")

```

Output:

```
learning rate 0.001  loss start    11.813 -> end        9.696   learns
learning rate 0.05   loss start    11.813 -> end      0.01021   learns
learning rate 0.4    loss start    11.813 -> end     0.009382   learns
learning rate 1.2    loss start    11.813 -> end     7.28e+31   diverges (loss explodes)
```

## Sweep three rates

Try your rate, half and double it on a small run; pick the best validation loss, not the lowest training loss.

**Quiz:** A loss that suddenly shoots up usually suggests what?

- [ ] The model is perfect
- [ ] Training is finished
- [x] The learning rate is too high (or the data/batch is bad)
- [ ] The GPU is faster

*Answer:* The learning rate is too high (or the data/batch is bad). Lower the learning rate and inspect the data and logs.
