# Supervised Fine-Tuning (SFT) in Plain Terms — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/h-sft

> What the training loop actually does to the model.

## Show examples, nudge weights, repeat

**Supervised fine-tuning (SFT)** trains a pretrained model on pairs of **input and desired output**. For each example the model predicts the answer tokens, a **loss** measures the distance from the target, **backpropagation** computes each weight's contribution, and an **optimiser** (such as Adam) nudges weights to reduce the loss. This repeats over batches for one or a few **epochs**. Because the model starts with strong general ability it needs far fewer examples than training from scratch, but big steps or poor data can distort what it knew. The loss is usually computed only on the **assistant answer tokens**.

## Small steps on a pretrained model

Fine-tuning continues gradient descent on your examples, ideally changing only a small part of the weights.

![Five ideas: SFT, learning rate, LoRA, preferences, distillation.](assets/figures/fine-tuning/section-3-map.svg) — Figure 3.1 — SFT, learning rate, LoRA, preferences and distillation.

## The training loop in one picture

Conceptual, not tied to one library.

```text
for epoch in 1..E:
    for batch in shuffled(training_examples):
        predictions = model(batch.input)                        # forward pass
        loss = cross_entropy(predictions, batch.answer_tokens)  # only the answer part counts
        gradients = backward(loss)                              # how each weight affects the loss
        weights -= learning_rate * adjust(gradients)            # small step (Adam etc.)
    check validation loss on held-out examples                  # stop early if it stops improving
```

## Mask the prompt

Train on the answer tokens only; otherwise the model also spends capacity learning to imitate user prompts.

**Quiz:** What does SFT optimise?

- [ ] The price per token
- [ ] The length of the user prompt
- [x] The weights, to reduce loss on desired example outputs
- [ ] The GPU fan speed

*Answer:* The weights, to reduce loss on desired example outputs. Training adjusts weights to make the target answers more likely.
