# Distillation: Teaching a Small Model From a Large One — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/h-distill

> Use a strong model to label data that trains a cheaper one.

## Copy the behaviour, not the size

**Distillation** trains a small **student** to imitate a larger **teacher**. A common LLM recipe: collect realistic inputs, have a strong model answer them (with checks), review a sample, and fine-tune a smaller model on those pairs. On a **narrow task** the student can approach the teacher at far lower cost and latency. Cautions: the student inherits the teacher's mistakes; check the provider's **terms** on training other models with outputs; measure on **human-verified** test data, not just agreement with the teacher; keep the data diverse.

## A student helped by teacher labels, run

I ran this with Python, numpy 2.5.3 and scikit-learn 1.9.1, with fixed random seeds. It trains small classical models, not a language model: the mechanics (gradient descent, learning rate, overfitting, forgetting, low-rank updates) are the same ideas that apply to fine-tuning an LLM, but the numbers are not LLM results. The linear teacher scores 0.885 on held-out data. A small tree trained on 300 noisy real labels scores 0.723; given teacher labels for 4,700 more unlabelled inputs it scores 0.791, improving but staying below the teacher.

```python
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier

rng = np.random.default_rng(6)
X = rng.normal(size=(8000, 12)); w = rng.normal(size=12)
y = (X @ w + rng.normal(size=8000) * 1.2 > 0).astype(int)      # noisy real labels
Xtr, ytr, Xte, yte = X[:300], y[:300], X[5000:], y[5000:]

teacher = LogisticRegression(max_iter=2000).fit(Xtr, ytr)                       # the "big" model trained on the labelled data
Xunl = X[300:5000]                                                               # plenty of UNLABELLED inputs
teacher_labels = teacher.predict(Xunl)

student_alone = DecisionTreeClassifier(max_depth=6, random_state=0).fit(Xtr, ytr)
student_distilled = DecisionTreeClassifier(max_depth=6, random_state=0).fit(np.vstack([Xtr, Xunl]), np.concatenate([ytr, teacher_labels]))
print("teacher                         :", round(teacher.score(Xte, yte), 3))
print("small student, real labels only :", round(student_alone.score(Xte, yte), 3))
print("small student + teacher labels  :", round(student_distilled.score(Xte, yte), 3))

```

Output:

```
teacher                         : 0.885
small student, real labels only : 0.723
small student + teacher labels  : 0.791
```

## Sample and read teacher outputs

Read at least a few dozen teacher answers before training; every systematic error you skip becomes the student's habit.

**Quiz:** How should a distilled student be evaluated?

- [ ] Only by how closely it copies the teacher
- [x] On human-verified test data, not only agreement with the teacher
- [ ] Not at all
- [ ] Only by model size

*Answer:* On human-verified test data, not only agreement with the teacher. Matching the teacher's errors is not the goal; being correct is.
