# Learning Curves: When Training Overtakes Examples in the Prompt — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/p-curve

> See why more labelled data can beat a handful of shots.

## A classical analogy for the trade-off

In a classical toy task, one method keeps labelled examples and matches each new input to its nearest example (like **few-shot prompting**: it cannot learn a global rule), while the other **trains** a classifier on all examples (like fine-tuning). With very few examples they are similar; as data grows the trained model keeps improving while example matching flattens. Real LLM prompting is far more capable than nearest-neighbour matching, so this only illustrates the **shape** of the curves. Check the curve for your own task: if accuracy is still rising as you add data, training may help.

## Examples as lookup versus training, run

I ran this with Python, numpy 2.5.3 and scikit-learn 1.9.1, with fixed random seeds. It trains small classical models, not a language model: the mechanics (gradient descent, learning rate, overfitting, forgetting, low-rank updates) are the same ideas that apply to fine-tuning an LLM, but the numbers are not LLM results. With 3 labelled examples both score about 0.46; at 1,500 examples nearest-example reaches 0.689 while the trained classifier reaches 0.846. The toy only shows curve shapes.

```python
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier

# A toy "task": tell 3 ticket categories apart from 20 noisy numeric features (a stand-in for text features).
rng = np.random.default_rng(0)
centers = rng.normal(size=(3, 20)) * 1.0
def make(n): 
    y = rng.integers(0, 3, n); X = centers[y] + rng.normal(size=(n, 20)) * 2.2
    return X, y
Xtest, ytest = make(3000)

print("labelled examples | few-shot nearest-example (prompt-like) | trained classifier (fine-tune-like)")
for n in (3, 6, 15, 60, 300, 1500):
    X, y = make(n)
    knn = KNeighborsClassifier(n_neighbors=1).fit(X, y)
    clf = LogisticRegression(max_iter=2000).fit(X, y)
    print(f"{n:17d} | {knn.score(Xtest, ytest):37.3f} | {clf.score(Xtest, ytest):.3f}")

```

Output:

```
labelled examples | few-shot nearest-example (prompt-like) | trained classifier (fine-tune-like)
                3 |                                 0.458 | 0.462
                6 |                                 0.549 | 0.575
               15 |                                 0.631 | 0.726
               60 |                                 0.635 | 0.717
              300 |                                 0.662 | 0.826
             1500 |                                 0.689 | 0.846
```

## Plot before you decide

Score your task with 10, 50, 200 examples in the prompt or training set; the slope tells you whether more data is worth collecting.

**Quiz:** What does the learning curve suggest?

- [ ] Both lines are identical
- [ ] More data always hurts
- [ ] Training never helps
- [x] A trained model can keep improving where example matching flattens

*Answer:* A trained model can keep improving where example matching flattens. Check your own curve: if accuracy still climbs with data, training may pay off.
