SkillByAIOpen interactive version →

Lesson 14 / 27

Label Quality: Noise Costs Accuracy

Wrong or inconsistent examples teach wrong or inconsistent behaviour.

Garbage in, confident garbage out

Fine-tuning copies the patterns in your data, including its mistakes, inconsistent style and hidden biases. A few hundred clean, consistent examples usually beat thousands of noisy ones. Common problems: contradictory labels for the same kind of input, answers that mix formats, copy-pasted boilerplate, outdated facts and examples that are really the assistant refusing or hedging. Review a random sample by hand, write down labelling rules, have two people label a subset and measure how often they agree, and remove or fix disagreements before training.

Quality, format and clean splits

Most fine-tuning outcomes are decided by the data, not by the training run.

Figure 4.1 — Labels, format and splits.

Test accuracy as labels get noisier, run

I ran this with Python, numpy 2.5.3 and scikit-learn 1.9.1, with fixed random seeds. It trains small classical models, not a language model: the mechanics (gradient descent, learning rate, overfitting, forgetting, low-rank updates) are the same ideas that apply to fine-tuning an LLM, but the numbers are not LLM results. Flipping a share of the training labels lowers held-out accuracy: 0.881 with clean labels, about 0.85 at 10 to 20% wrong, and 0.782 at 40% wrong. The size of the loss depends on the model and task.

import numpy as np
from sklearn.linear_model import LogisticRegression

rng = np.random.default_rng(5)
centers = rng.normal(size=(3, 20))
def make(n): y = rng.integers(0, 3, n); return centers[y] + rng.normal(size=(n, 20)) * 2.0, y
Xtr, ytr = make(400); Xte, yte = make(3000)

print("share of wrong labels in the training data | test accuracy")
for frac in (0.0, 0.1, 0.2, 0.4):
    y = ytr.copy(); flip = rng.random(len(y)) < frac; y[flip] = (y[flip] + rng.integers(1, 3, flip.sum())) % 3
    acc = LogisticRegression(max_iter=3000).fit(Xtr, y).score(Xte, yte)
    print(f"{frac:>42.0%} | {acc:.3f}")

Output:

share of wrong labels in the training data | test accuracy
                                        0% | 0.881
                                       10% | 0.849
                                       20% | 0.850
                                       40% | 0.782

Read 50 examples yourself

Before every training run, read 50 random rows; most data bugs are obvious by eye and invisible in aggregate metrics.

Quick check: Which is usually better for fine-tuning?

  • Duplicated boilerplate answers
  • Thousands of contradictory examples
  • Unreviewed scraped text
  • Hundreds of clean, consistent examples
Answer

Hundreds of clean, consistent examples — Quality and consistency matter more than raw volume.