Lesson 7 / 25

Data Validation

Reject bad batches before they reach the model.

Schema, ranges, freshness, volume

Validate every batch used for training or scoring: schema (columns and types), missing values, ranges and allowed categories, freshness (is the data from today?), volume (row counts within expected bounds) and distribution checks. Decide per check whether to block, quarantine bad rows or alert. Libraries such as Great Expectations, Pandera or TensorFlow Data Validation formalise this; even a few hand-written checks catch most broken exports.

Validate, align, time-travel

Most production ML incidents start with data: bad batches, skew between training and serving, and leakage from the future.

Three ideas: validation, skew, point-in-time features.
Figure 3.1 — Validation, skew and point-in-time features.

Row-level checks on a batch, run

I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. Four records are checked: a missing age, an impossible age of 212, a negative income and an unknown country code are reported, and with 75% of rows bad the batch is rejected (threshold 5%).

rows = [
    {"age": 34, "income": 52000, "country": "IN", "signup_days": 120},
    {"age": None, "income": 61000, "country": "IN", "signup_days": 30},
    {"age": 212, "income": 48000, "country": "IN", "signup_days": 15},
    {"age": 45, "income": -5, "country": "XX", "signup_days": 400},
]
checks = {
    "age": lambda v: v is not None and 0 < v < 120,
    "income": lambda v: v is not None and v >= 0,
    "country": lambda v: v in {"IN", "US", "GB"},
    "signup_days": lambda v: v is not None and v >= 0,
}
failures = [(i, col, r[col]) for i, r in enumerate(rows) for col, ok in checks.items() if not ok(r[col])]
for i, col, v in failures:
    print(f"row {i}: {col} = {v!r} failed")
rate = len({i for i, *_ in failures}) / len(rows)
print(f"bad row rate {rate:.0%} -> {'REJECT batch' if rate > 0.05 else 'accept'}")

Output:

row 1: age = None failed
row 2: age = 212 failed
row 3: income = -5 failed
row 3: country = 'XX' failed
bad row rate 75% -> REJECT batch

Validate before training too

A model retrained on a corrupted batch is a silent incident; run the same checks in the training pipeline.

Quick check: Which is a typical data validation check?

  • Values within allowed ranges and categories
  • Model accuracy on the test set
  • Number of engineers on call
  • GPU temperature
Answer

Values within allowed ranges and categories — Check the data before trusting predictions.