# Data Validation — MLOps

Source: https://www.skillbyai.com/en/mlops/d-validate

> Reject bad batches before they reach the model.

## Schema, ranges, freshness, volume

Validate every batch used for training or scoring: **schema** (columns and types), **missing values**, **ranges** and allowed categories, **freshness** (is the data from today?), **volume** (row counts within expected bounds) and **distribution** checks. Decide per check whether to block, quarantine bad rows or alert. Libraries such as Great Expectations, Pandera or TensorFlow Data Validation formalise this; even a few hand-written checks catch most broken exports.

## Validate, align, time-travel

Most production ML incidents start with data: bad batches, skew between training and serving, and leakage from the future.

![Three ideas: validation, skew, point-in-time features.](assets/figures/mlops/section-3-map.svg) — Figure 3.1 — Validation, skew and point-in-time features.

## Row-level checks on a batch, run

I ran this with Python 3, MLflow 3.16.1, scikit-learn 1.9.1, scipy 1.18.1 and numpy 2.5.3, using bundled or seeded synthetic data and a local SQLite tracking store. Four records are checked: a missing age, an impossible age of 212, a negative income and an unknown country code are reported, and with 75% of rows bad the batch is rejected (threshold 5%).

```python
rows = [
    {"age": 34, "income": 52000, "country": "IN", "signup_days": 120},
    {"age": None, "income": 61000, "country": "IN", "signup_days": 30},
    {"age": 212, "income": 48000, "country": "IN", "signup_days": 15},
    {"age": 45, "income": -5, "country": "XX", "signup_days": 400},
]
checks = {
    "age": lambda v: v is not None and 0 < v < 120,
    "income": lambda v: v is not None and v >= 0,
    "country": lambda v: v in {"IN", "US", "GB"},
    "signup_days": lambda v: v is not None and v >= 0,
}
failures = [(i, col, r[col]) for i, r in enumerate(rows) for col, ok in checks.items() if not ok(r[col])]
for i, col, v in failures:
    print(f"row {i}: {col} = {v!r} failed")
rate = len({i for i, *_ in failures}) / len(rows)
print(f"bad row rate {rate:.0%} -> {'REJECT batch' if rate > 0.05 else 'accept'}")
```

Output:

```
row 1: age = None failed
row 2: age = 212 failed
row 3: income = -5 failed
row 3: country = 'XX' failed
bad row rate 75% -> REJECT batch
```

## Validate before training too

A model retrained on a corrupted batch is a silent incident; run the same checks in the training pipeline.

**Quiz:** Which is a typical data validation check?

- [x] Values within allowed ranges and categories
- [ ] Model accuracy on the test set
- [ ] Number of engineers on call
- [ ] GPU temperature

*Answer:* Values within allowed ranges and categories. Check the data before trusting predictions.
