# Testing ML Code, Data and Models — MLOps

Source: https://www.skillbyai.com/en/mlops/g-tests

> More than unit tests.

## Four kinds of tests

ML CI runs: **code tests** (feature functions, preprocessing, API contracts); **data tests** (schema, ranges, null rates on fresh data); **model tests** (minimum quality on a fixed evaluation set, per-slice floors, invariance tests such as "changing the customer name does not change the prediction", directional tests such as "higher debt should not lower risk"); and **infrastructure tests** (the model loads, latency and memory within limits). Keep a small, fast suite on every commit and a full evaluation in the training pipeline.

## Example model tests (sketch)

Pytest-style checks; adapt to your model. Not run here.

```python
def test_minimum_quality(model, eval_set):
    assert model.score(eval_set.X, eval_set.y) >= 0.90

def test_slice_floor(model, eval_set):
    for name, (X, y) in eval_set.slices.items():
        assert model.score(X, y) >= 0.85, name

def test_invariance_to_name(model, row):
    a = model.predict_one({**row, "name": "Asha"})
    b = model.predict_one({**row, "name": "Rahul"})
    assert a == b

def test_latency(model, row):
    assert timed(model.predict_one, row) < 0.050   # seconds
```

## Turn incidents into tests

Every production failure should become a test case so it cannot silently return.

**Quiz:** What is an invariance test for a model?

- [ ] Checking the number of rows
- [ ] Checking that training finishes
- [ ] Checking the GPU type
- [x] Checking that an irrelevant change in input does not change the prediction

*Answer:* Checking that an irrelevant change in input does not change the prediction. Behavioural tests catch unwanted dependencies.
