पाठ 11 / 25

Testing ML Code, Data and Models

More than unit tests.

Four kinds of tests

ML CI runs: code tests (feature functions, preprocessing, API contracts); data tests (schema, ranges, null rates on fresh data); model tests (minimum quality on a fixed evaluation set, per-slice floors, invariance tests such as "changing the customer name does not change the prediction", directional tests such as "higher debt should not lower risk"); and infrastructure tests (the model loads, latency and memory within limits). Keep a small, fast suite on every commit and a full evaluation in the training pipeline.

Example model tests (sketch)

Pytest-style checks; adapt to your model. Not run here.

def test_minimum_quality(model, eval_set):
    assert model.score(eval_set.X, eval_set.y) >= 0.90

def test_slice_floor(model, eval_set):
    for name, (X, y) in eval_set.slices.items():
        assert model.score(X, y) >= 0.85, name

def test_invariance_to_name(model, row):
    a = model.predict_one({**row, "name": "Asha"})
    b = model.predict_one({**row, "name": "Rahul"})
    assert a == b

def test_latency(model, row):
    assert timed(model.predict_one, row) < 0.050   # seconds

Turn incidents into tests

Every production failure should become a test case so it cannot silently return.

त्वरित जाँच: What is an invariance test for a model?

  • Checking the number of rows
  • Checking that training finishes
  • Checking the GPU type
  • Checking that an irrelevant change in input does not change the prediction
Answer

Checking that an irrelevant change in input does not change the prediction — Behavioural tests catch unwanted dependencies.