पाठ 22 / 25

Anomaly Detection

Find the rare and unusual.

Scores, thresholds and review

Anomaly detection flags points that look unlike the rest: fraudulent orders, faulty sensors, unusual logins. Methods such as Isolation Forest score how easily a point can be isolated by random splits; rare, extreme points isolate quickly. You must choose how many to flag (a contamination rate or score threshold), and some flags will be normal-but-unusual cases. Treat anomalies as candidates for review, measure precision against labelled incidents when you have them, and tune the threshold to the review capacity.

Isolation Forest on orders, run

I ran this with Python 3, numpy 2.5.3 and scikit-learn 1.9.1, using fixed random seeds. Among 500 normal orders plus three odd ones, an Isolation Forest set to flag about 1% marks six rows: the three injected odd orders (rows 500 to 502) and three unusual but normal orders. A human review step would sort them out.

import numpy as np
from sklearn.ensemble import IsolationForest
rng = np.random.default_rng(0)
normal = rng.normal(loc=[50, 2], scale=[10, 0.5], size=(500, 2))     # amount, items per order
odd = np.array([[400, 1], [5, 30], [300, 25]])
X = np.vstack([normal, odd])
flags = IsolationForest(contamination=0.01, random_state=0).fit_predict(X)
print("flagged rows:", np.where(flags == -1)[0].tolist())
print("the three injected odd orders are rows 500, 501, 502")

Output:

flagged rows: [151, 206, 239, 500, 501, 502]
the three injected odd orders are rows 500, 501, 502

Match flags to review capacity

If analysts can check 50 cases a day, set the threshold so roughly 50 cases a day are flagged.

त्वरित जाँच: Why should anomaly flags usually be reviewed?

  • Review makes training faster
  • Anomaly detectors are always perfect
  • Flags are random
  • Some flagged points are unusual but legitimate
Answer

Some flagged points are unusual but legitimate — Unusual is not the same as bad.