पाठ 18 / 25
Automatic Rollback Triggers
Decide the response in advance.
Thresholds mapped to actions
Write rules that map metric breaches to actions: severe breaches (error spikes, safety flags, data leakage) automatically flip the kill switch; moderate ones (latency or cost overruns) page the on-call; softer quality signals pause the ramp and start a review. Automation matters because AI problems can spread quickly and incidents often happen outside working hours. Test the triggers, and make sure an automatic disable also notifies the owners.
Evaluating rollback rules on a metrics window, run
I ran this with Python 3 (scipy 1.18.1 where imported) on example numbers, not data from a real product. In this example window, the error rate (3.1%) exceeds its 2% limit, which is an auto-disable rule, and the thumbs-down rate exceeds 5%, which pauses the ramp. The decision is to flip the kill switch and notify the owners.
window = {"error_rate": 0.031, "p95_latency_ms": 4200, "thumbs_down_rate": 0.052,
"safety_flags_per_1k": 0.4, "cost_per_request": 0.0091}
rules = [ # (metric, limit, action)
("error_rate", 0.02, "auto-disable"), ("safety_flags_per_1k", 1.0, "auto-disable"),
("p95_latency_ms", 5000, "page on-call"), ("thumbs_down_rate", 0.05, "pause ramp + review"),
("cost_per_request", 0.012, "page on-call")]
triggered = [(m, window[m], lim, act) for m, lim, act in rules if window[m] > lim]
for m, v, lim, act in triggered:
print(f"{m:<20} {v} > {lim} -> {act}")
severity = "auto-disable" if any(a == "auto-disable" for *_, a in triggered) else "none"
print("decision:", "flip kill switch, notify owners" if severity == "auto-disable" else "continue")
Output:
error_rate 0.031 > 0.02 -> auto-disable thumbs_down_rate 0.052 > 0.05 -> pause ramp + review decision: flip kill switch, notify owners
Rehearse a rollback
Run a drill during the beta: trigger a rule on purpose and confirm that the feature turns off and the right people are told.
त्वरित जाँच: Which breach should usually trigger an automatic kill switch?
- A spike in errors or safety flags
- A small dip in a success metric for one hour
- A new feature request
- A positive review
Answer
A spike in errors or safety flags — Severe harm warrants automatic action.