Lesson 18 / 25

Designing Good Alerts

Symptoms, SLOs and burn rates.

Page on user-facing symptoms

Page people for symptoms users feel (high error ratio, high latency, site unreachable), not for every possible cause (high CPU on one node). Every page should be urgent, actionable and real; send everything else to tickets or dashboards. Define service level objectives (for example 99.9% of requests succeed over 30 days) and alert on the error budget burn rate: a fast burn over short windows pages immediately, a slow burn over long windows opens a ticket. Review noisy alerts regularly and delete or tune them.

A multi-window burn-rate alert

For a 99.9% SLO (error budget 0.1%); thresholds follow the widely used Google SRE workbook approach.

# assumes recording rules for the error ratio over 1h and 5m windows
# page: burning the 30-day budget about 14.4 times too fast
# (about 2% of the budget per hour), confirmed on a short window
(
  job:http_requests_error_ratio:rate1h{job="orders-api"} > (14.4 * 0.001)
 and
  job:http_requests_error_ratio:rate5m{job="orders-api"} > (14.4 * 0.001)
)

Track alert quality

Count pages per week and the share that needed action; aim to remove alerts nobody acts on.

Quick check: Which is the better paging alert?

  • Checkout error ratio above the SLO burn threshold
  • CPU above 70% on one node for 1 minute
  • Disk 50% full
  • A deployment started
Answer

Checkout error ratio above the SLO burn threshold — Page on symptoms users feel.