# Designing Good Alerts — Prometheus + Grafana

Source: https://www.skillbyai.com/en/prometheus-grafana/al-design

> Symptoms, SLOs and burn rates.

## Page on user-facing symptoms

Page people for **symptoms users feel** (high error ratio, high latency, site unreachable), not for every possible cause (high CPU on one node). Every page should be urgent, actionable and real; send everything else to tickets or dashboards. Define **service level objectives** (for example 99.9% of requests succeed over 30 days) and alert on the **error budget burn rate**: a fast burn over short windows pages immediately, a slow burn over long windows opens a ticket. Review noisy alerts regularly and delete or tune them.

## A multi-window burn-rate alert

For a 99.9% SLO (error budget 0.1%); thresholds follow the widely used Google SRE workbook approach.

```promql
# assumes recording rules for the error ratio over 1h and 5m windows
# page: burning the 30-day budget about 14.4 times too fast
# (about 2% of the budget per hour), confirmed on a short window
(
  job:http_requests_error_ratio:rate1h{job="orders-api"} > (14.4 * 0.001)
 and
  job:http_requests_error_ratio:rate5m{job="orders-api"} > (14.4 * 0.001)
)
```

## Track alert quality

Count pages per week and the share that needed action; aim to remove alerts nobody acts on.

**Quiz:** Which is the better paging alert?

- [x] Checkout error ratio above the SLO burn threshold
- [ ] CPU above 70% on one node for 1 minute
- [ ] Disk 50% full
- [ ] A deployment started

*Answer:* Checkout error ratio above the SLO burn threshold. Page on symptoms users feel.
