पाठ 25 / 25
A Monitoring Checklist
Review before relying on it.
Questions to ask
Does every service expose RED metrics with bounded labels? Are latency histograms bucketed around the SLO? Are scrape targets discovered automatically and up monitored? Are heavy queries recorded? Do paging alerts focus on user-facing symptoms with runbooks, and are lower-priority alerts routed to tickets? Are dashboards provisioned from version control? Is cardinality tracked and limited? Are retention, high availability and long-term storage planned, and is Prometheus itself monitored?
The checklist
Use it in reviews.
[ ] RED metrics on every service; route templates, no IDs in labels
[ ] histograms with buckets around SLO thresholds
[ ] targets via service discovery; alert on up == 0 and scrape errors
[ ] rate() before sum(); windows >= 4x scrape interval
[ ] recording rules for dashboard and alert queries
[ ] paging alerts on symptoms / SLO burn, with runbook links
[ ] Alertmanager grouping, inhibition, routing tested
[ ] Grafana dashboards provisioned from git; variables for reuse
[ ] cardinality monitored; sample_limit set
[ ] retention, HA pair, remote_write for long-term dataTest alerts deliberately
Use promtool test rules and occasional game days to confirm alerts fire and reach the right people.
त्वरित जाँच: Which item belongs on a monitoring checklist?
- Sum counters before rate()
- Use user IDs as labels
- Paging alerts have runbook links
- Edit dashboards only by hand in production
Answer
Paging alerts have runbook links — Actionable, maintainable monitoring.