Lesson 24 / 25
RED and USE Methods
What to monitor.
Services and resources
Two methods give good coverage. RED for request-driven services: Rate, Errors, Duration. USE for resources such as CPU, memory, disks and networks: Utilisation, Saturation (queued work), Errors. Google's four golden signals (latency, traffic, errors, saturation) combine both views. Build a standard dashboard and alert set from these for every service, so on-call engineers know where to look.
USE queries with node_exporter
Host resource health.
# CPU utilisation
1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))
# CPU saturation: load per CPU
node_load5 / count by (instance) (node_cpu_seconds_total{mode="idle"})
# memory utilisation
1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes
# network errors
sum by (instance) (rate(node_network_receive_errs_total[5m]))Standardise across services
Shared metric names and a common dashboard template make every service equally easy to debug.
Quick check: What does the S in the USE method stand for?
- Success
- Speed
- Saturation
- Storage
Answer
Saturation — Utilisation, Saturation, Errors.