Lesson 24 / 25

RED and USE Methods

What to monitor.

Services and resources

Two methods give good coverage. RED for request-driven services: Rate, Errors, Duration. USE for resources such as CPU, memory, disks and networks: Utilisation, Saturation (queued work), Errors. Google's four golden signals (latency, traffic, errors, saturation) combine both views. Build a standard dashboard and alert set from these for every service, so on-call engineers know where to look.

USE queries with node_exporter

Host resource health.

# CPU utilisation
1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))

# CPU saturation: load per CPU
node_load5 / count by (instance) (node_cpu_seconds_total{mode="idle"})

# memory utilisation
1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes

# network errors
sum by (instance) (rate(node_network_receive_errs_total[5m]))

Standardise across services

Shared metric names and a common dashboard template make every service equally easy to debug.

Quick check: What does the S in the USE method stand for?

  • Success
  • Speed
  • Saturation
  • Storage
Answer

Saturation — Utilisation, Saturation, Errors.