SkillByAIOpen interactive version →

Lesson 24 / 32

Logging, Metrics, Tracing & Alerting

Monitor systems with logs, metrics and traces, track latency, traffic, errors and saturation, and alert on user-facing symptoms.

The three pillars

Logs are timestamped discrete events (good for 'what exactly happened here'). Metrics are aggregated numbers over time (good for trends and dashboards — request rate, error rate, latency). Traces follow one request across every service it touched (good for finding where in a call chain the time went).

The golden signals

For any service, watch latency (p50/p95/p99, not just average), traffic (requests/sec), errors (rate of failed requests), and saturation (how full is CPU, memory, queue depth). These four catch most incidents early.

A car dashboard

Metrics are the speedometer and fuel gauge — glance and know your state. Logs are the trip computer's detailed event history. Tracing is like a black box recorder that shows the exact sequence of what each part of the car did leading up to a specific trip.

Alert on symptoms, not causes

Alert on user-facing symptoms (error rate spiked, p99 latency breached SLO) rather than every internal cause (one server's CPU is high). Symptom-based alerts reduce noise and point straight at what users are experiencing.