Lesson 4 / 25

Counters and Gauges

Things that go up, things that go up and down.

Pick the type by behaviour

A counter only increases (or resets to zero when the process restarts): requests served, errors, bytes sent. You almost never graph a raw counter; you graph its rate. A gauge goes up and down: memory in use, queue length, temperature, active connections. Do not use a gauge for something you count, because you lose increments between scrapes and restarts.

Measure the right thing

Choosing the right metric type and instrumenting the key paths makes later queries and alerts simple.

Three ideas: counters and gauges, histograms and summaries, client libraries.
Figure 2.1 — Counters, histograms and instrumentation.

Examples by type

Common choices.

counter  http_requests_total, jobs_failed_total, bytes_sent_total
         -> query with rate() / increase()
gauge    queue_depth, memory_used_bytes, active_sessions, last_success_timestamp_seconds
         -> query directly, or with avg_over_time()/max_over_time()

Record timestamps as gauges

A last_success_timestamp_seconds gauge lets you alert when a batch job has not succeeded recently: time() - metric > threshold.

Quick check: Which type fits "number of jobs waiting in a queue"?

  • Counter
  • Gauge
  • Histogram only
  • Summary only
Answer

Gauge — It goes up and down.