Lesson 5 / 25
Histograms and Summaries
Distributions such as latency.
Buckets for percentiles
Averages hide slow requests, so latency is recorded as a distribution. A histogram counts observations into cumulative buckets (le labels such as 0.1, 0.25, 0.5 seconds) and exposes _bucket, _sum and _count series; percentiles are estimated at query time with histogram_quantile, and buckets can be aggregated across instances. A summary calculates quantiles inside the client, which cannot be meaningfully aggregated across instances. Prefer histograms, and choose buckets around your latency targets. Newer Prometheus versions also support native histograms with automatic buckets.
Histogram series for one route
Illustrative exposition output.
http_request_duration_seconds_bucket{route="/api/orders",le="0.1"} 9200
http_request_duration_seconds_bucket{route="/api/orders",le="0.25"} 10100
http_request_duration_seconds_bucket{route="/api/orders",le="0.5"} 10380
http_request_duration_seconds_bucket{route="/api/orders",le="1"} 10440
http_request_duration_seconds_bucket{route="/api/orders",le="+Inf"} 10449
http_request_duration_seconds_sum{route="/api/orders"} 1187.4
http_request_duration_seconds_count{route="/api/orders"} 10449Put a bucket at your SLO threshold
If the target is 300 ms, a 0.3 bucket lets you count good requests exactly.
Quick check: Why prefer histograms over summaries for multi-instance services?
- Summaries cannot record latency
- Histogram buckets can be aggregated across instances before computing percentiles
- Histograms need no buckets
- Summaries are not supported by Prometheus
Answer
Histogram buckets can be aggregated across instances before computing percentiles — You cannot average percentiles.