# Percentiles With histogram_quantile — Prometheus + Grafana

Source: https://www.skillbyai.com/en/prometheus-grafana/a-quantile

> p95 and p99 latency.

## Rate the buckets, keep le, then compute

To estimate a percentile, take the rate of the `_bucket` series, aggregate while **keeping the `le` label**, and pass the result to `histogram_quantile(0.95, ...)`. The result is interpolated within a bucket, so its accuracy depends on bucket boundaries. Average latency is `rate(_sum) / rate(_count)`. The share of requests faster than a threshold uses the bucket at that boundary divided by the total count.

## Percentiles, joins and precomputation

Histogram quantiles, vector matching and recording rules make complex queries fast and correct.

![Three ideas: percentiles, vector matching, recording rules.](assets/figures/prometheus-grafana/section-5-map.svg) — Figure 5.1 — Percentiles, vector matching and recording rules.

## Latency queries

PromQL.

```promql
# p95 latency per route over 5 minutes
histogram_quantile(0.95,
  sum by (route, le) (rate(http_request_duration_seconds_bucket{job="orders-api"}[5m])))

# average latency
sum(rate(http_request_duration_seconds_sum[5m])) / sum(rate(http_request_duration_seconds_count[5m]))

# fraction of requests faster than 300 ms (needs a 0.3 bucket)
sum(rate(http_request_duration_seconds_bucket{le="0.3"}[5m]))
  / sum(rate(http_request_duration_seconds_count[5m]))
```

## Never drop le before histogram_quantile

Aggregating without the le label destroys the bucket structure and produces meaningless results.

**Quiz:** Which label must be kept when aggregating buckets for histogram_quantile?

- [ ] instance
- [x] le
- [ ] job
- [ ] __name__

*Answer:* le. le identifies the bucket boundary.
