पाठ 14 / 25

Metric Aggregations: avg, percentiles, cardinality

Summarise numbers.

Exact and approximate metrics

Metric aggregations compute values over a set of documents: avg, sum, min, max, value_count, and stats for several at once. Some metrics are approximate by design to stay fast and memory-bounded at scale: percentiles uses the TDigest algorithm (very accurate at the extremes like p99, less so in the middle on huge datasets), and cardinality (distinct count) uses HyperLogLog++, exact up to roughly the configurable precision_threshold and approximate beyond it. Metrics can sit inside bucket aggregations, and pipeline aggregations such as bucket_sort or derivative work on the output of other aggregations.

Latency and unique users per endpoint

Kibana Dev Tools console syntax; send the same requests with curl or a client library.

GET /requests/_search
{
  "size": 0,
  "aggs": {
    "by_endpoint": {
      "terms": { "field": "endpoint", "size": 20 },
      "aggs": {
        "avg_ms":        { "avg": { "field": "duration_ms" } },
        "latency":       { "percentiles": { "field": "duration_ms", "percents": [ 50, 95, 99 ] } },
        "unique_users":  { "cardinality": { "field": "user_id", "precision_threshold": 3000 } }
      }
    }
  }
}

Do not report approximate as exact

If finance needs an exact distinct count, compute it in the system of record. Cardinality from Elasticsearch is ideal for dashboards and trends, not invoices.

त्वरित जाँच: What does the cardinality aggregation return?

  • The number of shards
  • The exact number of documents
  • The largest value in a field
  • An approximate count of distinct values
Answer

An approximate count of distinct values — It uses HyperLogLog++ with a precision_threshold.