पाठ 13 / 25

Bucket Aggregations: terms, date_histogram, range

Group documents.

Putting documents into buckets

Aggregations run over the documents matched by the query. Bucket aggregations group them: terms creates a bucket per distinct value of a keyword or numeric field (top 10 by count unless you set size), date_histogram creates time buckets using calendar_interval (month, week, ...) or fixed_interval (30m, 1d), and range/date_range use explicit boundaries. Buckets can contain sub-aggregations, such as revenue per brand per month. Set "size": 0 on the search when you only need aggregation results. terms counts on multi-shard indices can be approximate, because each shard returns its own top terms; the response reports doc_count_error_upper_bound, and raising shard_size improves accuracy. For paging through all buckets, use the composite aggregation.

Analytics alongside search

Aggregations group matching documents into buckets and compute metrics over them, in the same request as the search.

Three ideas: bucket aggregations, metric aggregations, faceted search.
Figure 5.1 — Matching documents grouped into buckets, then summarised by metrics.

Orders per month, per status, by price band

Kibana Dev Tools console syntax; send the same requests with curl or a client library.

GET /orders/_search
{
  "size": 0,
  "query": { "range": { "created_at": { "gte": "now-1y/d" } } },
  "aggs": {
    "per_month": {
      "date_histogram": { "field": "created_at", "calendar_interval": "month" },
      "aggs": {
        "by_status": { "terms": { "field": "status", "size": 5 } }
      }
    },
    "price_bands": {
      "range": {
        "field": "total",
        "ranges": [ { "to": 50 }, { "from": 50, "to": 200 }, { "from": 200 } ]
      }
    }
  }
}

Aggregate on keyword, not text

A terms aggregation on a text field fails by default (fielddata is disabled). Use the keyword field or a .keyword / .raw sub-field.

त्वरित जाँच: Why can terms aggregation counts be approximate?

  • Elasticsearch samples 1% of documents
  • Each shard returns only its own top terms before results are merged
  • Counts are rounded to tens
  • Replicas return stale counts
Answer

Each shard returns only its own top terms before results are merged — shard_size and doc_count_error_upper_bound address this.