पाठ 21 / 25
Design a Metrics and Logging Pipeline
Collect, buffer, store, query.
High-volume telemetry
Requirements: ingest metrics and logs from many hosts, query recent data quickly for dashboards and alerts, keep older data cheaply, and survive spikes. Pipeline: agents on each host collect and batch data; a durable log such as Kafka buffers ingest and decouples producers from consumers; stream processors parse, enrich and pre-aggregate; data lands in stores suited to each shape: a time-series database for metrics (series identified by name plus labels, compressed by time) and an indexed log store or columnar store for logs. Apply retention and downsampling (raw for days, rolled-up aggregates for longer) and move cold data to object storage. Watch cardinality: labels such as user id multiply the number of series. An alerting service evaluates rules on recent data.
Pipeline and back-of-envelope
All numbers are assumptions for illustration.
Assume 50k hosts x 500 series x 1 sample / 10s -> 2.5M samples/s
Assume ~2 bytes/sample after compression (TSDBs compress heavily; check yours)
-> ~5 MB/s -> ~430 GB/day raw metrics
agents (batch, compress) -> load balancer -> ingest gateways
-> Kafka topics (partitioned by series hash / service)
-> stream jobs: parse, drop/limit high-cardinality labels, 1m rollups
-> TSDB (hot, 15 days raw) + object storage (rollups, 1 year)
-> log store (indexed, 7 days) + object storage (archive)
query layer -> dashboards; rule evaluator -> alert manager -> on-callShed load gracefully
Under overload, drop or sample debug logs before metrics used for alerting; tell the interviewer which data you would sacrifice first.
त्वरित जाँच: Why can adding a user id label to a metric be dangerous?
- Labels are not supported by time-series databases
- It can explode the number of time series (high cardinality)
- It makes metrics immutable
- It disables compression entirely
Answer
It can explode the number of time series (high cardinality) — Each unique label combination is a new series.