# Design a Metrics and Logging Pipeline — High-Level & Low-Level Design Interview Problems

Source: https://www.skillbyai.com/en/design-interviews/metrics-pipeline

> Collect, buffer, store, query.

## High-volume telemetry

Requirements: ingest metrics and logs from many hosts, query recent data quickly for dashboards and alerts, keep older data cheaply, and survive spikes. Pipeline: **agents** on each host collect and batch data; a **durable log** such as Kafka buffers ingest and decouples producers from consumers; **stream processors** parse, enrich and pre-aggregate; data lands in stores suited to each shape: a **time-series database** for metrics (series identified by name plus labels, compressed by time) and an **indexed log store** or columnar store for logs. Apply **retention and downsampling** (raw for days, rolled-up aggregates for longer) and move cold data to object storage. Watch **cardinality**: labels such as user id multiply the number of series. An alerting service evaluates rules on recent data.

## Pipeline and back-of-envelope

All numbers are assumptions for illustration.

```text
Assume 50k hosts x 500 series x 1 sample / 10s  ->  2.5M samples/s
Assume ~2 bytes/sample after compression (TSDBs compress heavily; check yours)
  -> ~5 MB/s -> ~430 GB/day raw metrics

agents (batch, compress) -> load balancer -> ingest gateways
   -> Kafka topics (partitioned by series hash / service)
   -> stream jobs: parse, drop/limit high-cardinality labels, 1m rollups
   -> TSDB (hot, 15 days raw) + object storage (rollups, 1 year)
   -> log store (indexed, 7 days) + object storage (archive)
query layer -> dashboards; rule evaluator -> alert manager -> on-call
```

## Shed load gracefully

Under overload, drop or sample debug logs before metrics used for alerting; tell the interviewer which data you would sacrifice first.

**Quiz:** Why can adding a user id label to a metric be dangerous?

- [ ] Labels are not supported by time-series databases
- [x] It can explode the number of time series (high cardinality)
- [ ] It makes metrics immutable
- [ ] It disables compression entirely

*Answer:* It can explode the number of time series (high cardinality). Each unique label combination is a new series.
