Lesson 23 / 25

Storage, Retention and Scaling

Beyond a single server.

Local TSDB plus long-term storage

A single Prometheus stores data on local disk with a configurable retention (--storage.tsdb.retention.time, 15 days by default, or a size limit). For high availability, run two identical Prometheus servers and let Alertmanager deduplicate alerts. For long-term storage, a global view across clusters and multi-tenancy, use remote_write to systems such as Thanos, Grafana Mimir, Cortex or VictoriaMetrics, or managed services. Federation can pull aggregated series from many servers. Monitor Prometheus itself (scrape failures, rule evaluation time, memory).

Sending samples to long-term storage

remote_write in prometheus.yml (endpoint is illustrative).

remote_write:
  - url: https://mimir.example.com/api/v1/push
    basic_auth:
      username: "tenant-1"
      password_file: /etc/prometheus/remote_write_password
    write_relabel_configs:
      - source_labels: [__name__]
        regex: "go_.*|process_.*"
        action: drop          # keep long-term storage focused

Keep local retention short with remote storage

Local data serves recent queries and alerts; long-term questions go to the remote store.

Quick check: What is the standard way to keep alerts working if one Prometheus fails?

  • Run two identical Prometheus servers and let Alertmanager deduplicate
  • Increase the scrape interval
  • Store data in Grafana
  • Disable alerting rules
Answer

Run two identical Prometheus servers and let Alertmanager deduplicate — Simple HA by duplication.