Lesson 23 / 25
Storage, Retention and Scaling
Beyond a single server.
Local TSDB plus long-term storage
A single Prometheus stores data on local disk with a configurable retention (--storage.tsdb.retention.time, 15 days by default, or a size limit). For high availability, run two identical Prometheus servers and let Alertmanager deduplicate alerts. For long-term storage, a global view across clusters and multi-tenancy, use remote_write to systems such as Thanos, Grafana Mimir, Cortex or VictoriaMetrics, or managed services. Federation can pull aggregated series from many servers. Monitor Prometheus itself (scrape failures, rule evaluation time, memory).
Sending samples to long-term storage
remote_write in prometheus.yml (endpoint is illustrative).
remote_write:
- url: https://mimir.example.com/api/v1/push
basic_auth:
username: "tenant-1"
password_file: /etc/prometheus/remote_write_password
write_relabel_configs:
- source_labels: [__name__]
regex: "go_.*|process_.*"
action: drop # keep long-term storage focusedKeep local retention short with remote storage
Local data serves recent queries and alerts; long-term questions go to the remote store.
Quick check: What is the standard way to keep alerts working if one Prometheus fails?
- Run two identical Prometheus servers and let Alertmanager deduplicate
- Increase the scrape interval
- Store data in Grafana
- Disable alerting rules
Answer
Run two identical Prometheus servers and let Alertmanager deduplicate — Simple HA by duplication.