# Storage, Retention and Scaling — Prometheus + Grafana

Source: https://www.skillbyai.com/en/prometheus-grafana/o-scale

> Beyond a single server.

## Local TSDB plus long-term storage

A single Prometheus stores data on local disk with a configurable retention (`--storage.tsdb.retention.time`, 15 days by default, or a size limit). For high availability, run two identical Prometheus servers and let Alertmanager deduplicate alerts. For long-term storage, a global view across clusters and multi-tenancy, use **remote_write** to systems such as Thanos, Grafana Mimir, Cortex or VictoriaMetrics, or managed services. Federation can pull aggregated series from many servers. Monitor Prometheus itself (scrape failures, rule evaluation time, memory).

## Sending samples to long-term storage

remote_write in prometheus.yml (endpoint is illustrative).

```yaml
remote_write:
  - url: https://mimir.example.com/api/v1/push
    basic_auth:
      username: "tenant-1"
      password_file: /etc/prometheus/remote_write_password
    write_relabel_configs:
      - source_labels: [__name__]
        regex: "go_.*|process_.*"
        action: drop          # keep long-term storage focused
```

## Keep local retention short with remote storage

Local data serves recent queries and alerts; long-term questions go to the remote store.

**Quiz:** What is the standard way to keep alerts working if one Prometheus fails?

- [x] Run two identical Prometheus servers and let Alertmanager deduplicate
- [ ] Increase the scrape interval
- [ ] Store data in Grafana
- [ ] Disable alerting rules

*Answer:* Run two identical Prometheus servers and let Alertmanager deduplicate. Simple HA by duplication.
