# Traffic and Storage Estimates — System Design Interview Prep

Source: https://www.skillbyai.com/en/system-design-interview/e-traffic

> QPS, storage, cache.

## Round aggressively, show your work

Convert daily volumes into **queries per second** (a day is about 86,400 seconds, roughly 10^5), multiply by a **peak factor** (2 to 5 times), and apply the read/write ratio. Storage is items per day times size times retention, then times the replication factor. The 80/20 rule suggests caching the hottest fraction of data. Aim for the right order of magnitude, not precision.

## Numbers that drive design

Quick estimates reveal whether you need one server or a thousand, and where the bottleneck is.

![Three ideas: traffic and storage, latency numbers, availability maths.](assets/figures/system-design-interview/section-2-map.svg) — Figure 2.1 — Traffic, latency and availability.

## Estimating a URL shortener, run

I ran this with Python 3.12.3 using only the standard library; inputs are fixed or seeded, so the output is reproducible. Ten million new links per day is only about 116 writes per second but over 11,000 reads per second; five years of links need about 9 TB before replication, and a cache for the hottest 20% of a day's links fits in about 1 GB.

```python
# Back-of-envelope: a URL shortener
DAU_WRITES = 10_000_000          # new short links per day (assumption)
READ_WRITE_RATIO = 100           # each link is read ~100 times
BYTES_PER_LINK = 500             # long URL + short code + metadata
YEARS = 5
SECONDS_PER_DAY = 86_400

write_qps = DAU_WRITES / SECONDS_PER_DAY
read_qps = write_qps * READ_WRITE_RATIO
links = DAU_WRITES * 365 * YEARS
storage_tb = links * BYTES_PER_LINK / 1e12

print(f"write QPS avg : {write_qps:,.0f}   peak (x3): {write_qps * 3:,.0f}")
print(f"read QPS avg  : {read_qps:,.0f}  peak (x3): {read_qps * 3:,.0f}")
print(f"links in {YEARS}y : {links:,}")
print(f"storage       : {storage_tb:.2f} TB (before replication)")
print(f"x3 replicas   : {storage_tb * 3:.2f} TB")
hot_gb = DAU_WRITES * 0.2 * BYTES_PER_LINK / 1e9  # 80/20: cache the hottest 20%
print(f"cache for 20% of a day's links: {hot_gb:.1f} GB")
```

Output:

```
write QPS avg : 116   peak (x3): 347
read QPS avg  : 11,574  peak (x3): 34,722
links in 5y : 18,250,000,000
storage       : 9.12 TB (before replication)
x3 replicas   : 27.38 TB
cache for 20% of a day's links: 1.0 GB
```

## Let numbers change the design

Say what the estimate implies: "11k reads/s fits a cache cluster easily; 9 TB means we need to partition the database."

**Quiz:** Roughly how many seconds are in a day for quick estimates?

- [ ] About 1,000
- [x] About 100,000 (86,400)
- [ ] About 10 million
- [ ] About 3,600

*Answer:* About 100,000 (86,400). 86,400 rounds to 10^5.
