पाठ 4 / 25
Traffic and Storage Estimates
QPS, storage, cache.
Round aggressively, show your work
Convert daily volumes into queries per second (a day is about 86,400 seconds, roughly 10^5), multiply by a peak factor (2 to 5 times), and apply the read/write ratio. Storage is items per day times size times retention, then times the replication factor. The 80/20 rule suggests caching the hottest fraction of data. Aim for the right order of magnitude, not precision.
Numbers that drive design
Quick estimates reveal whether you need one server or a thousand, and where the bottleneck is.
Estimating a URL shortener, run
I ran this with Python 3.12.3 using only the standard library; inputs are fixed or seeded, so the output is reproducible. Ten million new links per day is only about 116 writes per second but over 11,000 reads per second; five years of links need about 9 TB before replication, and a cache for the hottest 20% of a day's links fits in about 1 GB.
# Back-of-envelope: a URL shortener
DAU_WRITES = 10_000_000 # new short links per day (assumption)
READ_WRITE_RATIO = 100 # each link is read ~100 times
BYTES_PER_LINK = 500 # long URL + short code + metadata
YEARS = 5
SECONDS_PER_DAY = 86_400
write_qps = DAU_WRITES / SECONDS_PER_DAY
read_qps = write_qps * READ_WRITE_RATIO
links = DAU_WRITES * 365 * YEARS
storage_tb = links * BYTES_PER_LINK / 1e12
print(f"write QPS avg : {write_qps:,.0f} peak (x3): {write_qps * 3:,.0f}")
print(f"read QPS avg : {read_qps:,.0f} peak (x3): {read_qps * 3:,.0f}")
print(f"links in {YEARS}y : {links:,}")
print(f"storage : {storage_tb:.2f} TB (before replication)")
print(f"x3 replicas : {storage_tb * 3:.2f} TB")
hot_gb = DAU_WRITES * 0.2 * BYTES_PER_LINK / 1e9 # 80/20: cache the hottest 20%
print(f"cache for 20% of a day's links: {hot_gb:.1f} GB")
Output:
write QPS avg : 116 peak (x3): 347 read QPS avg : 11,574 peak (x3): 34,722 links in 5y : 18,250,000,000 storage : 9.12 TB (before replication) x3 replicas : 27.38 TB cache for 20% of a day's links: 1.0 GB
Let numbers change the design
Say what the estimate implies: "11k reads/s fits a cache cluster easily; 9 TB means we need to partition the database."
त्वरित जाँच: Roughly how many seconds are in a day for quick estimates?
- About 1,000
- About 100,000 (86,400)
- About 10 million
- About 3,600
Answer
About 100,000 (86,400) — 86,400 rounds to 10^5.