पाठ 22 / 25

Back-of-the-Envelope Capacity Estimation

Estimate traffic, storage and bandwidth quickly for a design.

Rough numbers, quickly

Good designs start with rough numbers. Estimate daily active users and actions per user to get daily requests, divide by about 86,400 seconds (round to 100,000 for easy maths) for average requests per second, and multiply by a peak factor (2× to 10×, depending on how spiky traffic is) for peak load. Separate reads and writes, since read-heavy systems lean on caches and replicas while write-heavy systems need partitioning. Estimate storage as objects per day × size × retention, plus replication factor and indexes. Estimate bandwidth as requests × response size. Then compare with what one machine or database can do to see whether you need one server, ten or a thousand. Useful anchors: a well-tuned single relational database handles thousands to tens of thousands of simple queries per second; memory access is nanoseconds, a same-datacentre round trip is around half a millisecond, a cross-continent round trip is on the order of 100 to 150 ms.

From users to machines

Users and actions become requests per second, storage and bandwidth, and then a number of machines.

A funnel shape with small people icons at the top narrowing into a row of numbers and then into a small stack of server blocks.
Figure 8.1 — Estimation flows from users to infrastructure.

Estimating a photo-sharing feature

Show the working; reviewers care about the method more than the exact numbers.

assumptions
  10 M daily active users, each views 50 photos and uploads 0.2 photos per day
  average photo 400 KB after compression, kept forever, 3 copies
  peak factor 5x

reads   10 M x 50  = 500 M/day  -> ~5,800/s average -> ~29,000/s peak
writes  10 M x 0.2 = 2 M/day    -> ~23/s average    -> ~115/s peak

storage 2 M x 400 KB = 800 GB/day -> ~290 TB/year raw -> ~870 TB/year with 3 copies
egress  29,000/s x 400 KB = ~11.6 GB/s at peak -> must be served by a CDN

conclusion: read-heavy; CDN + object storage for images, cache metadata,
            a modest write path; storage growth is the main cost driver

State assumptions out loud

In interviews and design docs, write down every assumption before calculating. If an assumption is wrong, everyone can see which numbers change.

त्वरित जाँच: About how many requests per second is 864 million requests per day on average?

  • 100
  • 1,000
  • 10,000
  • 100,000
Answer

10,000 — 864,000,000 / 86,400 = 10,000 requests per second on average.