# Case Study: Scaling a Web App Step by Step — Scalability, Availability & Reliability

Source: https://www.skillbyai.com/en/scalability/p-journey

> Follow a typical growth path and the change that each stage requires.

## Grow the architecture with the traffic

Systems rarely need their final architecture on day one. A common path: **Stage 1**, one server running the app and database; fine for a prototype. **Stage 2**, move the database to its own managed instance with backups and point-in-time recovery; add monitoring. **Stage 3**, put a load balancer in front of two or more stateless app instances across zones; sessions move to Redis; static files move to object storage and a CDN. **Stage 4**, add an application cache and database read replicas for read-heavy pages; move slow work to queues and workers. **Stage 5**, introduce SLOs, autoscaling, canary deployments and on-call. **Stage 6**, as write volume or data size outgrows one primary, partition data or split specific domains into separate services with their own stores. **Stage 7**, multi-region only if users or availability targets require it. At every stage, measure first and fix the current bottleneck rather than the imagined one.

## Bottleneck-driven growth

Each change is triggered by a measured symptom.

```text
symptom                                        change
---------------------------------------------  --------------------------------------------
app and DB fight for CPU on one box            separate managed database
one app server at 80% CPU, no redundancy       LB + 2-3 stateless instances across zones
same product pages read 1000x per minute       cache + CDN; read replicas
uploads and emails slow down requests          queue + workers
releases cause outages                         canaries, feature flags, SLO alerts
single primary saturated on writes             partition by customer / split a domain
users on another continent see 300 ms+ RTT     edge caching first, then a second region
```

## Extending a house

You add a room when the family grows, not before you are married. But you lay foundations strong enough for a second floor (stateless app servers, backups, monitoring) from the start, because they are hard to add later.

**Quiz:** Product pages are read thousands of times per minute and the primary database is busy. What is usually the first step?

- [ ] Shard the database by product ID
- [x] Add caching (and a CDN) and serve reads from replicas
- [ ] Move to multi-region active-active
- [ ] Rewrite the app in another language

*Answer:* Add caching (and a CDN) and serve reads from replicas. Read-heavy load is cheapest to address with caching and read replicas before partitioning.
