पाठ 25 / 25
Revision and Interview Questions
Recall the key ideas quickly for exams and system design interviews.
Cheat sheet
Definitions: scalability (handle growth), availability (fraction of time serving), reliability (correct over time), durability (data survives); latency vs throughput. Targets: SLI → SLO → SLA; error budget = 1 − SLO; 99.9% ≈ 43 min/month, 99.99% ≈ 4.3 min/month. Models: Little's law L = λW; Amdahl caps parallel speed-up; queueing latency explodes near 100% utilisation, so keep headroom. Compute: scale up vs out; stateless services; L4/L7 load balancing, least-requests, power of two choices; health checks; autoscaling on the right metric with cooldowns. Data: replication (sync/async, lag, read-your-writes); partitioning (range, hash, consistent hashing, hot keys); CAP, PACELC, quorums R + W > N. Performance: caching layers and stampedes; queues, back-pressure, batching; percentiles, fan-out tail, hedging. Availability: remove SPOFs, N+1, active-active vs active-passive, split brain; zones vs regions, RTO/RPO, DR tiers; series multiply, parallel = 1 − (1 − a)^n. Reliability: slow/partial/cascading failures, retry storms; timeouts, backoff + jitter, circuit breakers, bulkheads, load shedding, degradation; canaries and rollbacks. Ops: golden signals, burn-rate alerts, incident roles, blameless postmortems, load tests, chaos engineering.
Common interview questions
Answer with a definition, an example and a trade-off.
1. What is the difference between availability and reliability?
2. How much downtime does a 99.95% monthly SLO allow? What is an error budget for?
3. Vertical vs horizontal scaling: when would you choose each?
4. How do you make a web tier stateless?
5. Explain replica lag and how to give users read-your-writes consistency.
6. Range vs hash partitioning; what is consistent hashing for?
7. Explain CAP and PACELC with an example of each choice.
8. Calculate availability for three 99.9% components in series, and two in parallel.
9. What causes cascading failures, and which patterns prevent them?
10. Choose a DR strategy for RTO = 15 minutes and RPO = 1 minute, and justify it.Always quantify
"We add replicas" is weak; "reads are 95% of 30,000 peak requests per second, so three replicas plus a cache with a 90% hit rate keep each replica under 1,000 queries per second" shows you can reason about scale.
त्वरित जाँच: Two independent replicas each have 99% availability, and either can serve requests. What is the combined availability?
- 98%
- 99%
- 99.5%
- 99.99%
Answer
99.99% — 1 − (1 − 0.99)² = 1 − 0.0001 = 0.9999.