पाठ 24 / 25
Production Readiness Review
Check a service against a reliability checklist before launch.
A checklist before go-live
A production readiness review (PRR) asks a standard set of questions before a service takes real traffic, so reliability does not depend on memory. Typical areas: ownership and on-call (a team, a rota, runbooks); SLOs defined and dashboards built; alerting on symptoms and burn rate; capacity tested under load with headroom and autoscaling; redundancy across zones with no known SPOFs; dependencies listed, each with timeouts, retries, circuit breakers and a degradation plan; data with backups, tested restores, defined RPO and RTO; deployment with canaries, feature flags and practised rollback; security basics (secrets management, least privilege, patched images); and operations such as log retention, tracing, rate limits and cost alerts. A service that cannot answer these questions is not ready, however good the code is.
A condensed readiness checklist
Use as a pull-request template or launch gate.
[ ] owning team + on-call rota + runbook links
[ ] SLIs/SLOs defined; dashboard with golden signals
[ ] burn-rate alerts route to on-call; every alert actionable
[ ] load-tested at 2x expected peak; p99 within target; headroom >= 30%
[ ] instances spread across >= 2 zones; N+1 capacity
[ ] every remote call has a timeout; retries with backoff + jitter; circuit breakers
[ ] graceful degradation for non-critical dependencies
[ ] backups automated; restore tested in the last 90 days; RPO/RTO documented
[ ] canary or blue-green deploys; rollback < 10 minutes and rehearsed
[ ] feature flags / kill switches for risky features
[ ] rate limits and load shedding at the edge
[ ] cost and quota alerts configuredMake the checklist a habit, not a gate
Readiness reviews work best as a conversation early in the project, not a surprise exam the week before launch. Share the checklist when the design is written.
त्वरित जाँच: Which item is part of a typical production readiness review?
- Tested backup restores with documented RPO and RTO
- Choosing a code formatter
- Picking the team logo
- Deciding variable names
Answer
Tested backup restores with documented RPO and RTO — Readiness reviews focus on operability: SLOs, alerts, capacity, redundancy, backups and safe deploys.