# Production Readiness Review — Scalability, Availability & Reliability

Source: https://www.skillbyai.com/en/scalability/p-readiness

> Check a service against a reliability checklist before launch.

## A checklist before go-live

A **production readiness review (PRR)** asks a standard set of questions before a service takes real traffic, so reliability does not depend on memory. Typical areas: **ownership and on-call** (a team, a rota, runbooks); **SLOs** defined and dashboards built; **alerting** on symptoms and burn rate; **capacity** tested under load with headroom and autoscaling; **redundancy** across zones with no known SPOFs; **dependencies** listed, each with timeouts, retries, circuit breakers and a degradation plan; **data** with backups, tested restores, defined RPO and RTO; **deployment** with canaries, feature flags and practised rollback; **security** basics (secrets management, least privilege, patched images); and **operations** such as log retention, tracing, rate limits and cost alerts. A service that cannot answer these questions is not ready, however good the code is.

## A condensed readiness checklist

Use as a pull-request template or launch gate.

```text
[ ] owning team + on-call rota + runbook links
[ ] SLIs/SLOs defined; dashboard with golden signals
[ ] burn-rate alerts route to on-call; every alert actionable
[ ] load-tested at 2x expected peak; p99 within target; headroom >= 30%
[ ] instances spread across >= 2 zones; N+1 capacity
[ ] every remote call has a timeout; retries with backoff + jitter; circuit breakers
[ ] graceful degradation for non-critical dependencies
[ ] backups automated; restore tested in the last 90 days; RPO/RTO documented
[ ] canary or blue-green deploys; rollback < 10 minutes and rehearsed
[ ] feature flags / kill switches for risky features
[ ] rate limits and load shedding at the edge
[ ] cost and quota alerts configured
```

## Make the checklist a habit, not a gate

Readiness reviews work best as a conversation early in the project, not a surprise exam the week before launch. Share the checklist when the design is written.

**Quiz:** Which item is part of a typical production readiness review?

- [x] Tested backup restores with documented RPO and RTO
- [ ] Choosing a code formatter
- [ ] Picking the team logo
- [ ] Deciding variable names

*Answer:* Tested backup restores with documented RPO and RTO. Readiness reviews focus on operability: SLOs, alerts, capacity, redundancy, backups and safe deploys.
