# Incident Response and Postmortems — Scalability, Availability & Reliability

Source: https://www.skillbyai.com/en/scalability/o-incidents

> Run incidents calmly and learn from them with blameless postmortems.

## A process for the worst day

Incidents go better with a known process. Declare an incident early, with a **severity** level. Assign roles: an **incident commander** who coordinates and decides, an **operations lead** who changes systems, and a **communications lead** who updates stakeholders and status pages. Prioritise **mitigation over diagnosis**: roll back the last deploy, fail over, shed load or disable a feature flag first, and investigate the root cause once users are no longer affected. Keep a timeline in a shared channel. Afterwards, write a **blameless postmortem**: what happened, impact, timeline, contributing factors (usually several), what went well, and **action items** with owners and due dates. Blameless does not mean nobody is responsible; it means asking why the system made the mistake easy, because people are rarely the root cause of systemic failures. Track metrics such as time to detect and time to mitigate.

## A postmortem template

Keep it short, factual and focused on improving the system.

```text
Title: Checkout errors after cache cluster upgrade - 2026-09-14
Severity: SEV-2   Duration: 38 min   Impact: 7% of checkouts failed

Summary
  A cache upgrade restarted all nodes at once; cold cache sent full load to the
  orders database, which saturated its connection pool.

Timeline (IST)
  14:02 upgrade started   14:05 error-rate page   14:11 incident declared
  14:19 traffic shed for recommendations   14:40 cache warm, errors normal

Contributing factors
  - upgrade procedure restarted nodes in parallel
  - database sized only for warm-cache load
  - no alert on cache hit ratio

Action items
  - [owner A, Oct 1] rolling, one-node-at-a-time cache upgrades
  - [owner B, Oct 8] load test database with cold cache; raise pool limit
  - [owner C, Oct 8] alert on hit ratio < 80%
```

## Air crash investigations

Aviation became extraordinarily safe by investigating every incident to fix systems (checklists, cockpit design, training) rather than only blaming pilots. Blameless postmortems borrow that culture.

**Quiz:** During a major incident caused shortly after a deploy, what should usually happen first?

- [ ] Find the exact line of code at fault
- [ ] Write the postmortem
- [ ] Wait for the next scheduled deploy
- [x] Mitigate, for example by rolling back the deploy, then investigate

*Answer:* Mitigate, for example by rolling back the deploy, then investigate. Restoring service for users comes before full diagnosis.
