Lesson 20 / 25

Incident Response and Postmortems

Run incidents calmly and learn from them with blameless postmortems.

A process for the worst day

Incidents go better with a known process. Declare an incident early, with a severity level. Assign roles: an incident commander who coordinates and decides, an operations lead who changes systems, and a communications lead who updates stakeholders and status pages. Prioritise mitigation over diagnosis: roll back the last deploy, fail over, shed load or disable a feature flag first, and investigate the root cause once users are no longer affected. Keep a timeline in a shared channel. Afterwards, write a blameless postmortem: what happened, impact, timeline, contributing factors (usually several), what went well, and action items with owners and due dates. Blameless does not mean nobody is responsible; it means asking why the system made the mistake easy, because people are rarely the root cause of systemic failures. Track metrics such as time to detect and time to mitigate.

A postmortem template

Keep it short, factual and focused on improving the system.

Title: Checkout errors after cache cluster upgrade - 2026-09-14
Severity: SEV-2   Duration: 38 min   Impact: 7% of checkouts failed

Summary
  A cache upgrade restarted all nodes at once; cold cache sent full load to the
  orders database, which saturated its connection pool.

Timeline (IST)
  14:02 upgrade started   14:05 error-rate page   14:11 incident declared
  14:19 traffic shed for recommendations   14:40 cache warm, errors normal

Contributing factors
  - upgrade procedure restarted nodes in parallel
  - database sized only for warm-cache load
  - no alert on cache hit ratio

Action items
  - [owner A, Oct 1] rolling, one-node-at-a-time cache upgrades
  - [owner B, Oct 8] load test database with cold cache; raise pool limit
  - [owner C, Oct 8] alert on hit ratio < 80%

Air crash investigations

Aviation became extraordinarily safe by investigating every incident to fix systems (checklists, cockpit design, training) rather than only blaming pilots. Blameless postmortems borrow that culture.

Quick check: During a major incident caused shortly after a deploy, what should usually happen first?

  • Find the exact line of code at fault
  • Write the postmortem
  • Wait for the next scheduled deploy
  • Mitigate, for example by rolling back the deploy, then investigate
Answer

Mitigate, for example by rolling back the deploy, then investigate — Restoring service for users comes before full diagnosis.