पाठ 20 / 25
Incident Response and Postmortems
Run incidents calmly and learn from them with blameless postmortems.
A process for the worst day
Incidents go better with a known process. Declare an incident early, with a severity level. Assign roles: an incident commander who coordinates and decides, an operations lead who changes systems, and a communications lead who updates stakeholders and status pages. Prioritise mitigation over diagnosis: roll back the last deploy, fail over, shed load or disable a feature flag first, and investigate the root cause once users are no longer affected. Keep a timeline in a shared channel. Afterwards, write a blameless postmortem: what happened, impact, timeline, contributing factors (usually several), what went well, and action items with owners and due dates. Blameless does not mean nobody is responsible; it means asking why the system made the mistake easy, because people are rarely the root cause of systemic failures. Track metrics such as time to detect and time to mitigate.
A postmortem template
Keep it short, factual and focused on improving the system.
Title: Checkout errors after cache cluster upgrade - 2026-09-14
Severity: SEV-2 Duration: 38 min Impact: 7% of checkouts failed
Summary
A cache upgrade restarted all nodes at once; cold cache sent full load to the
orders database, which saturated its connection pool.
Timeline (IST)
14:02 upgrade started 14:05 error-rate page 14:11 incident declared
14:19 traffic shed for recommendations 14:40 cache warm, errors normal
Contributing factors
- upgrade procedure restarted nodes in parallel
- database sized only for warm-cache load
- no alert on cache hit ratio
Action items
- [owner A, Oct 1] rolling, one-node-at-a-time cache upgrades
- [owner B, Oct 8] load test database with cold cache; raise pool limit
- [owner C, Oct 8] alert on hit ratio < 80%Air crash investigations
Aviation became extraordinarily safe by investigating every incident to fix systems (checklists, cockpit design, training) rather than only blaming pilots. Blameless postmortems borrow that culture.
त्वरित जाँच: During a major incident caused shortly after a deploy, what should usually happen first?
- Find the exact line of code at fault
- Write the postmortem
- Wait for the next scheduled deploy
- Mitigate, for example by rolling back the deploy, then investigate
Answer
Mitigate, for example by rolling back the deploy, then investigate — Restoring service for users comes before full diagnosis.