SkillByAIOpen interactive version →

Lesson 23 / 25

Discussing Trade-offs and Failure Modes

Show judgement, not just components.

Every choice has a cost

Strong candidates state the alternative they rejected and why: SQL versus NoSQL, push versus pull, strong versus eventual consistency, precompute versus compute on read, monolith versus services. Then walk through failure modes: what happens when a node, a dependency, a region or the network fails? Mention timeouts, retries with backoff and jitter, idempotency, circuit breakers, replication and failover, and graceful degradation (serve stale data, disable a feature). For LLD, the equivalent is edge cases and concurrency: invalid input, races on shared state, and how the design extends to a new requirement.

A failure-mode walk-through template

Apply to each critical component.

component: Payment Service
if it is slow     -> timeout at Xs (assumed), retry with same idempotency key, backoff + jitter
if it is down     -> orders stay PENDING, queue retries, show "processing" to user
if DB primary dies -> failover to replica; writes paused for failover window
if duplicated msg -> idempotency key dedupes
if data mismatch  -> nightly reconciliation flags for manual review
what we sacrifice -> slower confirmation during incidents, never double charge

A pilot pre-flight briefing

Pilots brief what they will do if an engine fails before it happens. Describing failure handling up front shows the same preparedness.

Quick check: Which statement best shows trade-off reasoning?

  • Choosing eventual consistency for the feed to gain availability, while keeping payments strongly consistent
  • Using the newest database because it is popular
  • Saying the system will never fail
  • Adding a cache everywhere without discussion
Answer

Choosing eventual consistency for the feed to gain availability, while keeping payments strongly consistent — Tie choices to requirements and costs.