Lesson 25 / 25
Revision and Interview Questions
Recall resilience and rate-limiting concepts for exams and interviews.
Cheat sheet
Why: failures are normal; slowness causes cascades; retries can amplify. Deadlines: every request has one; inner timeouts shorter than outer; propagate remaining budget. Timeouts: connect and request timeouts, from p99/p99.9; pool acquisition timeouts. Retries: only transient errors on idempotent operations; exponential backoff + jitter; cap attempts; honour Retry-After; one layer only; retry budgets; hedging for idempotent reads. Circuit breaker: closed → open (failure or slow-call rate over a window with a minimum number of calls) → half-open trial → closed/open; per dependency; client errors do not count; Resilience4j, Polly. Bulkheads: per-dependency concurrency or pools; cells. Fallbacks: cached, default, reduced feature, deferred; critical vs optional dependencies; kill switches and brownouts. Rate limiting: keys, 429 + Retry-After, fixed window, sliding log, sliding window counter, token bucket, leaky bucket; distributed with atomic Redis scripts; fail open vs closed. Overload: load shedding by priority and queue time; adaptive concurrency limits; back-pressure; client-side adaptive throttling (requests − K × accepts). Platform: Envoy/Istio timeouts, retries, outlier detection, connection limits; idempotency keys. Practice: fault injection, chaos, game days; metrics for every pattern.
Common interview questions
Answer each with mechanism, parameters and a trade-off.
1. Explain the three states of a circuit breaker and what moves it between them.
2. Which errors would you retry, and how do backoff and jitter help?
3. How do retries cause amplification across layers, and how do you prevent it?
4. What is a bulkhead? Give two ways to implement one.
5. Compare fixed window, sliding window and token bucket rate limiting.
6. How would you implement a rate limiter shared by 20 API instances?
7. What is load shedding, and why does it improve goodput under overload?
8. How do idempotency keys make POST requests safe to retry?
9. How should timeouts relate across a gateway, a service and a database?
10. Design resilience for a checkout flow with payments and fraud scoring.Name the numbers
Interviewers like concrete parameters: "300 ms timeout from a p99 of 180 ms, one retry with full jitter, breaker at 50% failures over 20 calls, 30 s open" shows you have configured these for real.
Quick check: Which combination best protects a caller from a dependency that suddenly becomes very slow?
- Longer timeouts and more retries
- Removing all timeouts
- Caching the dependency's errors forever
- Short timeouts, a circuit breaker with slow-call detection and a bulkhead
Answer
Short timeouts, a circuit breaker with slow-call detection and a bulkhead — Timeouts bound waiting, the breaker stops calls to a slow dependency and the bulkhead caps resource use.