Lesson 10 / 25
Bulkheads
Isolate resources per dependency so one failure cannot consume everything.
Watertight compartments
Ships are divided into watertight compartments (bulkheads) so a hole floods one section, not the whole hull. In software, a bulkhead gives each dependency or workload its own limited pool of resources: a separate thread pool, a semaphore limiting concurrent calls, or a separate connection pool. If the recommendations service hangs, it can occupy at most its 20 permitted concurrent calls; checkout's calls to payments still have their own capacity. Two common forms: semaphore bulkheads limit concurrent in-flight calls on the caller's threads (cheap, works well with async code); thread-pool bulkheads run calls on a dedicated pool with a bounded queue (stronger isolation, more overhead). Bulkheads also apply at larger scales: separate instance pools for internal and public traffic, separate clusters per tenant tier, or cell-based architecture, where each cell serves a subset of customers so a failure affects only that cell.
Compartments for each dependency
Each dependency gets its own slice of capacity, so one flood stays contained.
A semaphore bulkhead per dependency
When the limit is reached, calls are rejected immediately instead of piling up.
import asyncio
class BulkheadFull(Exception): pass
class Bulkhead:
def __init__(self, max_concurrent):
self.sem = asyncio.Semaphore(max_concurrent)
async def run(self, coro_fn):
if self.sem.locked():
raise BulkheadFull() # reject rather than queue unboundedly
async with self.sem:
return await coro_fn()
payments_bulkhead = Bulkhead(max_concurrent=50)
recs_bulkhead = Bulkhead(max_concurrent=20)
async def get_recommendations(user_id):
try:
return await recs_bulkhead.run(lambda: recs_client.get(user_id, timeout=0.3))
except (BulkheadFull, TimeoutError):
return [] # degrade gracefullySize bulkheads with Little's law
Concurrency needed ≈ request rate × latency. For 200 calls per second at 100 ms, about 20 concurrent calls are normal; a limit of 40 leaves headroom while still capping damage if latency explodes.
Quick check: What does a bulkhead limit?
- The number of retries
- The size of HTTP responses
- The share of resources (such as concurrent calls) one dependency or workload can consume
- The cache TTL
Answer
The share of resources (such as concurrent calls) one dependency or workload can consume — Bulkheads cap resource usage per dependency so failures stay contained.