# Bulkheads — Rate Limiting, Circuit Breakers & Resilience Patterns

Source: https://www.skillbyai.com/en/resilience-patterns/b-bulkheads

> Isolate resources per dependency so one failure cannot consume everything.

## Watertight compartments

Ships are divided into watertight compartments (bulkheads) so a hole floods one section, not the whole hull. In software, a **bulkhead** gives each dependency or workload its **own limited pool of resources**: a separate thread pool, a semaphore limiting concurrent calls, or a separate connection pool. If the recommendations service hangs, it can occupy at most its 20 permitted concurrent calls; checkout's calls to payments still have their own capacity. Two common forms: **semaphore bulkheads** limit concurrent in-flight calls on the caller's threads (cheap, works well with async code); **thread-pool bulkheads** run calls on a dedicated pool with a bounded queue (stronger isolation, more overhead). Bulkheads also apply at larger scales: separate **instance pools** for internal and public traffic, separate **clusters** per tenant tier, or **cell-based architecture**, where each cell serves a subset of customers so a failure affects only that cell.

## Compartments for each dependency

Each dependency gets its own slice of capacity, so one flood stays contained.

![A ship hull shape divided into four compartments, one filled with dark water and the others dry.](assets/figures/resilience-patterns/section-4-map.svg) — Figure 4.1 — Bulkheads keep one failing dependency from sinking the service.

## A semaphore bulkhead per dependency

When the limit is reached, calls are rejected immediately instead of piling up.

```python
import asyncio

class BulkheadFull(Exception): pass

class Bulkhead:
    def __init__(self, max_concurrent):
        self.sem = asyncio.Semaphore(max_concurrent)

    async def run(self, coro_fn):
        if self.sem.locked():
            raise BulkheadFull()          # reject rather than queue unboundedly
        async with self.sem:
            return await coro_fn()

payments_bulkhead = Bulkhead(max_concurrent=50)
recs_bulkhead = Bulkhead(max_concurrent=20)

async def get_recommendations(user_id):
    try:
        return await recs_bulkhead.run(lambda: recs_client.get(user_id, timeout=0.3))
    except (BulkheadFull, TimeoutError):
        return []                          # degrade gracefully
```

## Size bulkheads with Little's law

Concurrency needed ≈ request rate × latency. For 200 calls per second at 100 ms, about 20 concurrent calls are normal; a limit of 40 leaves headroom while still capping damage if latency explodes.

**Quiz:** What does a bulkhead limit?

- [ ] The number of retries
- [ ] The size of HTTP responses
- [x] The share of resources (such as concurrent calls) one dependency or workload can consume
- [ ] The cache TTL

*Answer:* The share of resources (such as concurrent calls) one dependency or workload can consume. Bulkheads cap resource usage per dependency so failures stay contained.
