पाठ 23 / 25
Observing Resilience Patterns
Instrument breakers, retries, limiters and shedding so incidents are diagnosable.
Patterns must be visible
Resilience patterns change system behaviour in ways that confuse on-call engineers unless they are visible. Emit metrics for: timeouts per dependency; retry attempts and retry budget exhaustion; circuit breaker state (closed, open, half-open) and transitions; bulkhead rejections and concurrency in use; rate limiter allows and rejections per key or tier; load shedding rejections by reason and priority; and fallback usage per feature. Log breaker transitions and shedding decisions with context. Add the attempt number and fallback used as attributes on trace spans. Build a dashboard per service showing each dependency's latency, errors and breaker state side by side, and alert on sustained open breakers, rising shed rates and fallback usage for critical features. Watch for silent degradation: a fallback that has been serving cached prices for three days is an incident nobody noticed.
Resilience4j metrics in Prometheus queries
Spring Boot with Micrometer exports breaker and retry metrics automatically.
# breaker state per instance (1 for the current state)
resilience4j_circuitbreaker_state{name="payments", state="open"}
# calls by outcome over 5 minutes
sum by (kind) (rate(resilience4j_circuitbreaker_calls_seconds_count{name="payments"}[5m]))
# calls not permitted because the breaker was open
rate(resilience4j_circuitbreaker_not_permitted_calls_total{name="payments"}[5m])
# retries by outcome
sum by (kind) (rate(resilience4j_retry_calls_total{name="payments"}[5m]))Warning lights on a dashboard
A car that silently switches to limp mode without a warning light leaves you wondering why it is slow. Every resilience pattern that kicks in should light a warning on someone's dashboard.
त्वरित जाँच: Why should fallback usage be monitored?
- Fallbacks are always errors
- A fallback can hide a long-running failure, such as stale data being served for days
- Fallbacks increase CPU usage dramatically
- Fallbacks cannot be logged
Answer
A fallback can hide a long-running failure, such as stale data being served for days — Graceful degradation keeps users happy but must not let failures go unnoticed.