Implementing Service Resilience
1. Understanding Failure Modes
| Failure | Symptom |
|---|---|
| Crash | Process gone; restart |
| Slow / hang | Latency spike, threads blocked |
| Partial / Byzantine | Inconsistent behavior |
| Network partition | Split-brain risk |
| Resource exhaustion | OOM, file descriptors, threads |
2. Implementing Circuit Breaker Pattern
| State | Behavior |
|---|---|
| Closed | Calls pass through; track failures |
| Open | Fast-fail; no calls for cool-down |
| Half-Open | Probe small sample; close on success |
| Threshold | e.g. 50% failures over 20 calls |
| Tools | Resilience4j, Polly, gobreaker, Envoy |
3. Implementing Retry Pattern
| Element | Detail |
|---|---|
| Retry only on | Idempotent ops + transient errors |
| Backoff | Exponential with jitter |
| Cap attempts | 3–5 |
| Budget | Limit retries as % of total RPS |
| Avoid | Retry storms (combine with circuit breaker) |
4. Implementing Timeout Pattern
| Layer | Recommendation |
|---|---|
| Connect | 1–2s |
| Per-request | 1–10s based on SLA |
| Background job | Explicit per-job |
| Cascading | Caller TO > downstream TO |
| Cancel | Propagate context cancellation |
5. Implementing Bulkhead Pattern
| Type | Detail |
|---|---|
| Thread pool | Per-dependency executor |
| Semaphore | Cap concurrent calls |
| Connection pool | Per-upstream cap |
| Process / cluster | Workload isolation |
6. Implementing Graceful Degradation
| Strategy | Example |
|---|---|
| Fallback value | Cached recommendations on ML failure |
| Reduced feature set | Disable autocomplete; main search works |
| Stale-while-revalidate | Serve last good response |
| Read-only mode | Block writes during partial outage |
7. Using Failure Detection
| Detector | Detail |
|---|---|
| Health probes | K8s liveness/readiness |
| Outlier detection | Mesh ejects bad endpoints |
| Heartbeats | Periodic pings |
| Phi accrual | Adaptive failure detector (Akka, Cassandra) |
8. Handling Cascading Failures
| Defense | Detail |
|---|---|
| Timeouts | Free callers when downstream slow |
| Circuit breaker | Cut traffic to failing dep |
| Bulkheads | Isolate one slow dep from others |
| Load shedding | Reject low-priority work first |
| Backpressure | Signal upstream to slow down |
9. Implementing Self-Healing
| Mechanism | Detail |
|---|---|
| Auto-restart | K8s restarts on liveness fail |
| Auto-replace | Replace unhealthy nodes (ASG, MIG) |
| Auto-scale | HPA / KEDA on metrics |
| Auto-rollback | Argo Rollouts on bad metrics |
10. Managing Partial System Availability
| Tactic | Detail |
|---|---|
| Feature flags | Disable affected features |
| Region failover | Route traffic away from sick region |
| Quotas / fairness | Per-tenant isolation |
| Communicate | Status page, in-app banner |
11. Using Redundancy
| Type | Detail |
|---|---|
| N+1 instances | Survive single-instance loss |
| Multi-AZ | Spread replicas across zones |
| Multi-region | DR + locality |
| Active-active | Both serve traffic |
| Active-passive | Standby; failover on incident |
12. Using Resilience Libraries
| Library | Language |
|---|---|
| Resilience4j | Java (modern, modular) |
| Polly | .NET |
| gobreaker / failsafe-go | Go |
| opossum | Node.js |
| Tenacity | Python |
| Hystrix DEPRECATED | Use Resilience4j instead |