Designing for High Availability
1. Designing Multi-Region Architecture
| Element | Detail |
|---|---|
| Stateless tier | Replicate per region |
| Data tier | Async global repl or per-region partitioned |
| Routing | Geo / latency DNS, anycast |
| Conflict handling | Last-writer-wins / CRDT / per-region ownership |
| Cost | Cross-region egress, duplicated infra |
2. Designing Multi-AZ Architecture
| Practice | Detail |
|---|---|
| Spread replicas | ≥3 AZs for quorum services |
| Topology spread | K8s topologySpreadConstraints |
| DB Multi-AZ | Sync standby (RDS, Aurora) |
| LB | Cross-zone enabled |
3. Designing Disaster Recovery Strategy
| Tier | RTO / RPO |
|---|---|
| Backup & Restore | Hours / hours |
| Pilot Light | ~1h / minutes |
| Warm Standby | Minutes / seconds |
| Multi-Site Active | Seconds / near-zero |
4. Designing Active-Active Configuration
| Aspect | Detail |
|---|---|
| Both regions serve | Load split via DNS / GSLB |
| Conflict resolution | Region affinity per partition |
| Pros | Capacity + resilience |
| Cons | Complex consistency |
5. Designing Active-Passive Configuration
| Aspect | Detail |
|---|---|
| Primary serves | Secondary on standby |
| Failover | DNS swap or Route53 health check |
| Pros | Simpler consistency |
| Cons | Idle capacity, slower failover |
6. Designing Failover Mechanisms
| Type | Detail |
|---|---|
| Automatic | Health check triggered |
| Manual | Operator-approved |
| DNS-based | TTL impacts time |
| Anycast IP | Near-instant |
| Split-brain prevention | Fencing, quorum |
7. Designing Database Replication for HA
| Mode | Detail |
|---|---|
| Sync | Strong durability; higher latency |
| Async | Lag risk; better throughput |
| Semi-sync | Wait for one replica ack |
| Auto-failover | Patroni, Orchestrator, Aurora |
8. Designing Zero-Downtime Deployments
| Element | Detail |
|---|---|
| Rolling updates | K8s default |
| Backwards-compatible APIs | Tolerate N/N+1 mix |
| DB migrations | Expand → migrate → contract |
| Connection draining | Honor terminationGracePeriod |
| Health probes | Don't route until ready |
9. Designing Health Monitoring
| Signal | Detail |
|---|---|
| RED | Rate, Errors, Duration |
| USE | Utilization, Saturation, Errors |
| Synthetic probes | External user-flow checks |
| SLO burn rate | Multi-window alerts |
10. Designing Service Dependency Management
| Practice | Detail |
|---|---|
| Dependency map | Track upstream/downstream |
| Critical path analysis | What's needed for P0 flows |
| Soft dependencies | Degrade gracefully if down |
| Vendor SLAs | Compose to your effective SLA |
11. Designing Chaos Engineering Practices
| Experiment | Goal |
|---|---|
| Kill pod | Verify replicas + health |
| Network partition | Validate timeouts/CB |
| AZ down | Cross-AZ failover |
| DB failover | Reconnection logic |
| Latency injection | Tail latency handling |
12. Designing SLA and Uptime Targets
| Uptime | Allowed downtime / year |
|---|---|
| 99% (2 nines) | 3.65 days |
| 99.9% (3 nines) | 8.76 hours |
| 99.95% | 4.38 hours |
| 99.99% (4 nines) | 52.6 minutes |
| 99.999% (5 nines) | 5.26 minutes |