Implementing Disaster Recovery
1. Designing Backup Strategies
| Element | Detail |
|---|---|
| 3-2-1 rule | 3 copies, 2 media, 1 off-site |
| Frequency | Continuous (WAL/CDC) + daily snapshot |
| Encrypt | At rest + in transit |
| Test restore | Quarterly; backup w/o test = no backup |
2. Implementing Recovery Point Objective
| RPO | Strategy |
|---|---|
| 0 | Synchronous replication |
| Seconds | WAL streaming, CDC |
| Minutes | Periodic snapshots |
| Hours/days | Daily backups (low value data) |
3. Implementing Recovery Time Objective
| RTO | Pattern |
|---|---|
| Seconds | Active-active multi-region |
| Minutes | Hot standby, automated failover |
| Hours | Pilot light (minimal warm) |
| Days | Backup & restore |
4. Using Multi-Region Deployment
| Topology | Detail |
|---|---|
| Active-passive | Standby region; failover on disaster |
| Active-active | Both serve traffic; geo-routing |
| Data | Spanner / Aurora Global / Cassandra |
| Caveats | Cross-region latency; conflict resolution |
5. Implementing Failover Mechanisms
| Element | Detail |
|---|---|
| DNS failover | Route 53 health checks |
| Anycast | BGP withdrawal on failure |
| DB promotion | Promote replica to primary |
| Avoid split-brain | Quorum / fencing |
6. Testing Disaster Recovery
| Drill | Detail |
|---|---|
| Tabletop | Walk-through scenarios |
| Game day | Live partial outage |
| Full DR | Cut traffic to DR region |
| Cadence | Quarterly minimum |
7. Managing Data Replication
| Mode | Detail |
|---|---|
| Synchronous | RPO=0; latency cost |
| Async | Lag possible; data loss risk |
| Semi-sync | ≥1 replica acks |
| Logical (CDC) | Selective tables across versions |
8. Implementing Circuit Breakers
| Use in DR | Detail |
|---|---|
| Fail fast | Don't pile up on dead region |
| Bypass | Route to fallback region |
| Recovery | Half-open probes |
9. Using Chaos Engineering
| Experiment | Detail |
|---|---|
| Kill pods / nodes | Verify reschedule |
| AZ / region failure | Test failover |
| DB failover | Validate apps reconnect |
| Tools | Chaos Mesh, Litmus, Gremlin, AWS FIS |
10. Maintaining Disaster Recovery Plans
| Doc Section | Detail |
|---|---|
| Scope | Services + RPO/RTO |
| Roles | Incident commander, comms, ops |
| Runbooks | Per scenario |
| Comms | Status page + customer comms templates |
| Review | After every drill / incident |
11. Implementing Active-Active vs Active-Passive
| Aspect | Active-Active |
|---|---|
| Cost | Higher (2× capacity in use) |
| RTO | Seconds |
| Complexity | Conflict resolution, global state |
| Active-Passive | Cheaper, RTO minutes, simpler |
12. Using DR Automation Tools
| Tool | Use |
|---|---|
| AWS Elastic DR / DRS | VM replication |
| Velero | K8s backup/restore |
| Terraform | Recreate infra |
| Argo CD | Restore workloads from Git |