Designing Disaster Recovery and Business Continuity
1. Designing RTO and RPO Strategy
| Term | Definition |
|---|---|
| RTO | Max acceptable downtime to recover |
| RPO | Max acceptable data loss |
| Drives | Backup frequency, replication mode, infra cost |
| Per-tier | Different RTO/RPO per service criticality |
2. Designing Backup Strategy
| Type | Detail |
|---|---|
| Full | Complete copy; weekly |
| Incremental | Changes since last backup |
| Differential | Changes since last full |
| 3-2-1 rule | 3 copies, 2 media, 1 offsite |
| Immutable | WORM / object-lock for ransomware defense |
3. Designing Multi-Region Failover
| Element | Detail |
|---|---|
| Trigger | Manual approval recommended |
| DNS | Route 53 / Cloud DNS health-based |
| Data promote | Replica → primary (one-way without care) |
| Verify | Smoke tests post-failover |
4. Designing Database Backup and Restore
| Practice | Detail |
|---|---|
| Snapshots | Block-level, fast |
| Logical (pg_dump) | Cross-version restore |
| WAL archiving | Continuous; PITR |
| Test restores | Regularly; backup is only as good as restore |
| Encryption | Backups encrypted at rest |
5. Designing Disaster Recovery Testing
| Test | Detail |
|---|---|
| Tabletop | Walk-through scenario |
| Partial | Restore one service |
| Full DR drill | End-to-end region failover |
| Frequency | Quarterly minimum |
| Document | Lessons → update runbooks |
6. Designing Business Continuity Planning
| Element | Detail |
|---|---|
| BIA | Business impact analysis per system |
| Critical functions | Prioritize recovery order |
| Communication plan | Internal + customer + regulator |
| Vendor contingency | Backup providers |
7. Designing Data Replication for DR
| Mechanism | Detail |
|---|---|
| Async DB replication | Cross-region; lag in seconds |
| Log shipping | WAL/binlog to remote |
| Object storage replication | S3 CRR / Cloud Storage replication |
| Snapshot copy | Periodic cross-region |
8. Designing Recovery Automation
| Tool | Detail |
|---|---|
| Runbook automation | Rundeck, Ansible, AWS SSM |
| IaC | Terraform spin-up secondary |
| GitOps | Argo CD reconciles new cluster |
| One-click failover | Encapsulated workflow |
9. Designing Point-in-Time Recovery
| Element | Detail |
|---|---|
| Base backup + WAL | Replay to any timestamp |
| Granularity | Seconds to minutes typical |
| Use cases | Accidental delete, bad migration |
| Retention | e.g., 7–35 days WAL |
10. Designing Cross-Region Data Synchronization
| Pattern | Detail |
|---|---|
| Single-leader async | Simple; lag |
| Multi-leader | Conflict resolution required |
| Globally distributed DB | Spanner, CockroachDB, Yugabyte |
| Event-bus replication | MirrorMaker for Kafka |
11. Designing Disaster Recovery Runbooks
| Section | Detail |
|---|---|
| Trigger criteria | When to invoke |
| Roles | Incident commander, comms, ops |
| Step-by-step | Commands with expected output |
| Rollback | How to undo each step |
| Comms templates | Status page, email, Slack |
12. Designing Failback Procedures
| Step | Detail |
|---|---|
| Reverse replication | DR → original primary catch-up |
| Verify parity | Compare row counts / checksums |
| Plan downtime | Or read-only window |
| Switch traffic | DNS / GSLB |
| Postmortem | Capture learnings |