Simple, hard limit, single point of failure, expensive
Horizontal (Scale Out)
Add more nodes behind LB
Linear cost, complex coordination, requires statelessness
Diagonal
Vertical until limit, then horizontal
Pragmatic hybrid, common in practice
Functional Decomposition
Split by feature/service
Microservices, independent scaling
Data Partitioning
Shard data across nodes
Scales storage + writes; cross-shard joins hard
Note: Scalability is measured by throughput (RPS), latency at p95/p99, and concurrent users. Always define scaling target before designing.
2. Understanding Availability Requirements
SLA Tier
Uptime %
Downtime/Year
Downtime/Month
Two nines
99%
3.65 days
7.2 hours
Three nines
99.9%
8.77 hours
43.8 min
Four nines
99.99%
52.6 min
4.38 min
Five nines
99.999%
5.26 min
26.3 sec
Example: Composite availability of dependent services
Service A (99.9%) → Service B (99.9%) → Service C (99.9%)Composite = 0.999 × 0.999 × 0.999 = 0.997 (99.7%)Mitigation: redundancy, async fallback, circuit breakers
3. Understanding Reliability Patterns
Pattern
Purpose
Example
Redundancy
Eliminate SPOFs
N+1 servers, multi-AZ DB
Replication
Data durability + read scaling
3-way replicated storage
Failover
Automatic switch on failure
Active-passive DB, VRRP
Health checks
Detect failure fast
K8s liveness/readiness
Retries + backoff
Handle transient errors
Exponential backoff + jitter
Circuit breaker
Stop cascading failure
Resilience4j, Hystrix
Bulkhead
Isolate failure domains
Per-tenant thread pools
4. Understanding Consistency Models
Model
Guarantee
Use Case
Strong (Linearizable)
All reads see latest write immediately
Banking, inventory
Sequential
All clients see same order
Distributed locks
Causal
Causally related ops in order
Comments, social feeds
Read-your-writes
Client sees own writes
User profile updates
Monotonic Reads
No going back in time
Timeline views
Eventual
Replicas converge eventually
DNS, S3, Cassandra
5. Understanding CAP Theorem
Choice
Sacrifice
Examples
CP
Availability during partition
HBase, MongoDB (default), ZooKeeper, etcd
AP
Strong consistency
Cassandra, DynamoDB, Riak, CouchDB
CA
Partition tolerance (only single-node)
RDBMS on single node
Warning: Network partitions are inevitable in distributed systems. CA is not realistic — you must choose CP or AP during partitions.
6. Understanding PACELC Theorem
System
Partition (P)
Else (E)
DynamoDB
PA
EL (low latency over consistency)
Cassandra
PA
EL
MongoDB
PC
EC
Spanner
PC
EC (TrueTime keeps latency low)
CockroachDB
PC
EC
Note: PACELC extends CAP: if Partition then A or C; else if no partition, choose Latency or Consistency.
7. Understanding Latency vs Throughput Trade-offs
Metric
Definition
Optimization
Latency
Time per single request (ms)
Caching, CDN, reduce hops, faster algorithms
Throughput
Requests per second (RPS)
Batching, parallelism, async I/O
p50/p95/p99
Percentile latency
Reduce tail: hedged requests, isolation
Little's Law
L = λ × W (concurrency = arrival × wait)
Reduce W or limit L for predictability
8. Understanding Fault Tolerance Principles
Principle
Implementation
Fail fast
Short timeouts, reject early
Fail safe
Default to safe state on failure
Graceful degradation
Reduced functionality vs total outage
Self-healing
Auto-restart, auto-replace nodes
Idempotency
Safe retries
Isolation
Bulkheads prevent cascade
9. Understanding Data Redundancy and Replication
Strategy
Consistency
Use Case
Synchronous replication
Strong
Financial txns; higher write latency
Asynchronous replication
Eventual
Read replicas; risk of data loss on failover
Semi-sync
Bounded staleness
Wait for ≥1 replica ack
Multi-leader
Eventual + conflicts
Multi-region writes
Leaderless (quorum)
Tunable (R+W>N)
Cassandra, Dynamo
10. Understanding Single Points of Failure (SPOF)
SPOF
Mitigation
Single DB instance
Primary + replicas, multi-AZ
Single load balancer
Pair with VRRP/keepalived; managed LB
Single region
Multi-region active-active or active-passive
DNS provider
Multi-DNS (Route53 + NS1)
Shared cache
Cluster mode + replication
Cert authority
Multiple CAs, automated rotation
11. Understanding Trade-offs in System Design
Trade-off
Pick A
Pick B
Consistency vs Availability
RDBMS, banking
Social feed, catalog
Latency vs Throughput
Real-time API
Batch ETL
Read-heavy vs Write-heavy
Cache, replicas
Sharding, LSM stores
SQL vs NoSQL
ACID, joins, schema
Scale, flexible schema
Sync vs Async
Immediate response
Decoupling, resilience
Build vs Buy
Core differentiator
Commodity infra
12. Understanding SLA and SLO Design
Term
Meaning
Example
SLI
Indicator (measurement)
Request success rate, p99 latency
SLO
Internal objective
99.95% success/30d, p99 < 200ms
SLA
External contract w/ penalties
99.9% uptime or refund
Error Budget
1 - SLO; budget for risk
0.05% downtime/month allowed
Note: SLA < SLO < SLI accuracy. Buffer SLA below SLO. Halt feature releases when error budget burns too fast.