Implementing Scalability Patterns
1. Scaling Horizontally
| Aspect | Detail |
|---|---|
| Add instances | Same app, more replicas |
| Requires | Stateless service or sticky session externalized |
| Tool | K8s Deployment + HPA |
2. Scaling Vertically
| Aspect | Detail |
|---|---|
| Bigger node | More CPU / RAM |
| Use | Stateful (DB primary), JVM heap headroom |
| Limit | Hardware ceiling, downtime to resize |
3. Implementing Stateless Services
| Rule | Detail |
|---|---|
| No local state | Use Redis / DB |
| Idempotent ops | Safe retry |
| Config externalized | Env / config service |
4. Using Database Sharding
| Strategy | Detail |
|---|---|
| Hash | Even distribution |
| Range | Range queries; risk of hot shards |
| Directory | Tenant → shard map |
| Resharding | Painful; consistent hash helps |
5. Implementing Read Replicas
| Aspect | Detail |
|---|---|
| Routing | Reads → replicas; writes → primary |
| Lag | Eventual consistency window |
| Read-your-writes | Route to primary briefly |
6. Using Event Sourcing
| Benefit | Detail |
|---|---|
| Append-only writes | Highly scalable |
| Multiple read models | Materialized views per use |
| Replayable | Rebuild projections |
7. Implementing CQRS
| Aspect | Detail |
|---|---|
| Separate models | Scale read/write independently |
| Read store choice | Per query (Elastic, Redis) |
8. Using Asynchronous Processing
| Aspect | Detail |
|---|---|
| Queue absorbs spikes | Smooth load on downstream |
| Workers scale | HPA on queue depth |
9. Implementing Auto-Scaling
| Type | Signal |
|---|---|
| HPA | CPU, mem, custom (RPS, queue) |
| VPA | Adjust requests/limits |
| Cluster Autoscaler / Karpenter | Add nodes |
| KEDA | Scale on events (Kafka, SQS) |
10. Managing Hot Partitions
| Tactic | Detail |
|---|---|
| Better key | Add suffix / salt |
| Split partition | DynamoDB adaptive capacity |
| Cache hot keys | In-memory write-back |
11. Using Database Connection Pooling
| Aspect | Detail |
|---|---|
| App pool | HikariCP per pod |
| Proxy pool | pgbouncer (transaction mode) |
| Total conns | pods × maxPool ≤ DB max |
12. Implementing Queue-Based Load Leveling
| Element | Detail |
|---|---|
| Producer | Always fast (write to queue) |
| Consumer | Drains at safe rate |
| Backpressure | Reject 503 if queue full |