Designing for Scalability
1. Implementing Horizontal Scaling
| Requirement | Solution |
| Statelessness | Move session to Redis/JWT |
| Service discovery | Consul, K8s DNS, Eureka |
| Load balancing | L4/L7 LB, round-robin, least-conn |
| Shared storage | S3, EFS, distributed DB |
| Config | Env vars, central config store |
2. Implementing Vertical Scaling
| Resource | Limit | When to Use |
| CPU/RAM | Single-node ceiling | Low-traffic, simple ops |
| SSD/NVMe IOPS | Disk bandwidth | DB write workloads |
| NIC throughput | 10/25/100GbE | Network-bound services |
| Cost curve | Super-linear at top tier | Switch to horizontal |
3. Designing Auto-Scaling Strategies
| Strategy | Trigger | Use Case |
| Reactive (target tracking) | CPU/RPS over threshold | General workloads |
| Step scaling | Different sizes per band | Bursty traffic |
| Scheduled | Time-of-day | Predictable patterns |
| Predictive | ML forecast | Holiday spikes |
| Queue-based | Queue depth per worker | Async workers |
Note: Always set min/max bounds; scale-out faster than scale-in (cooldowns) to avoid flapping.
4. Designing Stateless vs Stateful Services
| Aspect | Stateless | Stateful |
| Scaling | Trivially horizontal | Requires data partitioning |
| Failover | Replace instance | Replicate state, leader election |
| Examples | API servers, web frontends | Databases, game servers, caches |
| K8s primitive | Deployment | StatefulSet, PVC |
5. Implementing Service Decomposition
| Strategy | Split By |
| Business capability | Order, Inventory, Billing |
| Subdomain (DDD) | Bounded contexts |
| Verb/use case | Checkout service, Search service |
| Data ownership | One service owns each entity |
| Volatility | Isolate frequently-changing parts |
6. Implementing Connection Pooling at Scale
| Setting | Guideline |
| Pool size | ≈ (CPU × 2) + effective spindles (Postgres rule) |
| Max DB connections | App pool × instances ≤ DB max |
| Connection proxy | PgBouncer, RDS Proxy for high fan-out |
| Acquire timeout | 200–500ms; fail fast |
| Idle eviction | 5–10 min |
7. Managing Distributed State
| Approach | Description |
| Externalize | Move state to Redis/DB |
| Sticky sessions | LB routes user to same node |
| Consistent hashing | Stable mapping under node changes |
| Leader election | Single owner per resource (etcd, ZK) |
| CRDTs | Conflict-free merging across replicas |
8. Designing Scalability Testing Strategy
| Test | Goal |
| Load | Verify SLO at expected peak |
| Stress | Find breaking point |
| Spike | Sudden 10× traffic |
| Soak | Detect leaks over hours/days |
| Tools | k6, Gatling, JMeter, Locust |
9. Identifying Scalability Bottlenecks
| Layer | Symptom | Tool |
| CPU | High %util, runqueue depth | top, perf, async-profiler |
| Memory | GC pauses, OOM | jstat, pprof |
| Disk I/O | iowait, queue depth | iostat, fio |
| Network | Retransmits, full TX queue | ss, tcpdump |
| DB | Slow queries, lock waits | pg_stat_activity, EXPLAIN |
| Application | Lock contention, thread pool | flame graphs, APM |
10. Designing Capacity Planning Strategy
Capacity Planning Process
- Forecast demand (RPS, GB/day) over horizon
- Benchmark single-instance throughput
- Compute required instances + headroom (≥30%)
- Validate via load test at projected scale
- Re-evaluate quarterly with actuals
11. Designing Sharding Strategy
| Strategy | Pros | Cons |
| Range | Range queries efficient | Hot spots on sequential keys |
| Hash | Even distribution | Range queries scatter |
| Consistent hash | Min reshuffling on resize | Slightly uneven; needs vnodes |
| Geo | Locality, compliance | Cross-region joins |
| Directory | Flexible mapping | Lookup is SPOF |
12. Designing Partitioning Strategy
| Type | Description |
| Horizontal | Rows split across partitions (sharding) |
| Vertical | Columns split (rare; column stores) |
| Functional | Tables split by feature |
| Hybrid | Composite key hashing + range |
Warning: Choose partition key carefully — changing it later requires full data migration. Avoid keys with hot partitions (e.g., timestamp as primary).