Designing Rate Limiting and Throttling
1. Designing Rate Limiting Algorithms
| Algorithm | Pros | Cons |
|---|---|---|
| Fixed window | Simple, low memory | Edge bursts at boundaries |
| Sliding log | Accurate | Memory per request |
| Sliding window counter | Good balance | Approximate |
| Token bucket | Allows bursts | Tune burst + rate |
| Leaky bucket | Smooths output | No burst |
2. Designing Fixed Window Rate Limiting
Example: Redis fixed window
key = "rl:{user}:{minute}"
INCR key
EXPIRE key 60 NX
if value > LIMIT then deny
3. Designing Sliding Window Rate Limiting
| Variant | Detail |
|---|---|
| Sliding log | ZSET of timestamps; trim old |
| Weighted counter | previous*overlap + current |
| Memory | Counter ≪ log |
4. Designing Distributed Rate Limiting
| Approach | Detail |
|---|---|
| Central Redis | Atomic INCR / Lua script |
| Local + sync | Each node holds quota slice |
| Gossip | Eventually consistent counters |
| Envoy global ratelimit | gRPC service |
| Consistency vs latency | Trade-off explicitly |
5. Designing Rate Limit Response Headers
| Header | Meaning |
|---|---|
| X-RateLimit-Limit | Quota |
| X-RateLimit-Remaining | Remaining |
| X-RateLimit-Reset | Epoch when reset |
| Retry-After | Seconds (with 429) |
| RFC 9239 standard | RateLimit-Policy / RateLimit fields |
6. Designing Per-User Rate Limiting
| Strategy | Detail |
|---|---|
| Key by user_id | Authenticated users |
| Tier-based | Free vs paid limits |
| Multiple windows | req/sec, req/min, req/day combined |
7. Designing Per-IP Rate Limiting
| Concern | Detail |
|---|---|
| NAT / shared IPs | Set higher limit; combine with cookie |
| IPv6 | Limit by /64 prefix |
| Trusted proxy | Use X-Forwarded-For correctly |
8. Designing API Quotas and Tiers
| Tier | Limits |
|---|---|
| Free | 60 req/min, 10K/day |
| Pro | 1K req/min, 1M/day |
| Enterprise | Custom, SLA-backed |
| Overage | Pay-per-use vs hard cap |
9. Designing Graceful Degradation Under Load
| Strategy | Detail |
|---|---|
| Shed non-critical | Disable analytics, recommendations |
| Serve stale | From cache |
| Reduce features | Read-only mode |
| Queue + 202 | Defer heavy work |
10. Designing Backpressure Mechanisms
| Mechanism | Detail |
|---|---|
| Bounded queues | Reject when full (fail fast) |
| Reactive streams | Demand-driven |
| Concurrency limits | Semaphore per dependency |
| Adaptive concurrency (AIMD) | Netflix concurrency-limits |
11. Designing Adaptive Rate Limiting
| Signal | Action |
|---|---|
| Latency rising | Tighten limits |
| Error rate up | Throttle harder |
| CPU / queue depth | Scale or shed |
| Per-user anomaly | Step-down that user |
12. Designing Rate Limit Bypass Strategies
| Use Case | Mechanism |
|---|---|
| Internal services | Allow-list by mTLS identity |
| VIP customers | Higher tier or no limit |
| Health checks | Excluded paths |
| Audit | Always log bypass usage |