Implementing Monitoring and Alerting Patterns
1. Synthetic Monitoring Pattern
| Aspect | Detail |
| Definition | Scripted requests run on schedule from external locations |
| Tools | Pingdom, Datadog Synthetics, Checkly, Grafana k6 |
| Use | Detect outage before users report; SLO measurement |
| Cadence | Every 1-5 minutes per critical journey |
2. Real User Monitoring Pattern
| Aspect | Detail |
| Source | Actual user sessions (browser/mobile SDK) |
| Captures | Page load, errors, Core Web Vitals (LCP, FID, CLS, INP) |
| Tools | Datadog RUM, New Relic Browser, Sentry, Cloudflare |
| Privacy | Mask PII; comply with GDPR/CCPA |
3. Health Dashboard Pattern
| Section | Content |
| Service Overview | RED metrics per service |
| SLO Compliance | Error budget burn-down |
| Dependency Health | Upstream service status |
| Resource Usage | CPU, memory, disk, network |
| Recent Deploys | Annotation overlay on charts |
| Tools | Grafana, Datadog, New Relic, Honeycomb |
4. Alerting Pattern
| Principle | Detail |
| Alert on Symptoms | User-facing impact (high latency, errors), not causes |
| Severity Levels | P1 page, P2 ticket, P3 ignore until business hours |
| Burn Rate Alerts | Multi-window SLO burn (1h fast, 6h slow) |
| Avoid | Static thresholds without context; alerts on every blip |
| Routing | PagerDuty, Opsgenie; escalation policies |
Warning: Alert fatigue kills response. Every alert must be actionable; if no action exists, delete the alert.
5. Anomaly Detection Pattern
| Approach | Detail |
| Statistical | Z-score, MAD, EWMA |
| Seasonal | Compare to same time last week (holt-winters) |
| ML-Based | Isolation Forest, Prophet, LSTM forecasting |
| Tools | Datadog Watchdog, AWS Lookout for Metrics, Anomaly |
6. SLA Monitoring Pattern
| Term | Detail |
| SLI | Indicator: measured signal (e.g., success rate) |
| SLO | Objective: target (99.9% over 30d) |
| SLA | Agreement: external contract with consequences |
| Error Budget | 1 − SLO; freeze releases if exhausted |
| Burn Rate | Rate at which budget is being consumed |
7. Latency Monitoring Pattern
| Metric | Detail |
| Avg / mean | Misleading; hides outliers |
| p50 (median) | Typical user experience |
| p95 / p99 | Tail latency; SLO targets |
| p99.9 | Worst-case for premium |
| Histogram | Use Prometheus histograms; aggregate via histogram_quantile() |
8. Error Rate Monitoring Pattern
| Calculation | Detail |
| Rate | errors / total requests over window |
| 5xx vs 4xx | 5xx server fault; 4xx often client |
| Threshold | SLO-derived (e.g., < 0.1%) |
| Multi-Window Burn | Alert if burn rate too high |
9. Resource Utilization Monitoring
| Resource | Key Metrics |
| CPU | Utilization %, throttling, run queue |
| Memory | RSS, heap, GC pauses, OOM |
| Disk | IOPS, latency, % full, queue depth |
| Network | Bandwidth, retransmits, conn count |
| JVM | Heap, GC, threads, class loading |
| DB | Connections, slow queries, replication lag |
10. Business Metrics Monitoring
| Metric | Example |
| Conversion Rate | Orders / sessions |
| Revenue per Minute | $ throughput |
| Active Users | DAU, MAU |
| Funnel Drop-off | Step-by-step abandonment |
| Why | Detect business impact before infra metrics show it |