Implementing Monitoring and Alerting Patterns

1. Synthetic Monitoring Pattern

AspectDetail
DefinitionScripted requests run on schedule from external locations
ToolsPingdom, Datadog Synthetics, Checkly, Grafana k6
UseDetect outage before users report; SLO measurement
CadenceEvery 1-5 minutes per critical journey

2. Real User Monitoring Pattern

AspectDetail
SourceActual user sessions (browser/mobile SDK)
CapturesPage load, errors, Core Web Vitals (LCP, FID, CLS, INP)
ToolsDatadog RUM, New Relic Browser, Sentry, Cloudflare
PrivacyMask PII; comply with GDPR/CCPA

3. Health Dashboard Pattern

SectionContent
Service OverviewRED metrics per service
SLO ComplianceError budget burn-down
Dependency HealthUpstream service status
Resource UsageCPU, memory, disk, network
Recent DeploysAnnotation overlay on charts
ToolsGrafana, Datadog, New Relic, Honeycomb

4. Alerting Pattern

PrincipleDetail
Alert on SymptomsUser-facing impact (high latency, errors), not causes
Severity LevelsP1 page, P2 ticket, P3 ignore until business hours
Burn Rate AlertsMulti-window SLO burn (1h fast, 6h slow)
AvoidStatic thresholds without context; alerts on every blip
RoutingPagerDuty, Opsgenie; escalation policies
Warning: Alert fatigue kills response. Every alert must be actionable; if no action exists, delete the alert.

5. Anomaly Detection Pattern

ApproachDetail
StatisticalZ-score, MAD, EWMA
SeasonalCompare to same time last week (holt-winters)
ML-BasedIsolation Forest, Prophet, LSTM forecasting
ToolsDatadog Watchdog, AWS Lookout for Metrics, Anomaly

6. SLA Monitoring Pattern

TermDetail
SLIIndicator: measured signal (e.g., success rate)
SLOObjective: target (99.9% over 30d)
SLAAgreement: external contract with consequences
Error Budget1 − SLO; freeze releases if exhausted
Burn RateRate at which budget is being consumed

7. Latency Monitoring Pattern

MetricDetail
Avg / meanMisleading; hides outliers
p50 (median)Typical user experience
p95 / p99Tail latency; SLO targets
p99.9Worst-case for premium
HistogramUse Prometheus histograms; aggregate via histogram_quantile()

8. Error Rate Monitoring Pattern

CalculationDetail
Rateerrors / total requests over window
5xx vs 4xx5xx server fault; 4xx often client
ThresholdSLO-derived (e.g., < 0.1%)
Multi-Window BurnAlert if burn rate too high

9. Resource Utilization Monitoring

ResourceKey Metrics
CPUUtilization %, throttling, run queue
MemoryRSS, heap, GC pauses, OOM
DiskIOPS, latency, % full, queue depth
NetworkBandwidth, retransmits, conn count
JVMHeap, GC, threads, class loading
DBConnections, slow queries, replication lag

10. Business Metrics Monitoring

MetricExample
Conversion RateOrders / sessions
Revenue per Minute$ throughput
Active UsersDAU, MAU
Funnel Drop-offStep-by-step abandonment
WhyDetect business impact before infra metrics show it