Designing Monitoring and Observability
1. Designing Application Metrics Collection
| Metric Type | Use |
|---|---|
| Counter | Monotonic events |
| Gauge | Current value (queue depth) |
| Histogram | Latency distributions |
| Summary | Pre-computed quantiles |
| Tools | Prometheus, OpenTelemetry, StatsD |
2. Designing Distributed Tracing
| Concept | Detail |
|---|---|
| Trace | Tree of spans across services |
| Span | One unit of work; parent_id linkage |
| Context propagation | W3C Trace Context headers |
| Sampling | Head / tail / probabilistic |
| Tools | Jaeger, Tempo, Honeycomb, Datadog APM |
3. Designing Centralized Logging
| Element | Detail |
|---|---|
| Collector | Fluent Bit, Vector, OpenTelemetry Collector |
| Backend | ELK, Loki, Splunk, Datadog |
| Format | Structured JSON |
| Correlation | trace_id, span_id, request_id |
| Retention tiers | Hot 7d, warm 30d, archive |
4. Designing Health Check Architecture
| Check | Detail |
|---|---|
| Liveness | Process responsive |
| Readiness | Dependencies ready |
| Deep health | For dashboards, not LB |
| Composition | Service-A health excludes B's health |
5. Designing Alert and Notification System
| Practice | Detail |
|---|---|
| SLO-based | Multi-window burn-rate alerts |
| Symptom over cause | Alert on user impact |
| Severity | P1 page / P2 email / P3 ticket |
| Routing | PagerDuty, Opsgenie |
| Anti-flap | For-duration thresholds |
6. Designing Application Performance Monitoring (APM)
| Capability | Detail |
|---|---|
| Auto-instrumentation | Frameworks, DB, HTTP |
| Slow transaction | Drill-down to span |
| Code-level profiling | Flamegraphs (Pyroscope, Parca) |
| RUM | Browser/mobile real user metrics |
| Tools | Datadog, New Relic, Dynatrace, Elastic APM |
7. Designing Monitoring Dashboards
| Practice | Detail |
|---|---|
| Service overview | RED/USE per service |
| SLO dashboard | Error budget remaining |
| Drilldown links | To traces and logs |
| Tools | Grafana, Kibana, Datadog, Looker |
8. Designing Error Tracking
| Tool | Detail |
|---|---|
| Sentry / Rollbar / Bugsnag | Aggregate exceptions; dedup |
| Source maps | For minified frontends |
| Release tagging | Identify regressions |
| User context | Without PII |
9. Designing Log Aggregation
| Layer | Detail |
|---|---|
| Agent | Tail files / stdout |
| Buffer | Kafka for spikes |
| Index | Hot search store |
| Cold | S3/GCS Parquet for long-term |
| Cost control | Sample / drop noisy logs |
10. Designing SLO and SLA Monitoring
| Term | Detail |
|---|---|
| SLI | What you measure (success rate) |
| SLO | Internal target (99.9%) |
| SLA | Customer commitment |
| Error budget | 1 - SLO; freeze releases when burned |
| Burn-rate alerts | Multi-window: 1h fast / 6h slow |
11. Designing Synthetic Monitoring
| Capability | Detail |
|---|---|
| Probes | Periodic from multiple regions |
| User flow tests | Login → checkout |
| API checks | Endpoint contract |
| Tools | Datadog Synthetics, Pingdom, Checkly |
12. Designing Observability Data Correlation
| Pillar | Detail |
|---|---|
| Three pillars | Metrics, traces, logs (+ profiles, events) |
| Common keys | service, env, version, trace_id |
| Exemplars | Link metric data point → trace |
| Single pane | Unified UI for triage |