Implementing Observability Practices

1. Understanding Observability Pillars

PillarQuestion Answered
MetricsWhat is broken? (aggregate health)
LogsWhat happened? (events)
TracesWhere in the request flow? (causality)
Profiles (4th)Why is it slow at code level?

2. Implementing Distributed Tracing

StepDetail
SDKOpenTelemetry auto + manual instrumentation
PropagationW3C traceparent everywhere
BackendTempo / Jaeger / Honeycomb

3. Collecting Metrics

StepDetail
InstrumentRED + USE + business
ExportOTLP or /metrics
Cardinality budgetDefine per service

4. Centralizing Logs

StepDetail
Stdout JSONContainer-friendly
AgentFluent Bit / Vector / OTel collector
BackendLoki / OpenSearch / Elastic

5. Creating Dashboards

PatternDetail
Single paneMetrics + logs + traces (Grafana)
Drill-downClick span → related logs
VersionedDashboards-as-code (Jsonnet/Terraform)

6. Setting Up Alerts

SourceDetail
MetricsSLO burn rate
LogsError pattern explosion
TracesAnomalous latency
RoutingAlertmanager → PagerDuty / Opsgenie

7. Implementing SLO Monitoring

ElementDetail
Multi-window burn rate5m+1h, 1h+6h pairs
ToolsSloth, Pyrra, Nobl9

8. Using Correlation IDs

FieldDetail
traceIdAcross services
requestIdPer HTTP request
In all logsAuto via MDC/context

9. Implementing Error Tracking

ToolCapability
Sentry / Rollbar / BugsnagGroup, dedupe, alert
Release taggingRegression detection
Source maps / symbolsReadable stacks

10. Creating Service Dependency Maps

SourceDetail
TracesAuto service graph
Service meshKiali (Istio)
CMDB / BackstageDeclared

11. Implementing Audit Trails

ElementDetail
Whatactor, action, resource, before/after, ts
StorageAppend-only / WORM
RetentionPer regulation

12. Using OpenTelemetry for Standardization

ComponentDetail
SDKsVendor-neutral instrumentation
OTLPWire protocol (gRPC/HTTP)
CollectorReceive → process → export to any backend
Semantic conventionsShared attribute names