What does observability mean and how do the three pillars differ?
Observability is the ability to understand a system's internal state from its outputs, so you can debug novel problems without shipping new code.
- Metrics: cheap numeric time series (latency, error rate, saturation, throughput). Great for dashboards and alerting; use histograms/percentiles, not just averages.
- Logs: discrete timestamped events, ideal for detail and forensics. Use structured JSON logs, correlation IDs and appropriate retention/cost control.
- Traces: end-to-end request paths across services showing where time is spent (OpenTelemetry, Jaeger). Essential for microservices and N+1 detection.
Tie it together with SLIs/SLOs and error budgets. Alert on symptoms users feel (high latency, error budget burn) rather than every noisy cause, and drive incident response with runbooks and postmortems.