Monitoring tells you when something is broken. Observability tells you why. In distributed systems, you need both.
Monitoring vs Observability
| Aspect | Monitoring | Observability |
|---|---|---|
| Question | "Is it broken?" | "Why is it broken?" |
| Approach | Predefined dashboards & alerts | Explore unknown unknowns |
| Data | Metrics, uptime checks | Metrics + Logs + Traces |
| When useful | Known failure modes | Novel, unexpected failures |
The Three Pillars
1. Metrics (Numbers Over Time)
CPU: 78% → 85% → 92% → 99% 🔥
Request rate: 500/s → 200/s → 50/s
Error rate: 0.1% → 5% → 25%Best for: Dashboards, alerts, trends, SLOs.
Prometheus example:
# Error rate
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# P99 latency
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)2. Logs (Events with Context)
{
"timestamp": "2026-01-29T14:23:01Z",
"level": "error",
"service": "payment-service",
"trace_id": "abc123def456",
"user_id": "usr_789",
"message": "Payment failed: card declined",
"error": "stripe_card_declined",
"amount": 2900,
"currency": "EUR"
}Best for: Debugging specific requests, audit trails, error details.
3. Traces (Request Flow Across Services)
[Client] → [API Gateway] → [Auth Service] → [Payment Service] → [Database]
0ms 5ms 15ms 45ms 80ms
↑
Slow query: 35msBest for: Finding bottlenecks in distributed systems, understanding request flow.
Master this topic with hands-on labs
Go beyond reading — build real projects in sandboxed environments with expert video guidance.
Browse Courses →OpenTelemetry (The Standard)
OpenTelemetry provides a single SDK for all three signals:
// Auto-instrumentation (Node.js)
import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-http';
import { OTLPMetricExporter } from '@opentelemetry/exporter-metrics-otlp-http';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
const sdk = new NodeSDK({
traceExporter: new OTLPTraceExporter({
url: 'http://otel-collector:4318/v1/traces',
}),
metricReader: new PeriodicExportingMetricReader({
exporter: new OTLPMetricExporter({
url: 'http://otel-collector:4318/v1/metrics',
}),
}),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();This automatically instruments: - HTTP requests (incoming and outgoing) - Database queries (pg, mysql, redis) - gRPC calls - Express/Fastify routes
Custom Spans
import { trace } from '@opentelemetry/api';
const tracer = trace.getTracer('payment-service');
async function processPayment(order: Order) {
return tracer.startActiveSpan('process-payment', async (span) => {
span.setAttribute('order.id', order.id);
span.setAttribute('order.amount', order.amount);
try {
const result = await stripe.charges.create({
amount: order.amount,
currency: 'eur',
});
span.setAttribute('payment.status', 'success');
return result;
} catch (error) {
span.recordException(error);
span.setStatus({ code: SpanStatusCode.ERROR });
throw error;
} finally {
span.end();
}
});
}The Observability Stack
Prometheus + Grafana + Loki + Tempo
# docker-compose.yml
services:
prometheus:
image: prom/prometheus:latest
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
grafana:
image: grafana/grafana:latest
ports:
- "3000:3000"
loki:
image: grafana/loki:latest
tempo:
image: grafana/tempo:latest
otel-collector:
image: otel/opentelemetry-collector-contrib:latest
volumes:
- ./otel-config.yaml:/etc/otel/config.yaml
command: ["--config=/etc/otel/config.yaml"]OTel Collector Config
# otel-config.yaml
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
exporters:
prometheus:
endpoint: 0.0.0.0:8889
loki:
endpoint: http://loki:3100/loki/api/v1/push
otlp/tempo:
endpoint: tempo:4317
tls:
insecure: true
service:
pipelines:
metrics:
receivers: [otlp]
exporters: [prometheus]
logs:
receivers: [otlp]
exporters: [loki]
traces:
receivers: [otlp]
exporters: [otlp/tempo]Structured Logging
import pino from 'pino';
const logger = pino({
level: process.env.LOG_LEVEL || 'info',
formatters: {
level: (label) => ({ level: label }),
},
});
// Always include context
logger.info({ orderId: order.id, userId: user.id, amount: order.total }, 'Order placed');
logger.error({ err, orderId: order.id }, 'Payment failed');Get weekly IT automation tips
Docker, Ansible, Terraform, MLOps — curated insights delivered to your inbox. No spam.
Subscribe Free →Correlating the Three Pillars
The real power is connecting metrics → logs → traces:
- Alert fires: Error rate > 5% (metric)
- Filter logs:
level=error AND service=payment-service AND time=last_5m - Find trace: Click
trace_idin log entry - See full request flow: API Gateway → Auth → Payment → Stripe API (timeout at 30s)
- Root cause: Stripe API latency spike
This workflow requires consistent trace_id propagation across all services.
SLOs (Service Level Objectives)
Define what "good" looks like:
| SLI | SLO | Error Budget |
|---|---|---|
| Availability | 99.95% over 30 days | 21.9 min downtime/month |
| Latency (P99) | < 500ms | 0.05% of requests can exceed |
| Error rate | < 0.1% | ~4,320 errors/month at 100 RPS |
# Error budget burn rate
1 - (
sum(rate(http_requests_total{status!~"5.."}[1h]))
/ sum(rate(http_requests_total[1h]))
) / (1 - 0.9995)What's Next?
Our MLflow for Kubernetes MLOps course covers observability for ML systems. Docker Fundamentals teaches container monitoring basics. First lessons are free. -e ---
Ready to go deeper? Explore our hands-on DevOps courses — practical labs covering Docker, Ansible, Terraform, and more.
Ready to learn by doing?
Stop reading tutorials — start building. Expert video courses with hands-on labs in real sandboxed environments.
Related Articles
OpenTelemetry Getting Started Guide
OpenTelemetry is the standard for observability instrumentation. Learn how to add traces, metrics, and logs to your applications with OTel SDKs.
Prometheus Monitoring Beginner Guide
Get started with Prometheus monitoring. Metrics collection, PromQL queries, Grafana dashboards, and alert configuration step by step.
Infrastructure Monitoring Stack
Build a complete monitoring stack. Prometheus for metrics, Grafana for dashboards, Loki for logs, and Alertmanager for alerting.
OPA Gatekeeper Kubernetes Policies
OPA Gatekeeper enforces custom policies in Kubernetes at admission time. Learn how to write ConstraintTemplates, enforce security standards, and prevent.
OpenClaw + Ansible Automation
Combine OpenClaw's AI agent with Ansible for intelligent infrastructure automation — writing playbooks, debugging tasks, and orchestrating deployments.
OpenClaw Architecture Deep Dive
Deep dive into OpenClaw's architecture — the gateway, sessions, tool system, memory layer, and channel providers that make it tick.
Explore topics
Browse more articles on the topics covered here.