Prometheus is the standard for monitoring Kubernetes and cloud-native applications. This guide gets you from zero to dashboards and alerts.
What Prometheus Does
Prometheus scrapes metrics from your applications and infrastructure at regular intervals, stores them as time series data, and lets you query, visualize, and alert on them.
Your App → /metrics endpoint → Prometheus scrapes → Stores time series
↓
PromQL queries → Grafana dashboards
↓
Alert rules → Alertmanager → Slack/PagerDutyQuick Start with Docker Compose
services:
prometheus:
image: prom/prometheus:latest
ports:
- "9090:9090"
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
- prometheus-data:/prometheus
grafana:
image: grafana/grafana:latest
ports:
- "3000:3000"
environment:
- GF_SECURITY_ADMIN_PASSWORD=admin
node-exporter:
image: prom/node-exporter:latest
ports:
- "9100:9100"
volumes:
prometheus-data:prometheus.yml:
global:
scrape_interval: 15s
scrape_configs:
- job_name: prometheus
static_configs:
- targets: ['localhost:9090']
- job_name: node
static_configs:
- targets: ['node-exporter:9100']docker compose up -d
# Prometheus: http://localhost:9090
# Grafana: http://localhost:3000 (admin/admin)Master this topic with hands-on labs
Go beyond reading — build real projects in sandboxed environments with expert video guidance.
Browse Courses →Instrumenting Your Application
Node.js / Express
npm install prom-clientimport { collectDefaultMetrics, Registry, Counter, Histogram } from 'prom-client';
const register = new Registry();
collectDefaultMetrics({ register });
// Custom metrics
const httpRequests = new Counter({
name: 'http_requests_total',
help: 'Total HTTP requests',
labelNames: ['method', 'path', 'status'],
registers: [register],
});
const httpDuration = new Histogram({
name: 'http_request_duration_seconds',
help: 'HTTP request duration',
labelNames: ['method', 'path'],
buckets: [0.01, 0.05, 0.1, 0.5, 1, 5],
registers: [register],
});
// Middleware
app.use((req, res, next) => {
const end = httpDuration.startTimer({ method: req.method, path: req.route?.path || req.path });
res.on('finish', () => {
httpRequests.inc({ method: req.method, path: req.route?.path || req.path, status: res.statusCode });
end();
});
next();
});
// Metrics endpoint
app.get('/metrics', async (req, res) => {
res.set('Content-Type', register.contentType);
res.send(await register.metrics());
});Python / Flask
from prometheus_client import Counter, Histogram, generate_latest
import time
REQUEST_COUNT = Counter('http_requests_total', 'Total requests', ['method', 'endpoint', 'status'])
REQUEST_LATENCY = Histogram('http_request_duration_seconds', 'Request latency', ['method', 'endpoint'])
@app.before_request
def before_request():
request.start_time = time.time()
@app.after_request
def after_request(response):
latency = time.time() - request.start_time
REQUEST_COUNT.labels(request.method, request.path, response.status_code).inc()
REQUEST_LATENCY.labels(request.method, request.path).observe(latency)
return response
@app.route('/metrics')
def metrics():
return generate_latest(), 200, {'Content-Type': 'text/plain'}Essential PromQL Queries
Request Rate
# Requests per second (last 5 min)
rate(http_requests_total[5m])
# Requests per second by status
sum by (status) (rate(http_requests_total[5m]))Error Rate
# Error percentage
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
* 100Latency (P50, P90, P99)
# 50th percentile
histogram_quantile(0.50, rate(http_request_duration_seconds_bucket[5m]))
# 99th percentile
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))Resource Usage
# CPU usage percentage
100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100)
# Memory usage percentage
(1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100
# Disk usage percentage
(1 - node_filesystem_avail_bytes / node_filesystem_size_bytes) * 100Alerting Rules
alert-rules.yml:
groups:
- name: application
rules:
- alert: HighErrorRate
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m])) > 0.05
for: 5m
labels:
severity: critical
annotations:
summary: "High error rate ({{ $value | humanizePercentage }})"
- alert: HighLatency
expr: |
histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m])) > 1
for: 5m
labels:
severity: warning
annotations:
summary: "P99 latency above 1s"
- name: infrastructure
rules:
- alert: HighCPU
expr: 100 - (avg(rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 80
for: 10m
labels:
severity: warning
- alert: DiskSpaceLow
expr: (1 - node_filesystem_avail_bytes / node_filesystem_size_bytes) > 0.85
for: 5m
labels:
severity: criticalGet weekly IT automation tips
Docker, Ansible, Terraform, MLOps — curated insights delivered to your inbox. No spam.
Subscribe Free →Grafana Dashboard
1. Add Prometheus as data source: http://prometheus:9090
2. Import community dashboards:
- Node Exporter Full: ID 1860
- Kubernetes Cluster: ID 6417
3. Build custom panels with PromQL queries above
The Four Golden Signals
Google SRE recommends monitoring these for every service:
| Signal | What to Measure | PromQL |
|---|---|---|
| Latency | Request duration | histogram_quantile(0.99, ...) |
| Traffic | Requests per second | sum(rate(http_requests_total[5m])) |
| Errors | Error rate | rate(http_requests_total{status=~"5.."}[5m]) |
| Saturation | Resource utilization | CPU, memory, disk, connections |
What's Next?
Our MLflow for Kubernetes MLOps course covers monitoring ML workloads with Prometheus and Grafana. Docker Fundamentals teaches container observability basics. First lessons are free. -e ---
Ready to go deeper? Explore our hands-on DevOps courses — practical labs covering Docker, Ansible, Terraform, and more.
Ready to learn by doing?
Stop reading tutorials — start building. Expert video courses with hands-on labs in real sandboxed environments.
Related Articles
Infrastructure Monitoring Stack
Build a complete monitoring stack. Prometheus for metrics, Grafana for dashboards, Loki for logs, and Alertmanager for alerting.
Grafana Dashboard Best Practices
Build effective Grafana dashboards for monitoring. Layout patterns, template variables, alert integration, and dashboard-as-code.
Thanos Long-Term Prometheus Storage
Thanos extends Prometheus with unlimited retention, global querying across clusters, and downsampling. Learn how to deploy Thanos Sidecar and Store Gateway.
Pulumi Infrastructure as Code Guide
Pulumi lets you write infrastructure in TypeScript, Python, Go, or C# with full programming language features. Learn how Pulumi works, how it compares.
Python for DevOps Automation
Python for DevOps automation. HTTP API clients, file processing, AWS boto3, subprocess management, and CLI tools for infrastructure.
Quality vs Cost in DevOps
The quality-cost tradeoff in DevOps is real but misunderstood. Learn why cutting quality to reduce cost usually increases total cost, and how to find.
Explore topics
Browse more articles on the topics covered here.