The Horizontal Pod Autoscaler (HPA) automatically adjusts the number of pod replicas based on observed metrics. Your application scales up during traffic spikes and scales down when demand drops.
How HPA Works
Metrics Server collects CPU/memory from kubelets
→ HPA controller checks metrics every 15 seconds
→ Compares current vs target utilization
→ Scales replicas up or downPrerequisites
# Metrics Server must be installed
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
# Verify
kubectl top nodes
kubectl top podsBasic HPA: CPU Scaling
apiVersion: apps/v1
kind: Deployment
metadata:
name: web-app
spec:
replicas: 2
selector:
matchLabels:
app: web-app
template:
metadata:
labels:
app: web-app
spec:
containers:
- name: app
image: my-app:latest
resources:
requests:
cpu: 200m # HPA needs requests defined!
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web-app
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web-app
minReplicas: 2
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70When average CPU across all pods exceeds 70% of the requested 200m (140m), HPA adds more pods.
Master this topic with hands-on labs
Go beyond reading — build real projects in sandboxed environments with expert video guidance.
Browse Courses →Multi-Metric HPA
Scale on both CPU and memory:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web-app
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web-app
minReplicas: 2
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 80
behavior:
scaleUp:
stabilizationWindowSeconds: 60
policies:
- type: Percent
value: 50
periodSeconds: 60
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 10
periodSeconds: 60The behavior section prevents flapping:
- Scale up: Wait 60s, then add up to 50% more pods per minute
- Scale down: Wait 5 minutes of stable low usage, then remove up to 10% per minute
Custom Metrics HPA
Scale on application-specific metrics (requires Prometheus Adapter):
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web-app
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web-app
minReplicas: 2
maxReplicas: 50
metrics:
- type: Pods
pods:
metric:
name: http_requests_per_second
target:
type: AverageValue
averageValue: "100"
- type: Object
object:
metric:
name: queue_depth
describedObject:
apiVersion: v1
kind: Service
name: rabbitmq
target:
type: Value
value: "50"This scales when: - Each pod handles more than 100 req/s - OR the RabbitMQ queue depth exceeds 50 messages
KEDA: Event-Driven Autoscaling
For more advanced scaling (queue depth, Kafka lag, cron schedules):
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: worker
spec:
scaleTargetRef:
name: worker-deployment
minReplicaCount: 0 # Scale to zero!
maxReplicaCount: 50
triggers:
- type: rabbitmq
metadata:
queueName: tasks
host: amqp://guest:guest@rabbitmq:5672/
queueLength: "10"
- type: cron
metadata:
timezone: Europe/Rome
start: 0 8 * * *
end: 0 20 * * *
desiredReplicas: "5"KEDA can scale to zero — no pods running when there is no work.
Get weekly IT automation tips
Docker, Ansible, Terraform, MLOps — curated insights delivered to your inbox. No spam.
Subscribe Free →Load Testing Your HPA
# Generate load
kubectl run load-test --image=busybox --rm -it -- \
sh -c "while true; do wget -q -O- http://web-app; done"
# Watch HPA in action
kubectl get hpa web-app --watch
# Watch pods scaling
kubectl get pods -l app=web-app --watchMonitoring HPA
# Current status
kubectl get hpa
# Detailed info
kubectl describe hpa web-app
# Events
kubectl get events --field-selector involvedObject.name=web-appKey Prometheus metrics:
# Current vs desired replicas
kube_horizontalpodautoscaler_status_current_replicas
kube_horizontalpodautoscaler_status_desired_replicas
# HPA condition
kube_horizontalpodautoscaler_status_condition{condition="ScalingActive"}Common Mistakes
| Mistake | Fix |
|---|---|
| No resource requests | HPA needs requests to calculate utilization |
| Min replicas = 1 | Use min 2 for high availability |
| No scale-down stabilization | Add stabilizationWindowSeconds to prevent flapping |
| Scaling on memory for JVM apps | JVM rarely releases memory; use CPU or custom metrics |
| Max too low | Set max high enough for peak traffic |
What's Next?
Our MLflow for Kubernetes MLOps course covers autoscaling ML inference workloads on Kubernetes. Docker Fundamentals builds the container foundation. First lessons are free. -e ---
Ready to go deeper? Explore our hands-on DevOps courses — from Docker and Terraform to MLflow on Kubernetes.
Ready to learn by doing?
Stop reading tutorials — start building. Expert video courses with hands-on labs in real sandboxed environments.
Related Articles
Kubernetes Cost Optimization Guide
Kubernetes clusters are often 60-70% over-provisioned. Learn practical cost optimization strategies: right-sizing, spot instances, autoscaling, namespace.
Kubernetes Services and Ingress
Expose Kubernetes workloads with ClusterIP, NodePort, LoadBalancer services and Ingress controllers. Practical examples.
kubectl Cheat Sheet for DevOps
Essential kubectl commands for pods, deployments, services, logs, and debugging Kubernetes clusters. Copy-paste ready DevOps reference.
Kubernetes Ingress Controllers Guide
Configure Kubernetes Ingress for HTTP routing. Nginx controller setup, TLS termination, path-based routing, and rate limiting.
Kubernetes Jobs and CronJobs Guide
Run batch workloads in Kubernetes with Jobs and CronJobs. Parallelism, backoff policies, TTL cleanup, and scheduling best practices.
Kubernetes Namespaces Multi-Tenancy
Kubernetes multi-tenancy with namespaces. Resource quotas, limit ranges, network policies, and RBAC for isolating teams.
Explore topics
Browse more articles on the topics covered here.