Resource requests and limits control how Kubernetes schedules pods and handles resource contention. Set them wrong and you get OOMKilled pods, throttled CPU, or wasted cluster capacity.
Requests vs Limits
resources:
requests:
cpu: 200m # Guaranteed minimum
memory: 256Mi # Guaranteed minimum
limits:
cpu: "1" # Maximum allowed
memory: 512Mi # Maximum allowed (OOMKilled if exceeded)| Setting | Purpose | Exceeding It |
|---|---|---|
| Request | Scheduling guarantee — node must have this available | N/A (minimum) |
| Limit | Maximum usage | CPU: throttled. Memory: OOMKilled |
CPU Units
cpu: "1" # 1 vCPU core
cpu: "0.5" # Half a core
cpu: 500m # 500 millicores = 0.5 cores
cpu: 100m # 100 millicores = 0.1 cores
cpu: 250m # Quarter coreCPU limits cause throttling, not killing. Your app gets slower, not terminated.
Memory Units
memory: 128Mi # 128 mebibytes (128 × 1024² bytes)
memory: 256Mi
memory: 1Gi # 1 gibibyte
memory: 512M # 512 megabytes (decimal, not binary)Memory limits cause OOMKilled — the pod is terminated immediately.
Master this topic with hands-on labs
Go beyond reading — build real projects in sandboxed environments with expert video guidance.
Browse Courses →Quality of Service (QoS) Classes
Kubernetes assigns QoS based on your resource configuration:
Guaranteed (highest priority)
# Requests == Limits for ALL containers
resources:
requests:
cpu: 500m
memory: 256Mi
limits:
cpu: 500m
memory: 256MiLast to be evicted under memory pressure.
Burstable
# Requests < Limits (or only requests set)
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 512MiEvicted after BestEffort pods.
BestEffort (lowest priority)
# No requests or limits set
resources: {}First to be evicted. Never use in production.
Common Patterns
Web Application
resources:
requests:
cpu: 100m
memory: 128Mi
limits:
cpu: 500m
memory: 256MiAPI Server
resources:
requests:
cpu: 250m
memory: 256Mi
limits:
cpu: "1"
memory: 512MiBackground Worker
resources:
requests:
cpu: 500m
memory: 512Mi
limits:
cpu: "2"
memory: 1GiDatabase (Guaranteed QoS)
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "1"
memory: 2GiRight-Sizing
Check Actual Usage
# Current usage
kubectl top pods
kubectl top pods --containers
# Detailed per-pod
kubectl top pod my-pod --containersPrometheus Queries
# Average CPU usage over 24h
avg_over_time(
rate(container_cpu_usage_seconds_total{pod="my-pod"}[5m])[24h:]
)
# Peak memory usage over 24h
max_over_time(
container_memory_working_set_bytes{pod="my-pod"}[24h]
)
# CPU request vs actual usage (over-provisioning)
sum(kube_pod_container_resource_requests{resource="cpu"})
/ sum(rate(container_cpu_usage_seconds_total[5m]))Vertical Pod Autoscaler (VPA)
Automatically adjusts requests based on usage:
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: api-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: api
updatePolicy:
updateMode: "Off" # Start with recommendations only
resourcePolicy:
containerPolicies:
- containerName: api
minAllowed:
cpu: 50m
memory: 64Mi
maxAllowed:
cpu: "2"
memory: 2Gi# Check recommendations
kubectl describe vpa api-vpaGet weekly IT automation tips
Docker, Ansible, Terraform, MLOps — curated insights delivered to your inbox. No spam.
Subscribe Free →Troubleshooting
OOMKilled
kubectl describe pod my-pod | grep -A5 "Last State"
# Reason: OOMKilled
# Exit Code: 137
# Fix: increase memory limit
kubectl set resources deployment my-app --limits=memory=1GiCPU Throttling
# Check throttling
kubectl exec my-pod -- cat /sys/fs/cgroup/cpu/cpu.stat
# nr_throttled: 12345 ← high number = significant throttling
# Fix: increase CPU limit (or remove it)
kubectl set resources deployment my-app --limits=cpu=2Pending Pod (Insufficient Resources)
kubectl describe pod my-pod
# Events:
# FailedScheduling: Insufficient cpu / Insufficient memory
# Check node capacity
kubectl describe nodes | grep -A5 "Allocated resources"Should You Set CPU Limits?
Controversial topic:
Set CPU limits when: - Running on shared clusters (prevent noisy neighbors) - Compliance requires it - Predictable performance needed
Skip CPU limits when: - You want pods to burst when CPU is available - Throttling causes latency spikes - You trust your request values
Many teams set CPU requests but not CPU limits — pods get guaranteed minimum CPU but can burst higher when capacity is available.
What's Next?
Our MLflow for Kubernetes MLOps course covers resource management for ML training workloads. Docker Fundamentals teaches container resource controls. First lessons are free. -e ---
Ready to go deeper? Explore our hands-on DevOps courses — from Docker and Terraform to MLflow on Kubernetes.
Ready to learn by doing?
Stop reading tutorials — start building. Expert video courses with hands-on labs in real sandboxed environments.
Related Articles
Kubernetes Services and Ingress
Expose Kubernetes workloads with ClusterIP, NodePort, LoadBalancer services and Ingress controllers. Practical examples.
kubectl Cheat Sheet for DevOps
Essential kubectl commands for pods, deployments, services, logs, and debugging Kubernetes clusters. Copy-paste ready DevOps reference.
Kubernetes Pod Troubleshooting
Debug Kubernetes pods systematically. Fix CrashLoopBackOff, ImagePullBackOff, pending pods, and OOMKilled with diagnostic commands.
Kubernetes Service Types Explained
Kubernetes service types explained. ClusterIP, NodePort, LoadBalancer, ExternalName, and headless services with YAML examples.
Kubernetes Storage PV and PVC Guide
Kubernetes persistent storage explained: PVs, PVCs, StorageClasses, dynamic provisioning, StatefulSets, and backup strategies.
Kubernetes Troubleshooting Checklist
Kubernetes troubleshooting checklist. Pod failures, CrashLoopBackOff, networking issues, DNS resolution, and storage fixes.
Explore topics
Browse more articles on the topics covered here.