etcd stores all Kubernetes cluster state ā deployments, services, secrets, configmaps, everything. Losing etcd without a backup means rebuilding from scratch.
What's in etcd
etcd contains:
āāā All Kubernetes objects (pods, services, deployments, etc.)
āāā RBAC configurations (roles, bindings)
āāā Secrets and ConfigMaps
āāā Custom Resources (CRDs and CRs)
āāā Namespace definitions
āāā Service account tokens
āāā Cluster configurationManual Snapshot
Find etcd Endpoint and Certs
# On a control plane node
kubectl -n kube-system get pods -l component=etcd
# Check etcd pod for cert paths
kubectl -n kube-system describe pod etcd-master-1 | grep -A5 CommandTypical cert locations:
ETCD_CERT=/etc/kubernetes/pki/etcd/server.crt
ETCD_KEY=/etc/kubernetes/pki/etcd/server.key
ETCD_CACERT=/etc/kubernetes/pki/etcd/ca.crt
ETCD_ENDPOINT=https://127.0.0.1:2379Take Snapshot
ETCDCTL_API=3 etcdctl snapshot save /backup/etcd-$(date +%Y%m%d-%H%M%S).db \
--endpoints=$ETCD_ENDPOINT \
--cert=$ETCD_CERT \
--key=$ETCD_KEY \
--cacert=$ETCD_CACERTVerify Snapshot
ETCDCTL_API=3 etcdctl snapshot status /backup/etcd-20260104-120000.db --write-table
# +----------+----------+------------+------------+
# | HASH | REVISION | TOTAL KEYS | TOTAL SIZE |
# +----------+----------+------------+------------+
# | abc12345 | 45678 | 1234 | 25 MB |
# +----------+----------+------------+------------+Master this topic with hands-on labs
Go beyond reading ā build real projects in sandboxed environments with expert video guidance.
Browse Courses āAutomated Backup CronJob
apiVersion: batch/v1
kind: CronJob
metadata:
name: etcd-backup
namespace: kube-system
spec:
schedule: "0 */6 * * *" # Every 6 hours
concurrencyPolicy: Forbid
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 3
jobTemplate:
spec:
template:
spec:
hostNetwork: true
nodeSelector:
node-role.kubernetes.io/control-plane: ""
tolerations:
- effect: NoSchedule
operator: Exists
containers:
- name: backup
image: bitnami/etcd:latest
command:
- /bin/sh
- -c
- |
set -e
FILENAME="etcd-$(date +%Y%m%d-%H%M%S).db"
etcdctl snapshot save "/backup/$FILENAME" \
--endpoints=https://127.0.0.1:2379 \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key \
--cacert=/etc/kubernetes/pki/etcd/ca.crt
etcdctl snapshot status "/backup/$FILENAME" --write-table
# Keep last 7 days
find /backup -name "etcd-*.db" -mtime +7 -delete
echo "Backup complete: $FILENAME"
env:
- name: ETCDCTL_API
value: "3"
volumeMounts:
- name: etcd-certs
mountPath: /etc/kubernetes/pki/etcd
readOnly: true
- name: backup
mountPath: /backup
volumes:
- name: etcd-certs
hostPath:
path: /etc/kubernetes/pki/etcd
- name: backup
persistentVolumeClaim:
claimName: etcd-backup-pvc
restartPolicy: OnFailureRestore from Snapshot
Stop kube-apiserver
# Move static pod manifests to stop API server and etcd
sudo mv /etc/kubernetes/manifests/kube-apiserver.yaml /tmp/
sudo mv /etc/kubernetes/manifests/etcd.yaml /tmp/Restore Snapshot
ETCDCTL_API=3 etcdctl snapshot restore /backup/etcd-20260104-120000.db \
--data-dir=/var/lib/etcd-restored \
--name=master-1 \
--initial-cluster=master-1=https://10.0.1.10:2380 \
--initial-cluster-token=etcd-cluster-1 \
--initial-advertise-peer-urls=https://10.0.1.10:2380Replace Data Directory
# Back up current (corrupted) data
sudo mv /var/lib/etcd /var/lib/etcd-corrupted
# Use restored data
sudo mv /var/lib/etcd-restored /var/lib/etcd
sudo chown -R etcd:etcd /var/lib/etcdRestart Components
# Restore static pod manifests
sudo mv /tmp/etcd.yaml /etc/kubernetes/manifests/
sudo mv /tmp/kube-apiserver.yaml /etc/kubernetes/manifests/
# Wait for etcd and API server to start
kubectl get nodes
kubectl get pods -ACluster Health
# Check cluster health
ETCDCTL_API=3 etcdctl endpoint health \
--endpoints=$ETCD_ENDPOINT \
--cert=$ETCD_CERT \
--key=$ETCD_KEY \
--cacert=$ETCD_CACERT
# Check cluster status
ETCDCTL_API=3 etcdctl endpoint status --write-table \
--endpoints=$ETCD_ENDPOINT \
--cert=$ETCD_CERT \
--key=$ETCD_KEY \
--cacert=$ETCD_CACERT
# List members
ETCDCTL_API=3 etcdctl member list --write-table \
--endpoints=$ETCD_ENDPOINT \
--cert=$ETCD_CERT \
--key=$ETCD_KEY \
--cacert=$ETCD_CACERT
# Defragment (reclaim space)
ETCDCTL_API=3 etcdctl defrag \
--endpoints=$ETCD_ENDPOINT \
--cert=$ETCD_CERT \
--key=$ETCD_KEY \
--cacert=$ETCD_CACERTGet weekly IT automation tips
Docker, Ansible, Terraform, MLOps ā curated insights delivered to your inbox. No spam.
Subscribe Free āBackup to S3
#!/bin/bash
set -e
FILENAME="etcd-$(date +%Y%m%d-%H%M%S).db"
BACKUP_DIR="/tmp/etcd-backup"
mkdir -p "$BACKUP_DIR"
ETCDCTL_API=3 etcdctl snapshot save "$BACKUP_DIR/$FILENAME" \
--endpoints=https://127.0.0.1:2379 \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key \
--cacert=/etc/kubernetes/pki/etcd/ca.crt
gzip "$BACKUP_DIR/$FILENAME"
aws s3 cp "$BACKUP_DIR/$FILENAME.gz" "s3://my-backups/etcd/$FILENAME.gz"
rm -f "$BACKUP_DIR/$FILENAME.gz"
echo "Uploaded $FILENAME.gz to S3"Disaster Recovery Checklist
| Step | Action |
|---|---|
| 1 | Stop kube-apiserver and etcd |
| 2 | Restore snapshot to new data directory |
| 3 | Replace etcd data directory |
| 4 | Start etcd, then kube-apiserver |
| 5 | Verify kubectl get nodes |
| 6 | Verify kubectl get pods -A |
| 7 | Check application health |
What's Next?
Our MLflow for Kubernetes MLOps course covers Kubernetes cluster management and operations. Docker Fundamentals teaches container orchestration basics. First lessons are free. -e ---
Ready to go deeper? Explore our hands-on DevOps courses ā from Docker and Terraform to MLflow on Kubernetes.
Ready to learn by doing?
Stop reading tutorials ā start building. Expert video courses with hands-on labs in real sandboxed environments.
Related Articles
Kubernetes Services and Ingress
Expose Kubernetes workloads with ClusterIP, NodePort, LoadBalancer services and Ingress controllers. Practical examples.
kubectl Cheat Sheet for DevOps
Essential kubectl commands for pods, deployments, services, logs, and debugging Kubernetes clusters. Copy-paste ready DevOps reference.
Kubernetes Pod Troubleshooting
Debug Kubernetes pods systematically. Fix CrashLoopBackOff, ImagePullBackOff, pending pods, and OOMKilled with diagnostic commands.
Kubernetes Helm Hooks and Testing
Use Helm hooks for database migrations, post-deploy verification, and rollback safety. Plus helm test and CI pipeline strategies.
Kubernetes HPA Autoscaling Guide
Configure Kubernetes HPA autoscaling. CPU and memory scaling, custom metrics, and practical YAML examples for production workloads.
Kubernetes Ingress Controllers Guide
Configure Kubernetes Ingress for HTTP routing. Nginx controller setup, TLS termination, path-based routing, and rate limiting.
Explore topics
Browse more articles on the topics covered here.