How to back up and restore etcd for a Kubernetes cluster
On a cluster whose control plane you run yourself, such as kubeadm, back up etcd with etcdctl snapshot save against one member over TLS, check the file with etcdutl snapshot status, then encrypt it and copy it off the node. To restore, stop the API server, run etcdutl snapshot restore into a new data directory and point the etcd static Pod at it. Managed Kubernetes runs etcd for you, so there you back up through the Kubernetes API instead.
When an etcd snapshot applies
Kubernetes keeps every API object in etcd: Deployments, Services, Secrets, ConfigMaps, RBAC rules, custom resources. Lose etcd and you lose the cluster's desired state, even while the nodes keep running. You can only snapshot etcd where you can reach it:
| Cluster | etcd snapshot? |
|---|---|
| kubeadm, stacked etcd | Yes, on a control-plane node. etcd runs as a static Pod from /etc/kubernetes/manifests/etcd.yaml. |
| kubeadm or hand-built, external etcd | Yes, on the etcd hosts, with that cluster's certificates |
| k3s with embedded etcd | Use k3s etcd-snapshot save and k3s's scheduled snapshots instead (5 kept by default) |
| EKS, GKE, AKS, DigitalOcean Kubernetes | No. The provider runs the control plane. |
Per their documentation: EKS runs etcd inside an EKS-managed VPC on private subnets; Azure operates etcd for AKS; GKE manages the control plane and keeps cluster state in etcd or in Spanner; DigitalOcean's control plane is fully managed and its configuration can't be changed. On those, back up through the API with Velero.
On k3s, use its own k3s etcd-snapshot command, or copy its SQLite datastore on a single server: see how to back up a k3s cluster.
What a snapshot does not include
- Volume data. A PersistentVolumeClaim is an object in etcd; the bytes on its disk are not. Back those up with Velero, CSI snapshots or a database dump.
- Files on the control-plane node. kubeadm keeps the cluster's certificate authorities and keys in
/etc/kubernetes/pki, and kubeconfig files and static Pod manifests under/etc/kubernetes. A rebuilt control plane needs the same CA, so copypkiwith each snapshot. - The encryption key. With encryption at rest on, Secrets are encrypted with the key in your
--encryption-provider-configfile. Kubernetes' docs say that losing every copy of the key means deleting the resources encrypted with it. Back that file up separately. - Images and anything outside the cluster: registries, DNS, load balancers, cloud disks.
By default the API server stores resources in etcd as plain text, so a snapshot holds every Secret in readable form. Kubernetes' etcd guide says to encrypt snapshot files. Treat one like a root password.
Find the endpoint, certificates and version
On a control-plane node, read etcd's settings from its manifest:
sudo grep -E 'image:|--name=|--data-dir=|--listen-client-urls=|--initial-advertise-peer-urls=|--initial-cluster=' /etc/kubernetes/manifests/etcd.yamlThe image tag carries the version: registry.k8s.io/etcd:3.7.1-0 is etcd 3.7.1. etcd listens on https://127.0.0.1:2379 and demands a client certificate. kubeadm's certificate guide names the files etcdctl should use:
| etcdctl flag | kubeadm file | Purpose |
|---|---|---|
--cacert | /etc/kubernetes/pki/etcd/ca.crt | Trusts etcd's server certificate |
--cert | /etc/kubernetes/pki/etcd/healthcheck-client.crt | A client certificate signed by the etcd CA |
--key | /etc/kubernetes/pki/etcd/healthcheck-client.key | Its private key |
Install etcdctl and etcdutl
kubeadm's packages don't include either tool. Take the release that matches your etcd version from the etcd project's GitHub releases and extract only the two clients:
ETCD_VER=v3.7.1curl -L https://github.com/etcd-io/etcd/releases/download/${ETCD_VER}/etcd-${ETCD_VER}-linux-amd64.tar.gz -o /tmp/etcd.tar.gzsudo tar -xzf /tmp/etcd.tar.gz -C /usr/local/bin --strip-components=1 --no-same-owner etcd-${ETCD_VER}-linux-amd64/etcdctl etcd-${ETCD_VER}-linux-amd64/etcdutletcdctl version && etcdutl versionetcdctl talks to a running etcd over the network; etcdutl works on files. etcd 3.6 removed snapshot restore and snapshot status from etcdctl, so save with etcdctl and check and restore with etcdutl. Older guides set ETCDCTL_API=3; that only matters for etcdctl releases before 3.4.
Take a snapshot
sudo install -d -m 700 /var/backups/etcdsudo etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/healthcheck-client.crt --key=/etc/kubernetes/pki/etcd/healthcheck-client.key snapshot save /var/backups/etcd/etcd-$(date +%F-%H%M).dbetcdctl writes to a .part file, syncs it and renames it when complete, then prints Snapshot saved at and the path. The file is readable only by its owner. Kubernetes' guide says taking a snapshot doesn't affect the member's performance.
Give it exactly one endpoint. With several, it stops with snapshot must be requested to one selected node, not multiple. Every member holds the whole keyspace, so on a cluster with several control-plane nodes one healthy member's snapshot is enough.
Check the snapshot
sudo etcdutl --write-out=table snapshot status /var/backups/etcd/etcd-2026-10-03-0200.dbIt prints the snapshot's hash, revision, total keys and size. A file that fails here won't restore either. Compare the key count with earlier snapshots: a sudden drop means an empty cluster or the wrong one.
Schedule it, encrypted and off the node
A snapshot left on the control-plane node dies with the node. The script below saves and checks a snapshot, encrypts it and a copy of /etc/kubernetes/pki to a GPG public key (set up as in encrypting backups), moves the encrypted files to object storage with rclone, and keeps three days of plain snapshots locally for quick restores.
#!/usr/bin/env bash
set -euo pipefail
KEY="A1D07C3D7D97CB70E2285FFD6A1D1AD200953CD5" # your GPG key's fingerprint
DIR=/var/backups/etcd
STAMP=$(date +%F-%H%M)
PKI=/etc/kubernetes/pki/etcd
install -d -m 700 "$DIR"
etcdctl --endpoints=https://127.0.0.1:2379 --cacert="$PKI/ca.crt" \
--cert="$PKI/healthcheck-client.crt" --key="$PKI/healthcheck-client.key" \
snapshot save "$DIR/etcd-$STAMP.db"
etcdutl snapshot status "$DIR/etcd-$STAMP.db"
gpg --batch --encrypt --recipient "$KEY" --output "$DIR/etcd-$STAMP.db.gpg" "$DIR/etcd-$STAMP.db"
tar -czf - -C /etc/kubernetes pki | gpg --batch --encrypt --recipient "$KEY" --output "$DIR/pki-$STAMP.tar.gz.gpg"
rclone move "$DIR" offsite:etcd-backups/cp1 --include "*.gpg"
find "$DIR" -name 'etcd-*.db' -mtime +3 -delete[Unit]
Description=etcd snapshot backup
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/etcd-backup.sh[Unit]
Description=Hourly etcd snapshot backup
[Timer]
OnCalendar=hourly
RandomizedDelaySec=5min
Persistent=true
[Install]
WantedBy=timers.targetsudo chmod 700 /usr/local/sbin/etcd-backup.sh && sudo systemctl daemon-reload && sudo systemctl enable --now etcd-backup.timerOnCalendar=hourly fires at the top of each hour, RandomizedDelaySec adds up to five minutes so several nodes don't upload at once, and Persistent=true runs a missed backup as soon as the node is back. systemctl list-timers etcd-backup.timer shows the next run; journalctl -u etcd-backup shows the output. A Kubernetes CronJob could run the same commands, but it needs the scheduler and API server, which are what's broken when you need a fresh snapshot most. Set retention on the bucket (retention policy).
Restore on a kubeadm control plane
Kubernetes' guide says to stop every API server, restore etcd, then start the API servers again, and to restart the scheduler, controller manager and kubelet so nothing acts on stale data. On a single control-plane node, as root:
1. Move the control-plane manifests out of the static Pod folder. The kubelet stops a static Pod when its file disappears; after about 20 seconds crictl ps no longer lists them.
mkdir /root/paused && mv /etc/kubernetes/manifests/kube-apiserver.yaml /etc/kubernetes/manifests/kube-controller-manager.yaml /etc/kubernetes/manifests/kube-scheduler.yaml /root/paused/2. Restore into a new data directory, with the name and peer URL from etcd.yaml (here cp1 and 10.0.0.10):
etcdutl snapshot restore /var/backups/etcd/etcd-2026-10-03-0200.db --data-dir /var/lib/etcd-restore --name cp1 --initial-cluster cp1=https://10.0.0.10:2380 --initial-advertise-peer-urls https://10.0.0.10:2380 --bump-revision 1000000000 --mark-compactedetcd's recovery guide highly recommends --bump-revision with --mark-compacted for Kubernetes. Without them, the revision goes backwards and controllers' caches may not refresh correctly; with them, every watch ends and those caches are invalidated. One billion covers a week-old snapshot if etcd takes fewer than 1,500 writes a second.
3. Point the etcd Pod's etcd-data volume at the new directory and restart the kubelet:
sed -i 's#path: /var/lib/etcd$#path: /var/lib/etcd-restore#' /etc/kubernetes/manifests/etcd.yamlsystemctl restart kubelet4. Wait until etcd answers:
etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/healthcheck-client.crt --key=/etc/kubernetes/pki/etcd/healthcheck-client.key endpoint health5. Put the manifests back, check kubectl get nodes and kubectl get pods -A, and run systemctl restart kubelet on every node:
mv /root/paused/*.yaml /etc/kubernetes/manifests/Keep the old /var/lib/etcd until the cluster checks out; it is your way back. With several control-plane nodes, stop the API server on all of them first, then restore the same snapshot on every member, each with its own --name and --initial-advertise-peer-urls and an --initial-cluster listing all members, as in etcd's disaster recovery guide. Start the API servers only when every member is healthy.
Test a snapshot without touching the cluster
Prove a snapshot restores on a spare machine with the same etcd release, never on a control-plane node, where port 2379 is taken. A restore without member flags makes a one-member cluster that plain etcd starts:
etcdutl snapshot restore etcd-2026-10-03-0200.db --data-dir /tmp/etcd-testetcd --data-dir /tmp/etcd-testIn a second terminal, list the namespaces the snapshot holds:
etcdctl get /registry/namespaces --prefix --keys-onlyThe API server stores objects under /registry by default, one key per object, so you should see one key per namespace. Values are binary; the keys are enough to show the snapshot holds your cluster. For a full rehearsal, build a throwaway kubeadm cluster at the same versions and run the restore steps. See testing backup restores.
Common errors
snapshot must be requested to one selected node, not multiple: pass one endpoint tosnapshot save.- A certificate error or a timeout when saving:
--cacert,--certor--keyis missing or wrong. kubeadm's etcd accepts only clients with a certificate from its CA. data-dir "/var/lib/etcd-restore" not empty or could not be read: etcdutl won't restore over existing data. Pick a new directory.snapshot missing hash but --skip-hash-check=false: the file was copied frommember/snap/db, not saved with etcdctl. Add--skip-hash-check, knowing a copied file can miss writes still in the WAL.--mark-compacted required if --revision-bump > 0: use--bump-revisionand--mark-compactedtogether.snapshot restoreorsnapshot statusmissing from etcdctl: it is 3.6 or newer. Use etcdutl.- Objects created after the snapshot are gone after a restore. That is the snapshot's point in time, so snapshot often enough (RPO and RTO).
Frequently asked questions
- Can I back up etcd on EKS, GKE or AKS?
- No. The provider runs the control plane, etcd included, and you only reach the API server. Back those clusters up through the Kubernetes API with a tool like Velero.
- Is ETCDCTL_API=3 still needed?
- Only for etcdctl releases before 3.4. Since 3.4 the v3 API is the default.
- Does an etcd snapshot include persistent volumes?
- No. It holds the PersistentVolumeClaim and PersistentVolume objects, not the data on the disks. Back up volume data separately.
- Should I restore with etcdctl or etcdutl?
- etcdutl. etcd 3.6 removed snapshot restore and snapshot status from etcdctl; for snapshots, etcdctl now only has save.
- How often should I snapshot etcd?
- As often as you can afford to lose cluster changes. A snapshot is a copy of the etcd database, usually small, so hourly is cheap on most clusters.
How this was checked
Commands, limits and prices were checked against these official pages, on October 4, 2026:
- Kubernetes documentation: Operating etcd clusters for Kubernetes
- Kubernetes documentation: PKI certificates and requirements
- Kubernetes documentation: Create static Pods
- Kubernetes documentation: kubeadm Configuration (v1beta4)
- Kubernetes documentation: Encrypting Confidential Data at Rest
- Kubernetes documentation: kube-apiserver (--etcd-prefix)
- kubeadm source: etcd static Pod manifest and supported etcd versions
- etcd documentation v3.7: Disaster recovery
- etcd documentation v3.7: Install
- etcd: etcdctl README (release-3.6)
- etcd: etcdutl README
- etcd v3.7.1 release
- K3s documentation: etcd-snapshot
- Amazon EKS Best Practices: EKS Control Plane
- Google Kubernetes Engine: GKE cluster architecture
- Azure Kubernetes Service: Core concepts
- DigitalOcean Kubernetes: Managed elements
- systemd.timer manual page