VPS Snaps

How to back up and restore etcd for a Kubernetes cluster

On a cluster whose control plane you run yourself, such as kubeadm, back up etcd with etcdctl snapshot save against one member over TLS, check the file with etcdutl snapshot status, then encrypt it and copy it off the node. To restore, stop the API server, run etcdutl snapshot restore into a new data directory and point the etcd static Pod at it. Managed Kubernetes runs etcd for you, so there you back up through the Kubernetes API instead.

8 min readUpdated Checked against official documentation

When an etcd snapshot applies

Kubernetes keeps every API object in etcd: Deployments, Services, Secrets, ConfigMaps, RBAC rules, custom resources. Lose etcd and you lose the cluster's desired state, even while the nodes keep running. You can only snapshot etcd where you can reach it:

Clusteretcd snapshot?
kubeadm, stacked etcdYes, on a control-plane node. etcd runs as a static Pod from /etc/kubernetes/manifests/etcd.yaml.
kubeadm or hand-built, external etcdYes, on the etcd hosts, with that cluster's certificates
k3s with embedded etcdUse k3s etcd-snapshot save and k3s's scheduled snapshots instead (5 kept by default)
EKS, GKE, AKS, DigitalOcean KubernetesNo. The provider runs the control plane.

Per their documentation: EKS runs etcd inside an EKS-managed VPC on private subnets; Azure operates etcd for AKS; GKE manages the control plane and keeps cluster state in etcd or in Spanner; DigitalOcean's control plane is fully managed and its configuration can't be changed. On those, back up through the API with Velero.

On k3s, use its own k3s etcd-snapshot command, or copy its SQLite datastore on a single server: see how to back up a k3s cluster.

What a snapshot does not include

  • Volume data. A PersistentVolumeClaim is an object in etcd; the bytes on its disk are not. Back those up with Velero, CSI snapshots or a database dump.
  • Files on the control-plane node. kubeadm keeps the cluster's certificate authorities and keys in /etc/kubernetes/pki, and kubeconfig files and static Pod manifests under /etc/kubernetes. A rebuilt control plane needs the same CA, so copy pki with each snapshot.
  • The encryption key. With encryption at rest on, Secrets are encrypted with the key in your --encryption-provider-config file. Kubernetes' docs say that losing every copy of the key means deleting the resources encrypted with it. Back that file up separately.
  • Images and anything outside the cluster: registries, DNS, load balancers, cloud disks.

By default the API server stores resources in etcd as plain text, so a snapshot holds every Secret in readable form. Kubernetes' etcd guide says to encrypt snapshot files. Treat one like a root password.

Find the endpoint, certificates and version

On a control-plane node, read etcd's settings from its manifest:

Terminal
sudo grep -E 'image:|--name=|--data-dir=|--listen-client-urls=|--initial-advertise-peer-urls=|--initial-cluster=' /etc/kubernetes/manifests/etcd.yaml

The image tag carries the version: registry.k8s.io/etcd:3.7.1-0 is etcd 3.7.1. etcd listens on https://127.0.0.1:2379 and demands a client certificate. kubeadm's certificate guide names the files etcdctl should use:

etcdctl flagkubeadm filePurpose
--cacert/etc/kubernetes/pki/etcd/ca.crtTrusts etcd's server certificate
--cert/etc/kubernetes/pki/etcd/healthcheck-client.crtA client certificate signed by the etcd CA
--key/etc/kubernetes/pki/etcd/healthcheck-client.keyIts private key

Install etcdctl and etcdutl

kubeadm's packages don't include either tool. Take the release that matches your etcd version from the etcd project's GitHub releases and extract only the two clients:

Terminal
ETCD_VER=v3.7.1
Terminal
curl -L https://github.com/etcd-io/etcd/releases/download/${ETCD_VER}/etcd-${ETCD_VER}-linux-amd64.tar.gz -o /tmp/etcd.tar.gz
Terminal
sudo tar -xzf /tmp/etcd.tar.gz -C /usr/local/bin --strip-components=1 --no-same-owner etcd-${ETCD_VER}-linux-amd64/etcdctl etcd-${ETCD_VER}-linux-amd64/etcdutl
Terminal
etcdctl version && etcdutl version

etcdctl talks to a running etcd over the network; etcdutl works on files. etcd 3.6 removed snapshot restore and snapshot status from etcdctl, so save with etcdctl and check and restore with etcdutl. Older guides set ETCDCTL_API=3; that only matters for etcdctl releases before 3.4.

Take a snapshot

Terminal
sudo install -d -m 700 /var/backups/etcd
Terminal
sudo etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/healthcheck-client.crt --key=/etc/kubernetes/pki/etcd/healthcheck-client.key snapshot save /var/backups/etcd/etcd-$(date +%F-%H%M).db

etcdctl writes to a .part file, syncs it and renames it when complete, then prints Snapshot saved at and the path. The file is readable only by its owner. Kubernetes' guide says taking a snapshot doesn't affect the member's performance.

Give it exactly one endpoint. With several, it stops with snapshot must be requested to one selected node, not multiple. Every member holds the whole keyspace, so on a cluster with several control-plane nodes one healthy member's snapshot is enough.

Check the snapshot

Terminal
sudo etcdutl --write-out=table snapshot status /var/backups/etcd/etcd-2026-10-03-0200.db

It prints the snapshot's hash, revision, total keys and size. A file that fails here won't restore either. Compare the key count with earlier snapshots: a sudden drop means an empty cluster or the wrong one.

Schedule it, encrypted and off the node

A snapshot left on the control-plane node dies with the node. The script below saves and checks a snapshot, encrypts it and a copy of /etc/kubernetes/pki to a GPG public key (set up as in encrypting backups), moves the encrypted files to object storage with rclone, and keeps three days of plain snapshots locally for quick restores.

/usr/local/sbin/etcd-backup.sh
#!/usr/bin/env bash
set -euo pipefail

KEY="A1D07C3D7D97CB70E2285FFD6A1D1AD200953CD5"   # your GPG key's fingerprint
DIR=/var/backups/etcd
STAMP=$(date +%F-%H%M)
PKI=/etc/kubernetes/pki/etcd

install -d -m 700 "$DIR"
etcdctl --endpoints=https://127.0.0.1:2379 --cacert="$PKI/ca.crt" \
  --cert="$PKI/healthcheck-client.crt" --key="$PKI/healthcheck-client.key" \
  snapshot save "$DIR/etcd-$STAMP.db"
etcdutl snapshot status "$DIR/etcd-$STAMP.db"

gpg --batch --encrypt --recipient "$KEY" --output "$DIR/etcd-$STAMP.db.gpg" "$DIR/etcd-$STAMP.db"
tar -czf - -C /etc/kubernetes pki | gpg --batch --encrypt --recipient "$KEY" --output "$DIR/pki-$STAMP.tar.gz.gpg"

rclone move "$DIR" offsite:etcd-backups/cp1 --include "*.gpg"
find "$DIR" -name 'etcd-*.db' -mtime +3 -delete
/etc/systemd/system/etcd-backup.service
[Unit]
Description=etcd snapshot backup

[Service]
Type=oneshot
ExecStart=/usr/local/sbin/etcd-backup.sh
/etc/systemd/system/etcd-backup.timer
[Unit]
Description=Hourly etcd snapshot backup

[Timer]
OnCalendar=hourly
RandomizedDelaySec=5min
Persistent=true

[Install]
WantedBy=timers.target
Terminal
sudo chmod 700 /usr/local/sbin/etcd-backup.sh && sudo systemctl daemon-reload && sudo systemctl enable --now etcd-backup.timer

OnCalendar=hourly fires at the top of each hour, RandomizedDelaySec adds up to five minutes so several nodes don't upload at once, and Persistent=true runs a missed backup as soon as the node is back. systemctl list-timers etcd-backup.timer shows the next run; journalctl -u etcd-backup shows the output. A Kubernetes CronJob could run the same commands, but it needs the scheduler and API server, which are what's broken when you need a fresh snapshot most. Set retention on the bucket (retention policy).

Restore on a kubeadm control plane

Kubernetes' guide says to stop every API server, restore etcd, then start the API servers again, and to restart the scheduler, controller manager and kubelet so nothing acts on stale data. On a single control-plane node, as root:

1. Move the control-plane manifests out of the static Pod folder. The kubelet stops a static Pod when its file disappears; after about 20 seconds crictl ps no longer lists them.

Terminal
mkdir /root/paused && mv /etc/kubernetes/manifests/kube-apiserver.yaml /etc/kubernetes/manifests/kube-controller-manager.yaml /etc/kubernetes/manifests/kube-scheduler.yaml /root/paused/

2. Restore into a new data directory, with the name and peer URL from etcd.yaml (here cp1 and 10.0.0.10):

Terminal
etcdutl snapshot restore /var/backups/etcd/etcd-2026-10-03-0200.db --data-dir /var/lib/etcd-restore --name cp1 --initial-cluster cp1=https://10.0.0.10:2380 --initial-advertise-peer-urls https://10.0.0.10:2380 --bump-revision 1000000000 --mark-compacted

etcd's recovery guide highly recommends --bump-revision with --mark-compacted for Kubernetes. Without them, the revision goes backwards and controllers' caches may not refresh correctly; with them, every watch ends and those caches are invalidated. One billion covers a week-old snapshot if etcd takes fewer than 1,500 writes a second.

3. Point the etcd Pod's etcd-data volume at the new directory and restart the kubelet:

Terminal
sed -i 's#path: /var/lib/etcd$#path: /var/lib/etcd-restore#' /etc/kubernetes/manifests/etcd.yaml
Terminal
systemctl restart kubelet

4. Wait until etcd answers:

Terminal
etcdctl --endpoints=https://127.0.0.1:2379 --cacert=/etc/kubernetes/pki/etcd/ca.crt --cert=/etc/kubernetes/pki/etcd/healthcheck-client.crt --key=/etc/kubernetes/pki/etcd/healthcheck-client.key endpoint health

5. Put the manifests back, check kubectl get nodes and kubectl get pods -A, and run systemctl restart kubelet on every node:

Terminal
mv /root/paused/*.yaml /etc/kubernetes/manifests/

Keep the old /var/lib/etcd until the cluster checks out; it is your way back. With several control-plane nodes, stop the API server on all of them first, then restore the same snapshot on every member, each with its own --name and --initial-advertise-peer-urls and an --initial-cluster listing all members, as in etcd's disaster recovery guide. Start the API servers only when every member is healthy.

Test a snapshot without touching the cluster

Prove a snapshot restores on a spare machine with the same etcd release, never on a control-plane node, where port 2379 is taken. A restore without member flags makes a one-member cluster that plain etcd starts:

Terminal
etcdutl snapshot restore etcd-2026-10-03-0200.db --data-dir /tmp/etcd-test
Terminal
etcd --data-dir /tmp/etcd-test

In a second terminal, list the namespaces the snapshot holds:

Terminal
etcdctl get /registry/namespaces --prefix --keys-only

The API server stores objects under /registry by default, one key per object, so you should see one key per namespace. Values are binary; the keys are enough to show the snapshot holds your cluster. For a full rehearsal, build a throwaway kubeadm cluster at the same versions and run the restore steps. See testing backup restores.

Common errors

  • snapshot must be requested to one selected node, not multiple: pass one endpoint to snapshot save.
  • A certificate error or a timeout when saving: --cacert, --cert or --key is missing or wrong. kubeadm's etcd accepts only clients with a certificate from its CA.
  • data-dir "/var/lib/etcd-restore" not empty or could not be read: etcdutl won't restore over existing data. Pick a new directory.
  • snapshot missing hash but --skip-hash-check=false: the file was copied from member/snap/db, not saved with etcdctl. Add --skip-hash-check, knowing a copied file can miss writes still in the WAL.
  • --mark-compacted required if --revision-bump > 0: use --bump-revision and --mark-compacted together.
  • snapshot restore or snapshot status missing from etcdctl: it is 3.6 or newer. Use etcdutl.
  • Objects created after the snapshot are gone after a restore. That is the snapshot's point in time, so snapshot often enough (RPO and RTO).

Frequently asked questions

Can I back up etcd on EKS, GKE or AKS?
No. The provider runs the control plane, etcd included, and you only reach the API server. Back those clusters up through the Kubernetes API with a tool like Velero.
Is ETCDCTL_API=3 still needed?
Only for etcdctl releases before 3.4. Since 3.4 the v3 API is the default.
Does an etcd snapshot include persistent volumes?
No. It holds the PersistentVolumeClaim and PersistentVolume objects, not the data on the disks. Back up volume data separately.
Should I restore with etcdctl or etcdutl?
etcdutl. etcd 3.6 removed snapshot restore and snapshot status from etcdctl; for snapshots, etcdctl now only has save.
How often should I snapshot etcd?
As often as you can afford to lose cluster changes. A snapshot is a copy of the etcd database, usually small, so hourly is cheap on most clusters.

How this was checked

Commands, limits and prices were checked against these official pages, on October 4, 2026: