How to back up and restore a k3s cluster
What to back up in k3s depends on its datastore. A single server uses SQLite unless told otherwise, and you copy /var/lib/rancher/k3s/server/db/ safely; embedded etcd takes its own snapshots, twice a day by default, which k3s server --cluster-reset --cluster-reset-restore-path restores. Either way, keep the server token from /var/lib/rancher/k3s/server/token: without it, a restore can't decrypt the cluster's certificates and keys.
Find out which datastore you have
| Datastore | When k3s uses it | How to back it up |
|---|---|---|
| Embedded SQLite | The default: no other datastore set and no etcd files on disk. One server only. | Copy the database safely (below) |
| Embedded etcd | Started with --cluster-init, joined to such a cluster, or etcd files found on disk | k3s etcd-snapshot and scheduled snapshots (below) |
| External database | --datastore-endpoint points at MySQL, MariaDB, PostgreSQL or etcd | That database's own backups, such as pg_dump or mysqldump |
On a server, list the database folder. state.db with -shm and -wal files means SQLite; an etcd folder means embedded etcd:
sudo ls /var/lib/rancher/k3s/server/dbAn external endpoint is set in the config file or in the service files the install script wrote:
sudo grep -rsi 'datastore.endpoint' /etc/rancher/k3s /etc/systemd/system/k3s.service /etc/systemd/system/k3s.service.envWhat to keep besides the datastore
| File | Why |
|---|---|
/var/lib/rancher/k3s/server/token | k3s encrypts the cluster's CA certificates and keys inside the datastore with a key derived from this token. Restore it, or pass it as --token, or the backup is unusable. |
/etc/rancher/k3s/config.yaml and config.yaml.d/ | Your server flags. Some, such as --cluster-cidr, must match on every server. |
/etc/systemd/system/k3s.service and k3s.service.env | Flags and K3S_ variables given to the install script are saved here |
/etc/rancher/node/password | The node's registration password. A node that rejoins under its old name must present the same one. |
/var/lib/rancher/k3s/server/manifests/ | Manifests you placed there for k3s to apply at startup |
/var/lib/rancher/k3s/storage/ | Data of volumes from the default local-path StorageClass (see below) |
k3s's documentation says the server token gives full administrator access to the cluster, and that the token plus a snapshot lets anyone extract the cluster CA's private keys and every Secret. Store them apart and encrypted. After k3s token rotate, snapshots taken earlier still need the old token, so keep it until they expire.
Back up a SQLite server
k3s's documentation says a copy of /var/lib/rancher/k3s/server/db/ is the backup. k3s runs SQLite in WAL mode, though, so recent writes sit in state.db-wal until a checkpoint, and a file copy made while k3s writes can miss them or be torn. Stop k3s for the copy, or let SQLite make it.
Stopping k3s. Your containers keep running when the k3s service stops; only the API and controllers pause while tar runs:
sudo systemctl stop k3ssudo tar -czf /var/backups/k3s-server-$(date +%F).tar.gz -C /var/lib/rancher/k3s/server db tokensudo systemctl start k3sWith SQLite's backup command, k3s keeps running. The sqlite3 shell is not part of k3s; install it from your distribution (details):
sudo sqlite3 /var/lib/rancher/k3s/server/db/state.db ".backup '/var/backups/k3s-state-$(date +%F).db'"sudo sqlite3 /var/backups/k3s-state-2026-10-04.db 'PRAGMA integrity_check;'It should print ok. Copy the token with it. To get k3s's snapshot tools instead, convert the server to embedded etcd: k3s's documentation says restarting it with --cluster-init is enough, and warns that etcd can be slow on SD cards and other slow disks.
Take an etcd snapshot
On a server with embedded etcd, ask the running k3s for a snapshot, for example before an upgrade:
sudo k3s etcd-snapshot save --name pre-upgradek3s logs Snapshot pre-upgrade-<node>-<unix-time> saved. and writes the file to /var/lib/rancher/k3s/server/db/snapshots/. --name sets the prefix (default on-demand), and --etcd-snapshot-compress stores it as a .zip. The command sends the request to the local k3s on port 6443 with the token file, so k3s must be running. On-demand snapshots are never deleted automatically. List them, and prune all but the newest three:
sudo k3s etcd-snapshot lssudo k3s etcd-snapshot prune --name pre-upgrade --etcd-snapshot-retention 3ls shows each snapshot's location (file:// or s3://), size and date. kubectl get etcdsnapshotfile lists the snapshots of every server at once.
Schedule snapshots and copy them to S3
With embedded etcd, scheduled snapshots are on by default: at 00:00 and 12:00 system time, 5 kept per server, on that server's own disk. A lost disk takes them with it. Add a schedule and a bucket to the config file on every server:
etcd-snapshot-schedule-cron: "0 */6 * * *"
etcd-snapshot-retention: 28
etcd-snapshot-compress: true
etcd-s3: true
etcd-s3-endpoint: "s3.eu-central-1.amazonaws.com"
etcd-s3-region: "eu-central-1"
etcd-s3-bucket: "acme-k3s-snapshots"
etcd-s3-folder: "prod"
etcd-s3-retention: 84
etcd-s3-access-key: "<access-key-id>"
etcd-s3-secret-key: "<secret-access-key>"etcd-snapshot-schedule-cron: when to snapshot, in cron syntax.0 */6 * * *is every six hours, four a day.etcd-snapshot-retention: snapshots kept on each server's disk. 28 at four a day is seven days.etcd-snapshot-compress: zip each snapshot.etcd-s3: upload every snapshot, scheduled or on demand, to the bucket as well.etcd-s3-endpoint: host name only, nohttps://(defaults3.amazonaws.com). For other S3-compatible storage, give its endpoint;etcd-s3-bucket-lookup-type: pathhelps providers without bucket subdomains.etcd-s3-bucketandetcd-s3-folder: the bucket must already exist. Give each cluster its own folder.etcd-s3-retention: snapshots kept in the folder for the whole cluster, not per server. Each server uploads every cycle, so k3s suggests servers times cycles: 3 servers x 28 = 84.etcd-s3-access-keyandetcd-s3-secret-key: a key limited to this bucket (bucket setup).
At the 5 MB per snapshot of k3s's own example, 84 copies take about 420 MB. The file now holds a secret key, so make it readable by root only, then restart k3s on each server, one at a time:
sudo chmod 600 /etc/rancher/k3s/config.yamlsudo systemctl restart k3sk3s can also read S3 settings from a Secret (etcd-s3-config-secret), but not during a restore, when the API is down, so keep the values at hand.
Restore a snapshot
On a single server:
sudo systemctl stop k3ssudo k3s server --cluster-reset --cluster-reset-restore-path=/var/lib/rancher/k3s/server/db/snapshots/pre-upgrade-k3s-1-1791079200k3s moves the current database to db/etcd-old-<timestamp>/, restores the snapshot, removes every other etcd member, and stops with Managed etcd cluster membership has been reset, restart without --cluster-reset flag now. Then:
sudo systemctl start k3sWhen the config file has S3 settings, --cluster-reset-restore-path takes only the snapshot's name and k3s downloads it; to use a local file then, add --etcd-s3=false and give the full path. Otherwise, pass --etcd-s3, --etcd-s3-bucket and the other S3 flags on the command line.
With three servers: stop k3s on all of them, run the reset on one and start it. Then, on each of the others, delete the old database and start k3s so it rejoins:
sudo rm -rf /var/lib/rancher/k3s/server/db/sudo systemctl start k3sRun that rm -rf only on the servers that rejoin, never on the one you restored. A restore takes the cluster back to the snapshot's moment: anything created after it is gone (RPO and RTO).
Restore on a new server
k3s restores a snapshot on the same version or a higher minor one. Install k3s without starting it, put back the config file, and restore with the old token (or set token: in the config file; a different token there stops k3s from starting):
curl -sfL https://get.k3s.io | INSTALL_K3S_VERSION=v1.36.5+k3s1 INSTALL_K3S_SKIP_START=true sh -sudo k3s server --cluster-reset --cluster-reset-restore-path=/root/pre-upgrade-k3s-1-1791079200 --token=<backed-up-token>sudo systemctl start k3sThe snapshot still lists the old machines as nodes; delete those that no longer exist with kubectl delete node <name>. Pods with local-path volumes are bound to the node that held them, so on a one-server cluster keep the old node name instead: give the new machine the same hostname (or node-name in the config file), restore /etc/rancher/node/password before the first start, and put /var/lib/rancher/k3s/storage back.
A SQLite server needs no reset. With k3s stopped (on a new machine, installed as above), move any old db folder aside, put state.db in /var/lib/rancher/k3s/server/db/ and the token at /var/lib/rancher/k3s/server/token, then start k3s.
What the datastore doesn't hold
A snapshot or SQLite copy holds Kubernetes objects such as Deployments, Secrets and PersistentVolumeClaims, not the bytes in your volumes or your images.
- local-path volumes live on one node's disk under
/var/lib/rancher/k3s/storage/, a directory per volume named<pv>_<namespace>_<pvc>. Archive that directory; for a database on it, take a dump, because a file copy of a running database can be inconsistent. - Velero backs up workloads and volume data through the API (Velero guide). Its File System Backup skips hostPath volumes with
is a hostPath volume which is not supported for pod volume backup, skipping. Since v1.29.4 and v1.30, k3s'slocal-pathStorageClass createslocalvolumes, which Velero supports; volumes created before that stay hostPath. - etcd itself: a k3s snapshot is a standard etcd snapshot, so
etcdutl snapshot statusreads it once unzipped. See backing up etcd.
Check which kind each volume is:
kubectl get pv -o custom-columns=NAME:.metadata.name,LOCAL:.spec.local.path,HOSTPATH:.spec.hostPath.pathTest a restore
Restore on a spare machine, never on a live server: follow the new-server steps with a copy of the snapshot and token, then check that namespaces and workloads came back:
sudo k3s kubectl get namespacessudo k3s kubectl get deployments,statefulsets -AA restored cluster starts the workloads it holds, CronJobs included, with your real Secrets. Block the test machine's outbound traffic to production databases, payment APIs and mail before its first start.
Time the run from install to healthy pods; that is your recovery time. See testing backup restores.
Common errors
bootstrap data already found and encrypted with different token: the token isn't the one the datastore was created with. Use the backed-up token and remove any othertoken:from the config file.no bootstrap data found in datastore - check server token value and verify datastore integrity: a wrong token, or a wrong or empty database./var/lib/rancher/k3s/server/token does not exist, please pass --token to complete the restoration: a restore on a new machine needs--token.cannot perform cluster-reset while server URL is set - remove server from configuration before resetting: this server joined withserver:in its config. Remove that line first.Managed etcd cluster membership was previously reset, please remove the cluster-reset flag and start k3s normally: a reset already ran. Start k3s normally, or delete/var/lib/rancher/k3s/server/db/reset-flagto reset again.bucket acme-k3s-snapshots does not exist: k3s doesn't create the bucket. Create it, or fix the name or endpoint.failed to validate server configuration: critical configuration value mismatch: a rejoining server's flags differ from the cluster's. Copy the config file from the restored server.
Frequently asked questions
- Does k3s back itself up automatically?
- With embedded etcd, yes: snapshots at 00:00 and 12:00, 5 kept per server, on the server's own disk. With SQLite, no; copy the database yourself.
- Can I use k3s etcd-snapshot with SQLite?
- No, it needs embedded etcd. Restarting a single SQLite server with --cluster-init converts it.
- Do I need the same k3s version to restore a snapshot?
- No. k3s restores a snapshot on the same version or a higher minor version.
- Where are k3s etcd snapshots stored?
- In /var/lib/rancher/k3s/server/db/snapshots/ on each server, and in your bucket under etcd-s3-folder when S3 is set up.
- Does a k3s snapshot include persistent volume data?
- No. It holds the volume objects, not their contents. Back up /var/lib/rancher/k3s/storage or use Velero.
How this was checked
Commands, limits and prices were checked against these official pages, on October 4, 2026:
- K3s documentation: Cluster Datastore
- K3s documentation: Backup and Restore
- K3s documentation: etcd-snapshot
- K3s documentation: server
- K3s documentation: token
- K3s documentation: High Availability Embedded etcd
- K3s documentation: Configuration Options
- K3s documentation: Environment Variables
- K3s documentation: Stopping K3s
- K3s documentation: Volumes and Storage
- K3s documentation: Architecture (node registration)
- K3s documentation: Managing Packaged Components
- k3s source: etcd-snapshot command flags
- k3s source: cluster reset checks
- k3s source: reset flag
- k3s source: bootstrap data and token errors
- k3s source: etcd snapshots
- k3s source: S3 client
- k3s source: bundled local-path StorageClass
- k3s v1.36.5+k3s1 release (stable channel)
- kine source: SQLite driver defaults
- Local Path Provisioner README (volume types)
- Velero v1.18 docs: File System Backup
- Velero source: pod volume backupper (hostPath volumes)
- etcd documentation v3.7: Disaster recovery (snapshot status)