VPS Snaps

How to back up Elasticsearch or OpenSearch with snapshots

Copying an Elasticsearch or OpenSearch data directory is not a backup; snapshots are the only supported way. Register a snapshot repository (a shared filesystem path listed in path.repo, or an S3 bucket), take snapshots with PUT _snapshot/<repo>/<snapshot> or on a schedule with Elasticsearch SLM or OpenSearch Snapshot Management, and restore with POST _snapshot/<repo>/<snapshot>/_restore. Snapshots are incremental, so taking them often costs little.

10 min readUpdated Checked against official documentation

Why copying the data directory doesn't work

A copy of the nodes' data directories isn't a consistent picture of the cluster at one moment, and Elastic says restoring one can fail with corruption or seem to work while silently losing data. Stopping the nodes first or taking an atomic filesystem snapshot doesn't help, because consistency spans the whole cluster. An LVM snapshot or a disk snapshot of /var/lib/elasticsearch is not a backup either.

Elasticsearch's source is available under AGPLv3, SSPL or the Elastic License 2.0, with Elastic's release builds under the Elastic License; 9.5 is current in October 2026. OpenSearch, derived from Elasticsearch 7.10.2, is Apache 2.0 licensed; 3.9 is current. Their snapshot APIs mostly match, and the differences are noted below.

The examples use an Elasticsearch Debian package install, so each request starts with curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic, which asks for the elastic password. On OpenSearch, use your admin user and your CA file, such as /etc/opensearch/root-ca.pem.

What a snapshot contains

Included by defaultNever included
All regular indices and data streams that are openClosed indices
Cluster state: persistent settings, index templates, ingest pipelines, ILM policies, stored scriptsTransient cluster settings
Feature states (Elasticsearch 7.12+): system indices for security, Kibana and othersRegistered snapshot repositories
Aliases of the included indicesNode config files, the keystore and TLS certificates
  • Incremental: a snapshot copies only index segments that the repository doesn't already hold, since segments never change. A week of hourly snapshots can take little more space than one at the end of the week.
  • Independent: deleting a snapshot through the API removes only the files no other snapshot uses. Never delete files from the repository by hand.
  • Not a single instant: each shard is captured at some point between the snapshot's start and end time. Indexing carries on while it runs.

Back up /etc/elasticsearch (or /etc/opensearch) separately with ordinary file backups, encrypted, since it holds the keystore and private keys.

Register a shared filesystem repository

On a multi-node cluster, mount one shared filesystem, usually NFS, at the same path on every master and data node; a single node can use a local directory. Give it to the service user (opensearch on OpenSearch):

Terminal
sudo mkdir -p /mnt/es-snapshots && sudo chown elasticsearch:elasticsearch /mnt/es-snapshots

List the path in the config on every master and data node:

/etc/elasticsearch/elasticsearch.yml
path:
  repo:
    - /mnt/es-snapshots

path.repo is a static setting, so restart each node, one at a time on a running cluster. OpenSearch reads the same setting from /etc/opensearch/opensearch.yml. Then register the repository; its location must sit under a path.repo entry:

Terminal
curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic -X PUT "https://localhost:9200/_snapshot/fs_backups" -H "Content-Type: application/json" -d '{"type": "fs", "settings": {"location": "/mnt/es-snapshots/fs_backups"}}'

Registration checks the repository on every master and data node. To run that check again later:

Terminal
curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic -X POST "https://localhost:9200/_snapshot/fs_backups/_verify"

On NFS, every node must run the service under the same numeric UID and GID, or verification fails with permission errors even though the user names match.

Or use an S3 repository

A bucket at another provider gives you an off-site copy with no extra step. Since Elasticsearch 8.0 the s3 repository type is built in; 7.x needed the repository-s3 plugin. OpenSearch still does: install it on every node from the OpenSearch home directory, then restart:

Terminal
sudo ./bin/opensearch-plugin install repository-s3

Store the bucket key in each node's keystore (on OpenSearch, ./bin/opensearch-keystore):

Terminal
sudo /usr/share/elasticsearch/bin/elasticsearch-keystore add s3.client.default.access_key
Terminal
sudo /usr/share/elasticsearch/bin/elasticsearch-keystore add s3.client.default.secret_key

Both settings are reloadable: call POST _nodes/reload_secure_settings instead of restarting. For storage other than AWS, set s3.client.default.endpoint to the full https:// URL and s3.client.default.region to the region the service expects, in the YAML config. Then register it:

Terminal
curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic -X PUT "https://localhost:9200/_snapshot/s3_backups" -H "Content-Type: application/json" -d '{"type": "s3", "settings": {"bucket": "acme-es-snapshots", "base_path": "prod"}}'
  • The bucket must exist first. Elastic's minimum policy grants s3:ListBucket, s3:GetBucketLocation, s3:ListBucketMultipartUploads and s3:ListBucketVersions on the bucket, and s3:GetObject, s3:PutObject, s3:DeleteObject, s3:AbortMultipartUpload and s3:ListMultipartUploadParts on its objects. See setting up an S3 bucket.
  • Never put lifecycle expiry, Glacier transitions or bucket replication on a repository bucket. Elastic warns all three can make the repository permanently unreadable. Let SLM retention delete old snapshots.
  • Elastic requires S3-compatible storage to behave exactly like AWS S3. The repository analysis API (POST _snapshot/s3_backups/_analyze) tests that, but Elastic says a realistic run needs max_total_data_size of at least 1tb, which you pay to write.

Take a snapshot by hand

Terminal
curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic -X PUT "https://localhost:9200/_snapshot/fs_backups/manual-$(date +%F)?wait_for_completion=true"

With no body, this saves every open index and data stream plus the cluster state; a body such as {"indices": "orders,logs-*", "include_global_state": false} narrows it. wait_for_completion=true holds the request until the snapshot finishes; on large clusters, leave it off and poll. Names must be unique in the repository. Check the result:

Terminal
curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic "https://localhost:9200/_snapshot/fs_backups/_all?pretty"

Each snapshot's state should be SUCCESS. PARTIAL means some shards weren't saved, which can only happen with partial: true. GET _snapshot/_status shows shard-by-shard progress of a running snapshot. OpenSearch won't start a snapshot while another is in progress.

Schedule snapshots: SLM and Snapshot Management

On Elasticsearch, a snapshot lifecycle management (SLM) policy takes snapshots on a schedule and deletes old ones:

slm-nightly.json
{
  "schedule": "0 30 1 * * ?",
  "name": "<nightly-snap-{now/d}>",
  "repository": "fs_backups",
  "config": {
    "indices": "*",
    "include_global_state": true
  },
  "retention": {
    "expire_after": "30d",
    "min_count": 5,
    "max_count": 50
  }
}
  • schedule uses Elasticsearch's cron, which starts with a seconds field: 0 30 1 * * ? is 01:30 UTC daily. Keep node clocks in sync.
  • name uses date math, so each snapshot gets the day's date, and SLM appends a UUID so names never clash.
  • retention deletes snapshots older than 30 days, but always keeps at least 5 and never more than 50. It only touches this policy's snapshots, and runs as a separate task at slm.retention_schedule, 01:30 UTC by default.
Terminal
curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic -X PUT "https://localhost:9200/_slm/policy/nightly-snapshots" -H "Content-Type: application/json" -d @slm-nightly.json

Run it once now with POST _slm/policy/nightly-snapshots/_execute, and check it with GET _slm/policy/nightly-snapshots, which shows the next run and the last success and failure. Alert on that last failure, as in backup failure alerts.

OpenSearch's equivalent is Snapshot Management, part of the Index Management plugin. Its cron has five fields and its own time zone:

sm-nightly.json
{
  "description": "Nightly snapshots, kept 30 days",
  "creation": {
    "schedule": { "cron": { "expression": "30 1 * * *", "timezone": "UTC" } },
    "time_limit": "1h"
  },
  "deletion": {
    "schedule": { "cron": { "expression": "0 2 * * *", "timezone": "UTC" } },
    "condition": { "max_age": "30d", "min_count": 5, "max_count": 50 }
  },
  "snapshot_config": {
    "repository": "fs_backups",
    "indices": "*,-.opendistro_security"
  }
}
Terminal
curl --cacert /etc/opensearch/root-ca.pem -u admin -X POST "https://localhost:9200/_plugins/_sm/policies/nightly" -H "Content-Type: application/json" -d @sm-nightly.json

Snapshots are named nightly-<date>-<random>. OpenSearch recommends leaving the .opendistro_security index out of snapshots, as above. Failed runs are retried up to three times; GET _plugins/_sm/policies/nightly/_explain shows the last result and any error.

Restore indices without overwriting live ones

A restore can't overwrite an open index. You have three options: restore under a new name, delete the index first, or close it first (only if the snapshot's copy has the same number of primary shards). Renaming is safest. Take the snapshot name from the listing above:

restore.json
{
  "indices": "orders",
  "rename_pattern": "(.+)",
  "rename_replacement": "restored-$1",
  "include_aliases": false
}
Terminal
curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic -X POST "https://localhost:9200/_snapshot/fs_backups/<snapshot>/_restore" -H "Content-Type: application/json" -d @restore.json
  • rename_pattern is a regular expression; $1 puts the original name back after restored-, giving restored-orders.
  • include_aliases: false stops the restored copy from claiming the live index's aliases.
  • Cluster state is restored only with include_global_state: true, which removes and replaces all persistent settings, index templates, ingest pipelines and ILM policies. Use it only to rebuild a whole cluster.
  • On Elasticsearch 8+, system indices come back only through feature_states. Restoring the security feature state overwrites the indices used for logins.
  • On OpenSearch with the Security plugin, you can't restore global state or .opendistro_security; send "include_global_state": false and leave that index out.

Cluster health stays yellow while primaries recover; follow it with GET _cluster/health and GET restored-orders/_recovery. Then compare the data, and either reindex restored-orders back to orders or point your alias at it.

Version compatibility

You can never restore a snapshot into an older version than the one that took it, and every restored index must be compatible with the target. For Elasticsearch, per Elastic's table:

Index created inRestores as a normal index intoInto 9.x
7.x7.x (same or newer minor) and 8.xOnly read-only, as archive indices or searchable snapshots
8.x8.x (same or newer minor)Yes
9.x9.x (same or newer minor)Yes
  • OpenSearch snapshots move forward by one major version only: a 2.x snapshot restores into 3.x, a 1.x snapshot doesn't.
  • Once a newer version writes to a repository, older versions may not read it. Snapshot before every upgrade; a pre-upgrade snapshot still restores into the old version if you roll back.
  • Elastic warns that Kibana data restored across a major version can fail when Kibana starts, even when Elasticsearch accepted the restore.

Copy the repository off the server

A filesystem repository on the cluster's own servers dies with them. Either use an S3 repository at another provider, or copy the repository directory elsewhere. Elastic allows that copy only while nothing writes to the repository, so switch it to read-only around the copy:

/usr/local/bin/es-repo-mode.sh
#!/bin/sh
# Usage: es-repo-mode.sh true|false
set -eu
case "$1" in true|false) ;; *) echo "usage: $0 true|false" >&2; exit 2 ;; esac
curl -sS --fail -o /dev/null --cacert /etc/elasticsearch/certs/http_ca.crt --netrc-file /root/.es-netrc -X PUT "https://localhost:9200/_snapshot/fs_backups" -H "Content-Type: application/json" -d '{"type": "fs", "settings": {"location": "/mnt/es-snapshots/fs_backups", "readonly": '"$1"'}}'

/root/.es-netrc holds one line, machine localhost login snapshot-admin password <password>, for a user whose role has the manage cluster privilege, readable only by root. --fail makes the script exit with an error if Elasticsearch refuses the change. Then open a window between snapshot runs and run your copy inside it, for example a restic backup of /mnt/es-snapshots/fs_backups to another provider at 03:00, per the 3-2-1 rule:

/etc/cron.d/es-repo-window
50 2 * * * root /usr/local/bin/es-repo-mode.sh true
0 4 * * * root /usr/local/bin/es-repo-mode.sh false

Cron runs in the server's time zone and SLM in UTC: keep snapshot creation and retention out of the window, and make the window longer than the copy takes. When restoring a copied repository, put every file back before you register it, then register it read-only and check it with POST _snapshot/fs_backups/_verify_integrity (Elasticsearch).

Test a restore and fix common errors

Build a scratch cluster on the same version or newer, register the repository there with "readonly": true (so two clusters never write to it), and restore everything. Compare document counts with production:

Terminal
curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic "https://localhost:9200/_cat/indices?v=true&h=index,docs.count&s=index"

Time it too: that is your real recovery time (RPO and RTO, testing a restore). Common restore errors:

  • open index with same name already exists: rename on restore, or delete or close the index first.
  • Cannot restore index [...] with [x] shards from a snapshot of index [...] with [y] shards: you closed an index whose shard count differs; rename instead.
  • system indices can only be restored as part of a feature state: drop the system index from indices and use feature_states.
  • Repository verification fails: the path isn't under path.repo on every node, the service user can't write it, or NFS UIDs differ.

Frequently asked questions

Can I back up Elasticsearch by copying /var/lib/elasticsearch?
No. Elastic says filesystem copies of data directories, even with the nodes stopped or from atomic snapshots, are not supported backups and may restore with silent data loss. Use snapshots.
Are Elasticsearch snapshots incremental?
Yes. Each snapshot copies only segments the repository doesn't already have, yet each one can be restored on its own. Delete old snapshots through the API so shared files are kept.
Do I need the repository-s3 plugin?
Not on Elasticsearch 8.0 or later, where s3 is built in. On Elasticsearch 7.x and on OpenSearch, install repository-s3 on every node and restart.
Can I restore a snapshot to an older Elasticsearch version?
No. A snapshot restores only into the same or a newer version, and each index must be compatible with the target, such as 8.x indices into 9.x.
What is the OpenSearch version of SLM?
Snapshot Management, configured through _plugins/_sm/policies, with a creation schedule, a deletion schedule and conditions such as max_age and max_count.

How this was checked

Commands, limits and prices were checked against these official pages, on October 4, 2026: