VPS Snaps

How to back up files and directories with tar

To back up a directory with tar, run sudo tar -czf /backups/www-2026-10-03.tar.gz -C /var www. That writes one compressed file holding everything under /var/www, with permissions, owners and timestamps. Below: the flags that matter for backups, how to check an archive before you need it, how to restore all or part of it, and how to make small daily incrementals.

9 min readUpdated Tested on Ubuntu 24.04 LTS, GNU tar 1.35

Create a compressed backup

Run tar as root (or with sudo) so it can read every file, including ones only root can open. This creates a gzip-compressed archive of /var/www with today's date in the name:

Terminal
sudo tar -czf /backups/www-$(date +%F).tar.gz -C /var www
FlagWhat it does
-cCreate a new archive.
-zCompress with gzip. GNU tar also has -J (xz) and --zstd (zstd).
-f FILEWrite to FILE. Keep f last in a bundle like -czf, because the next word is taken as the file name.
-C DIRChange to DIR first, so the archive stores www/... rather than var/www/..., and a restore goes wherever you point it.
-vPrint each file as it is added. Handy by hand, noise in a cron log.

$(date +%F) expands to a date such as 2026-10-03, so each run keeps its own file. If you pass an absolute path instead of using -C, tar still works, but warns that it is removing the leading / and stores names like var/www/html/index.php.

-a picks the compression from the file name: tar -caf /backups/www.tar.zst -C /var www writes a zstd archive. When reading, GNU tar detects the compression itself, so -tf and -xf work without -z.

Leave out caches, logs and dependencies

Skipping data you can rebuild makes backups smaller and faster. Patterns are matched against the names inside the archive, so write them relative to the -C directory:

Terminal
sudo tar -czf /backups/www-$(date +%F).tar.gz --exclude='www/html/cache' --exclude='node_modules' --exclude='*.log' -C /var www
  • --exclude='node_modules' has no slash, so it matches that name at any depth.
  • --exclude='www/html/cache' matches that one path. A nested www/html/uploads/cache is kept.
  • An absolute pattern such as --exclude='/var/www/html/cache' matches nothing, because names in the archive have no leading /.
  • Quote every pattern so the shell does not expand * first.
  • For a long list, put one pattern per line in a file and pass --exclude-from=/etc/backup/www.exclude.

Put --exclude before the directories you archive. In GNU tar 1.35 an exclude written after them is ignored: tar still writes the archive, includes the files you meant to skip, and exits with status 2.

Output
tar: The following options were used after non-option arguments.  These options are positional and affect only arguments that follow them.  Please, rearrange them properly.
tar: --exclude ‘*.log’ has no effect
tar: Exiting with failure status due to previous errors

See what is inside an archive

-t lists the contents without extracting anything. Add -v for permissions, owners, sizes and dates:

Terminal
tar -tvf /backups/www-2026-10-03.tar.gz
Output
drwxr-xr-x root/root         0 2026-10-03 02:15 www/
drwxr-xr-x root/root         0 2026-10-03 02:15 www/html/
-rw-r--r-- root/root        17 2026-10-03 02:15 www/html/index.php
drwxr-xr-x www-data/www-data 0 2026-10-03 02:15 www/html/uploads/
-rw-r--r-- www-data/www-data 4 2026-10-03 02:15 www/html/uploads/photo.jpg
-rw-r--r-- root/root        16 2026-10-03 02:15 www/html/wp-config.php

The names in this list are exactly what you pass later to extract a single file. To look for one, pipe the list through grep: tar -tf /backups/www-2026-10-03.tar.gz | grep wp-config.

Verify the backup

A backup you have not restored is a hope. Three checks, from quickest to most thorough.

Can it be read to the end? Listing the whole archive decompresses every byte, and a damaged file makes tar exit non-zero:

Terminal
tar -tzf /backups/www-2026-10-03.tar.gz > /dev/null && echo OK

On a copy with a few corrupted bytes, the same command printed this and exited with status 2:

Output
gzip: stdin: invalid compressed data--format violated
tar: Child returned status 1
tar: Error is not recoverable: exiting now

Does it match the files on disk? -d (--diff) compares every archived file with the live one. Run it right after the backup; anything changed since then is reported, and tar exits with status 1:

Terminal
sudo tar -dzf /backups/www-2026-10-03.tar.gz -C /var
Output
www/html/wp-config.php: Mod time differs
www/html/wp-config.php: Size differs

Did the copy arrive intact? Record a checksum next to the archive, then check it again wherever the copy ends up:

Terminal
cd /backups && sha256sum www-2026-10-03.tar.gz > www-2026-10-03.tar.gz.sha256
Terminal
sha256sum -c www-2026-10-03.tar.gz.sha256
Output
www-2026-10-03.tar.gz: OK
tar exit statusMeaning for a backup job
0Success.
1Some files differ. When creating, a file changed while tar was reading it (file changed as we read it), so its copy may be inconsistent. With -d, differences were found.
2Fatal error. Treat the archive as failed.

-W (--verify) only works on uncompressed archives; with -z, tar refuses with Cannot verify compressed archives. The checks above work on any archive. The real test is still a restore: see how to test a backup restore.

Restore to a different directory

Restore somewhere empty first, look at what you got, and only then copy files into place. -x extracts, and -C sets the target, which must already exist (otherwise tar stops with Cannot open: No such file or directory):

Terminal
sudo mkdir -p /restore/www-2026-10-03
Terminal
sudo tar -xzf /backups/www-2026-10-03.tar.gz -C /restore/www-2026-10-03

Compare the result with the live site. diff -r prints only what differs and exits 0 when nothing does:

Terminal
sudo diff -r /restore/www-2026-10-03/www /var/www

To restore one file, name it exactly as tar -tf shows it. For a folder or pattern, add --wildcards:

Terminal
sudo tar -xzf /backups/www-2026-10-03.tar.gz -C /restore/www-2026-10-03 www/html/wp-config.php
Terminal
sudo tar -xzf /backups/www-2026-10-03.tar.gz -C /restore/www-2026-10-03 --wildcards 'www/html/uploads/*'

--strip-components=1 drops the leading www/ from every name, so the contents land directly in the target directory.

Extracting with -C /var writes straight over the live files without asking. Files created after the backup stay where they are, so you get a mix of old and new. Restore into an empty directory unless overwriting is what you want.

Keep permissions, ownership and extended attributes

tar always records each file's mode, owner and group, by name and by number. What happens on extract depends on who runs it:

  • As root, tar restores the recorded owners and exact permissions by default. In our test, uploads/photo.jpg came back owned by www-data.
  • As an ordinary user, every file is owned by you, and your umask trims the permissions unless you add -p (--preserve-permissions).
  • --no-same-owner makes root extract files as root, which is handy when inspecting an archive from another machine.
  • --numeric-owner uses the stored user and group numbers instead of matching names. Use it when restoring a whole system into a chroot or onto a disk whose /etc/passwd is not the running system's.

ACLs, extended attributes and SELinux labels are not stored by default. If you rely on them, pass --acls --xattrs --selinux both when creating and when extracting. In our test, an attribute archived with --xattrs came back only when --xattrs was given on extract as well.

Incremental backups with --listed-incremental

A full archive every night repeats files that never change. With --listed-incremental (-g), tar keeps a snapshot file recording what it saw, and each later run stores only new and changed files, plus a record of what was deleted. The first run, with no snapshot file yet, is a full (level 0) backup:

Terminal
sudo tar -czf /backups/www-full.tar.gz --listed-incremental=/backups/www.snar -C /var www

Run the same command with a new archive name each day. tar compares against the snapshot file, archives only what changed since the previous run, and updates the snapshot:

Terminal
sudo tar -czf /backups/www-inc-$(date +%F).tar.gz --listed-incremental=/backups/www.snar -C /var www

Start a fresh chain each week by adding --level=0, which empties the snapshot file and makes a full backup again.

To make each daily archive hold everything changed since the full one, so a restore needs only two files: copy the snapshot right after the full run (cp /backups/www.snar /backups/www-full.snar), and before each daily run copy that back to a working file and point --listed-incremental at the copy.

To restore, extract the full archive, then every incremental in order, into the same empty directory. Pass --listed-incremental=/dev/null; tar does not need the snapshot file to extract:

Terminal
sudo tar -xzf /backups/www-full.tar.gz --listed-incremental=/dev/null -C /restore/www
Terminal
sudo tar -xzf /backups/www-inc-2026-10-03.tar.gz --listed-incremental=/dev/null -C /restore/www

Extracting an incremental archive deletes files in the target that did not exist when that archive was made. That is how deletions are replayed, and why you always extract into an empty directory, never over a live one. Lose one archive in the chain and every restore after it is incomplete.

For a weekly full and daily incremental rotation, restoring the chain in the right order, and what breaks a chain, see how to make incremental backups with tar.

Send the archive to another server over SSH

A backup on the same disk as the data is lost with it. With -f -, tar writes the archive to standard output, and ssh carries it to another machine without a temporary file:

Terminal
sudo tar -czf - -C /var www | ssh [email protected] "cat > /backups/www-$(date +%F).tar.gz"

The double quotes make $(date +%F) expand on your server before ssh runs. To restore, reverse the pipe:

Terminal
ssh [email protected] "cat /backups/www-2026-10-03.tar.gz" | sudo tar -xzf - -C /restore/www-2026-10-03

Use ssh -p 2222 for a non-standard port and -i to pick a dedicated key. In a bash script, add set -o pipefail: without it a pipeline's exit status is that of its last command, so a failed tar can look like success. An off-server copy is the core of the 3-2-1 backup rule.

An archive is readable by anyone who can read the bucket. Encrypt it first, or let restic encrypt and deduplicate for you. For which paths to include, see what to back up on a Linux server.

Frequently asked questions

What is the difference between .tar, .tar.gz and .tgz?
A .tar file is an uncompressed bundle of files. .tar.gz and .tgz are the same bundle compressed with gzip; the two suffixes are interchangeable. Create them with -z; when reading, GNU tar detects gzip on its own.
How do I extract a tar.gz file to a specific directory?
Use -C: tar -xzf backup.tar.gz -C /restore/target. The directory must exist first; if it does not, tar stops with Cannot open: No such file or directory.
Does tar preserve file permissions and ownership?
It always records them. When root extracts, owners and exact modes are restored by default. An ordinary user gets files owned by themselves, with the umask applied unless they pass -p.
Why does tar say "file changed as we read it"?
A file was written to while tar was archiving it, so its copy may be inconsistent, and tar exits with status 1. It is common with log files. For databases, never archive the live data files: dump them first with pg_dump or mysqldump.
How do I list the files in a tar.gz without extracting it?
tar -tzf backup.tar.gz prints the names. Add -v for sizes, owners and dates.

How this was checked

The commands were run on Ubuntu 24.04 LTS, GNU tar 1.35 on October 3, 2026. Any that need something this test server does not have, such as a second server, a cloud account or another database engine, were checked against the official pages below instead.

Sources, on October 3, 2026: