VPS Snaps

How to know when a server backup fails

A backup job can only report the failures it notices, so build three layers: a script that exits non-zero on any error (set -Eeuo pipefail plus size checks), a heartbeat that pings a monitor only after a successful upload, so silence raises the alert, and a separate check that the newest file in your storage is recent and a sensible size. Then break the job on purpose to prove each alert reaches you.

11 min readUpdated Checked against official documentation

Why backup failures go unnoticed

Most failed backups produce no error message that anyone reads. The usual causes:

What goes wrongWhy nobody hears about itWhat catches it
Cron output goes nowhereCron mails a job's output. With no mail transfer agent it is discarded, and some clouds block mail ports: DigitalOcean blocks 25, 465 and 587 on all Droplets by default.A heartbeat over HTTPS
The exit status is lost in a pipepg_dump appdb | gzip > appdb.sql.gz returns gzip's status: 0, with a 20-byte file, when pg_dump cannot connect.set -o pipefail, a size check
The dump is nearly emptyThe job reached the wrong database, or an exclude pattern skipped the main tables. Every command succeeded.A minimum size, a comparison with the last run
The disk fillsWithout set -e, the script carries on after No space left on device and uploads a truncated file.set -e, a temporary file name
The job never runscron is stopped, the server is off, or the crontab was lost in a rebuild. Nothing fails, so nothing reports.A heartbeat, a freshness check
The upload stops workingExpired keys or a deleted bucket. The local file looks fine.Ping only after the upload; check the bucket
The job hangsIt never exits, so it never reports.A heartbeat; timeout in the cron line

Several of these never run code that could send an alert. So alerting on errors is only the first layer. The second is a heartbeat: the job reports success to something outside the server, and silence raises the alarm. The third checks the result where it lands.

Make the script fail loudly

A pipeline's exit status is that of its last command, so a failure on the left disappears:

Terminal
bash -c 'false | gzip > test.gz; echo $?'
Output
0

gzip succeeded, so the line reports success, and test.gz is a 20-byte file holding nothing. With pipefail, a pipeline returns the status of the last command in it that failed:

Terminal
bash -c 'set -o pipefail; false | gzip > test.gz; echo $?'
Output
1

Start every backup script with set -Eeuo pipefail:

SettingWhat it does
-eExit as soon as a command fails.
-uTreat an unset variable as an error, so a typo such as $BACKUP_DRI stops the script instead of expanding to nothing.
-o pipefailA pipeline fails if any command in it fails, not only the last.
-ELet an ERR trap fire inside shell functions too.

-e has exceptions: it ignores a failure in an if test and on the left of && or ||. A line like upload && echo copied fails quietly and the script carries on. The same rule is what makes [ -s "$FILE" ] || fail "empty dump" safe. When a step's status matters, give it a line of its own.

A backup script that reports its own failure

This script dumps a PostgreSQL database, checks the result, uploads it, and only then reports success. On any failure it writes a line to its log and to standard error, and exits with status 1.

/usr/local/bin/backup-db.sh
#!/bin/bash
# /usr/local/bin/backup-db.sh: dump, check, upload, then report success.
set -Eeuo pipefail

DB=appdb
DIR=/var/backups/db
LOG=/var/log/backup-db.log
BUCKET=s3://example-backups/db
PING_URL=https://monitor.example.com/ping/backup-db
MIN_BYTES=100000

fail() { echo "$(date -Is) FAILED: $*" | tee -a "$LOG" >&2; exit 1; }
trap 'fail "line $LINENO exited with status $?"' ERR

mkdir -p "$DIR"
FILE="$DIR/$DB-$(date -u +%Y%m%dT%H%MZ).dump"

sudo -u postgres pg_dump -Fc "$DB" > "$FILE.part"
mv "$FILE.part" "$FILE"

size=$(stat -c %s "$FILE")
prev=$(cat "$DIR/last-size" 2>/dev/null || echo 0)
[ "$size" -ge "$MIN_BYTES" ] || fail "$FILE is only $size bytes"
[ "$size" -ge $((prev / 2)) ] || fail "$FILE is $size bytes; the last one was $prev"

aws s3 cp "$FILE" "$BUCKET/" --only-show-errors
echo "$size" > "$DIR/last-size"
find "$DIR" -name "$DB-*" -mtime +3 -delete
echo "$(date -Is) ok $FILE $size bytes" >> "$LOG"
curl -fsS -m 10 --retry 3 -o /dev/null "$PING_URL"
  • pg_dump writes to a .part file that gets its real name only when the dump finishes, so a half-written file never looks complete.
  • The ERR trap catches any failing command and logs its line and exit status.
  • MIN_BYTES catches an empty or nearly empty dump. Set it to a fraction of your normal size.
  • The next check fails a dump less than half the size of the last good one. If you deleted data on purpose, delete /var/backups/db/last-size.
  • Local copies are pruned only after a successful upload. find rounds ages down to whole days, so -mtime +3 removes files at least four days old.
  • The heartbeat ping is the last line, reached only when every step before it worked.

We ran it with stand-ins for pg_dump, the AWS CLI and curl: a normal run, a failed dump, a 20-byte dump, a dump half the previous size, a dump written to a full disk and a failed upload. Only the normal run reached the ping. The others exited 1 and logged why, such as FAILED: line 18 exited with status 1 for the failed dump. Schedule it with no output redirect:

/etc/cron.d/backup-db
SHELL=/bin/sh
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
[email protected]

17 2 * * * root /usr/local/bin/backup-db.sh

The script keeps its own log and prints nothing when it works, so anything cron captures is an error, and cron mails it to MAILTO. The cron guide covers the rest of the format, and flock and timeout for overlapping or hung runs.

Cron mail: useful, but not enough

Cron mails whatever a job prints to the address in MAILTO, or to the crontab's owner when it is not set; MAILTO="" turns mail off. Three limits:

  • Cron mails output, not exit status. A job that fails silently sends nothing.
  • It needs a mail transfer agent that relays through a mail provider on a port your cloud allows. Without one, Debian and Ubuntu cron discard the output.
  • A job that never starts sends no mail at all.

Tools that print progress on every run would mail you nightly until you stop reading. chronic, from the moreutils package, shows a command's output only when it exits non-zero or crashes:

crontab
17 3 * * * chronic /usr/local/bin/backup-files.sh

Prove mail works before relying on it: add * * * * * root echo "cron mail test from $(hostname)" to the cron file, wait for the message, then remove the line.

Heartbeat: alert when success stops arriving

A heartbeat monitor, also called a dead man's switch, gives each job a URL and an expected schedule. The job calls the URL when it succeeds. If no call arrives within the period plus a grace time, the monitor alerts you. That catches every row of the table above, including a stopped cron, a server that is off and a job that hangs.

Terminal
curl -fsS -m 10 --retry 3 -o /dev/null https://monitor.example.com/ping/backup-db
FlagWhat it does
-fTreat an HTTP error (400 or above) as a failure: exit code 22, no body.
-sSNo progress meter, but still print errors.
-m 10Give up after 10 seconds, so a slow monitor cannot hold up the job.
--retry 3Retry up to 3 times on a transient error, such as a timeout or an HTTP 429, 502 or 503.
-o /dev/nullDiscard the response body, so the script stays silent.
  • Period: the job's schedule, 24 hours for a nightly backup.
  • Grace: longer than the slowest normal run. If backups take 40 minutes but sometimes two hours, allow three.
  • Placement: after the upload, never before. A dump that never left the server is not a backup.
  • One monitor per job. A shared URL hides which job stopped.

Hosted cron monitors work this way, and so do the heartbeat checks in many uptime monitoring tools. Whatever you use, it must run somewhere other than the server it watches.

Check the bucket: is the newest backup recent and a sane size?

A heartbeat proves the script reached its last line. A freshness check looks at the storage itself, so it also catches a lifecycle rule that deletes too much or uploads landing under the wrong prefix. Run it from another machine with a read-only key; listing needs only the s3:ListBucket permission.

/usr/local/bin/check-backup-age.sh
#!/bin/bash
# /usr/local/bin/check-backup-age.sh: run it on a machine other than the one being backed up.
set -euo pipefail

BUCKET=example-backups
PREFIX=db/
MAX_AGE_HOURS=26
MIN_BYTES=100000
PING_URL=https://monitor.example.com/ping/backup-db-age

newest=$(aws s3api list-objects-v2 --bucket "$BUCKET" --prefix "$PREFIX" --output json \
  | jq -r '[.Contents[]?] | max_by(.LastModified) // empty | "\(.LastModified) \(.Size) \(.Key)"')

if [ -z "$newest" ]; then
  echo "No backups at all under s3://$BUCKET/$PREFIX" >&2
  exit 1
fi

read -r modified size key <<< "$newest"
age_hours=$(( ($(date +%s) - $(date -d "$modified" +%s)) / 3600 ))

if [ "$age_hours" -ge "$MAX_AGE_HOURS" ]; then
  echo "Newest backup is $age_hours hours old: s3://$BUCKET/$key" >&2
  exit 1
fi
if [ "$size" -lt "$MIN_BYTES" ]; then
  echo "Newest backup is only $size bytes: s3://$BUCKET/$key" >&2
  exit 1
fi

curl -fsS -m 10 --retry 3 -o /dev/null "$PING_URL"
  • S3 returns up to 1,000 objects per request; the AWS CLI requests every page and joins them.
  • [.Contents[]?] is an empty list when the prefix holds nothing, max_by(.LastModified) picks the newest object, and // empty turns "none" into no output, which the script reports.
  • date -d reads the ISO 8601 timestamp. MAX_AGE_HOURS=26 gives a nightly job two hours of slack.
  • If the AWS CLI fails, for example on bad credentials, pipefail stops the script with its error.

Against sample listings, a fresh backup reached the ping; a 30-hour-old one printed Newest backup is 30 hours old: s3://example-backups/db/appdb-20261002T0217Z.dump and exited 1, and so did an empty prefix and a 20-byte file.

Filtering with --query instead of jq? Use --output json. With --output text, the AWS CLI runs the query separately on each page of results, so sort_by(Contents, &LastModified)[-1] prints one "newest" object per page.

Put the date in the object name, as the backup script does. Keys are listed in lexicographical order, so the newest backup is also the last key, and anyone browsing the bucket sees the dates at once.

For other S3-compatible storage, add --endpoint-url with the provider's endpoint, as in the Cloudflare R2 and Backblaze B2 guides. Run the check hourly from cron on that other machine. It pings a heartbeat of its own, so if the checker stops, you hear about that too.

Test the alerts by breaking the job

An alert you have never seen fire is an assumption. Break each layer once, on purpose, and note how long each alert takes to reach you:

  1. A failed dump. Set DB=appdb_missing and run the script by hand: expect exit status 1 and a FAILED line in the log. Leave it broken for one scheduled run: expect cron's email, then the heartbeat alert when the grace time runs out.
  2. A shrunken dump. Write a number far above your real dump size into /var/backups/db/last-size and run the script. The size check must stop it before the upload.
  3. A full disk. Point the dump at /dev/full, where every write fails with No space left on device, and confirm nothing is uploaded and no ping is sent.
  4. A missed night. Comment out the line in /etc/cron.d/backup-db. The heartbeat should alert the next morning, and the freshness check soon after.
  5. An empty or stale bucket. Run the freshness check with PREFIX=nothing-here/, then with MAX_AGE_HOURS=1. Both must exit 1.

Undo each change and check that the next run pings again. Repeat the drill whenever you change mail settings, the monitor or the script. Your grace times also set how much data you can lose before anyone knows; compare them with your RPO.

Alerts tell you a backup is missing. Only a restore tells you a backup works, so schedule a restore test as well.

Frequently asked questions

Why does my cron backup succeed when the file is empty?
Usually a pipe. pg_dump | gzip returns gzip's exit status, so a dump that fails still exits 0 and leaves a 20-byte gzip file. Add set -o pipefail and check the file's size.
What is a dead man's switch for backups?
A monitor that expects a ping from the job on a schedule and alerts when the ping does not arrive. Because it alerts on silence, it catches jobs that never ran, hung, or ran on a server that is down.
Does cron email me when a job fails?
Only if the job prints something and the server can send mail. Cron mails output, not exit status, to MAILTO or the crontab's owner; without a mail transfer agent the output is discarded.
How do I find the age of the newest file in an S3 bucket?
List the prefix with aws s3api list-objects-v2 --output json, pick the newest LastModified with jq's max_by, and subtract it from the current time with date -d. The script in this guide does that and checks the size.

How this was checked

Commands, limits and prices were checked against these official pages, on October 4, 2026: