How to know when a server backup fails
A backup job can only report the failures it notices, so build three layers: a script that exits non-zero on any error (set -Eeuo pipefail plus size checks), a heartbeat that pings a monitor only after a successful upload, so silence raises the alert, and a separate check that the newest file in your storage is recent and a sensible size. Then break the job on purpose to prove each alert reaches you.
Why backup failures go unnoticed
Most failed backups produce no error message that anyone reads. The usual causes:
| What goes wrong | Why nobody hears about it | What catches it |
|---|---|---|
| Cron output goes nowhere | Cron mails a job's output. With no mail transfer agent it is discarded, and some clouds block mail ports: DigitalOcean blocks 25, 465 and 587 on all Droplets by default. | A heartbeat over HTTPS |
| The exit status is lost in a pipe | pg_dump appdb | gzip > appdb.sql.gz returns gzip's status: 0, with a 20-byte file, when pg_dump cannot connect. | set -o pipefail, a size check |
| The dump is nearly empty | The job reached the wrong database, or an exclude pattern skipped the main tables. Every command succeeded. | A minimum size, a comparison with the last run |
| The disk fills | Without set -e, the script carries on after No space left on device and uploads a truncated file. | set -e, a temporary file name |
| The job never runs | cron is stopped, the server is off, or the crontab was lost in a rebuild. Nothing fails, so nothing reports. | A heartbeat, a freshness check |
| The upload stops working | Expired keys or a deleted bucket. The local file looks fine. | Ping only after the upload; check the bucket |
| The job hangs | It never exits, so it never reports. | A heartbeat; timeout in the cron line |
Several of these never run code that could send an alert. So alerting on errors is only the first layer. The second is a heartbeat: the job reports success to something outside the server, and silence raises the alarm. The third checks the result where it lands.
Make the script fail loudly
A pipeline's exit status is that of its last command, so a failure on the left disappears:
bash -c 'false | gzip > test.gz; echo $?'0gzip succeeded, so the line reports success, and test.gz is a 20-byte file holding nothing. With pipefail, a pipeline returns the status of the last command in it that failed:
bash -c 'set -o pipefail; false | gzip > test.gz; echo $?'1Start every backup script with set -Eeuo pipefail:
| Setting | What it does |
|---|---|
-e | Exit as soon as a command fails. |
-u | Treat an unset variable as an error, so a typo such as $BACKUP_DRI stops the script instead of expanding to nothing. |
-o pipefail | A pipeline fails if any command in it fails, not only the last. |
-E | Let an ERR trap fire inside shell functions too. |
-e has exceptions: it ignores a failure in an if test and on the left of && or ||. A line like upload && echo copied fails quietly and the script carries on. The same rule is what makes [ -s "$FILE" ] || fail "empty dump" safe. When a step's status matters, give it a line of its own.
A backup script that reports its own failure
This script dumps a PostgreSQL database, checks the result, uploads it, and only then reports success. On any failure it writes a line to its log and to standard error, and exits with status 1.
#!/bin/bash
# /usr/local/bin/backup-db.sh: dump, check, upload, then report success.
set -Eeuo pipefail
DB=appdb
DIR=/var/backups/db
LOG=/var/log/backup-db.log
BUCKET=s3://example-backups/db
PING_URL=https://monitor.example.com/ping/backup-db
MIN_BYTES=100000
fail() { echo "$(date -Is) FAILED: $*" | tee -a "$LOG" >&2; exit 1; }
trap 'fail "line $LINENO exited with status $?"' ERR
mkdir -p "$DIR"
FILE="$DIR/$DB-$(date -u +%Y%m%dT%H%MZ).dump"
sudo -u postgres pg_dump -Fc "$DB" > "$FILE.part"
mv "$FILE.part" "$FILE"
size=$(stat -c %s "$FILE")
prev=$(cat "$DIR/last-size" 2>/dev/null || echo 0)
[ "$size" -ge "$MIN_BYTES" ] || fail "$FILE is only $size bytes"
[ "$size" -ge $((prev / 2)) ] || fail "$FILE is $size bytes; the last one was $prev"
aws s3 cp "$FILE" "$BUCKET/" --only-show-errors
echo "$size" > "$DIR/last-size"
find "$DIR" -name "$DB-*" -mtime +3 -delete
echo "$(date -Is) ok $FILE $size bytes" >> "$LOG"
curl -fsS -m 10 --retry 3 -o /dev/null "$PING_URL"- pg_dump writes to a
.partfile that gets its real name only when the dump finishes, so a half-written file never looks complete. - The
ERRtrap catches any failing command and logs its line and exit status. MIN_BYTEScatches an empty or nearly empty dump. Set it to a fraction of your normal size.- The next check fails a dump less than half the size of the last good one. If you deleted data on purpose, delete
/var/backups/db/last-size. - Local copies are pruned only after a successful upload. find rounds ages down to whole days, so
-mtime +3removes files at least four days old. - The heartbeat ping is the last line, reached only when every step before it worked.
We ran it with stand-ins for pg_dump, the AWS CLI and curl: a normal run, a failed dump, a 20-byte dump, a dump half the previous size, a dump written to a full disk and a failed upload. Only the normal run reached the ping. The others exited 1 and logged why, such as FAILED: line 18 exited with status 1 for the failed dump. Schedule it with no output redirect:
SHELL=/bin/sh
PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
[email protected]
17 2 * * * root /usr/local/bin/backup-db.sh
The script keeps its own log and prints nothing when it works, so anything cron captures is an error, and cron mails it to MAILTO. The cron guide covers the rest of the format, and flock and timeout for overlapping or hung runs.
Cron mail: useful, but not enough
Cron mails whatever a job prints to the address in MAILTO, or to the crontab's owner when it is not set; MAILTO="" turns mail off. Three limits:
- Cron mails output, not exit status. A job that fails silently sends nothing.
- It needs a mail transfer agent that relays through a mail provider on a port your cloud allows. Without one, Debian and Ubuntu cron discard the output.
- A job that never starts sends no mail at all.
Tools that print progress on every run would mail you nightly until you stop reading. chronic, from the moreutils package, shows a command's output only when it exits non-zero or crashes:
17 3 * * * chronic /usr/local/bin/backup-files.shProve mail works before relying on it: add * * * * * root echo "cron mail test from $(hostname)" to the cron file, wait for the message, then remove the line.
Heartbeat: alert when success stops arriving
A heartbeat monitor, also called a dead man's switch, gives each job a URL and an expected schedule. The job calls the URL when it succeeds. If no call arrives within the period plus a grace time, the monitor alerts you. That catches every row of the table above, including a stopped cron, a server that is off and a job that hangs.
curl -fsS -m 10 --retry 3 -o /dev/null https://monitor.example.com/ping/backup-db| Flag | What it does |
|---|---|
-f | Treat an HTTP error (400 or above) as a failure: exit code 22, no body. |
-sS | No progress meter, but still print errors. |
-m 10 | Give up after 10 seconds, so a slow monitor cannot hold up the job. |
--retry 3 | Retry up to 3 times on a transient error, such as a timeout or an HTTP 429, 502 or 503. |
-o /dev/null | Discard the response body, so the script stays silent. |
- Period: the job's schedule, 24 hours for a nightly backup.
- Grace: longer than the slowest normal run. If backups take 40 minutes but sometimes two hours, allow three.
- Placement: after the upload, never before. A dump that never left the server is not a backup.
- One monitor per job. A shared URL hides which job stopped.
Hosted cron monitors work this way, and so do the heartbeat checks in many uptime monitoring tools. Whatever you use, it must run somewhere other than the server it watches.
Check the bucket: is the newest backup recent and a sane size?
A heartbeat proves the script reached its last line. A freshness check looks at the storage itself, so it also catches a lifecycle rule that deletes too much or uploads landing under the wrong prefix. Run it from another machine with a read-only key; listing needs only the s3:ListBucket permission.
#!/bin/bash
# /usr/local/bin/check-backup-age.sh: run it on a machine other than the one being backed up.
set -euo pipefail
BUCKET=example-backups
PREFIX=db/
MAX_AGE_HOURS=26
MIN_BYTES=100000
PING_URL=https://monitor.example.com/ping/backup-db-age
newest=$(aws s3api list-objects-v2 --bucket "$BUCKET" --prefix "$PREFIX" --output json \
| jq -r '[.Contents[]?] | max_by(.LastModified) // empty | "\(.LastModified) \(.Size) \(.Key)"')
if [ -z "$newest" ]; then
echo "No backups at all under s3://$BUCKET/$PREFIX" >&2
exit 1
fi
read -r modified size key <<< "$newest"
age_hours=$(( ($(date +%s) - $(date -d "$modified" +%s)) / 3600 ))
if [ "$age_hours" -ge "$MAX_AGE_HOURS" ]; then
echo "Newest backup is $age_hours hours old: s3://$BUCKET/$key" >&2
exit 1
fi
if [ "$size" -lt "$MIN_BYTES" ]; then
echo "Newest backup is only $size bytes: s3://$BUCKET/$key" >&2
exit 1
fi
curl -fsS -m 10 --retry 3 -o /dev/null "$PING_URL"- S3 returns up to 1,000 objects per request; the AWS CLI requests every page and joins them.
[.Contents[]?]is an empty list when the prefix holds nothing,max_by(.LastModified)picks the newest object, and// emptyturns "none" into no output, which the script reports.date -dreads the ISO 8601 timestamp.MAX_AGE_HOURS=26gives a nightly job two hours of slack.- If the AWS CLI fails, for example on bad credentials,
pipefailstops the script with its error.
Against sample listings, a fresh backup reached the ping; a 30-hour-old one printed Newest backup is 30 hours old: s3://example-backups/db/appdb-20261002T0217Z.dump and exited 1, and so did an empty prefix and a 20-byte file.
Filtering with --query instead of jq? Use --output json. With --output text, the AWS CLI runs the query separately on each page of results, so sort_by(Contents, &LastModified)[-1] prints one "newest" object per page.
Put the date in the object name, as the backup script does. Keys are listed in lexicographical order, so the newest backup is also the last key, and anyone browsing the bucket sees the dates at once.
For other S3-compatible storage, add --endpoint-url with the provider's endpoint, as in the Cloudflare R2 and Backblaze B2 guides. Run the check hourly from cron on that other machine. It pings a heartbeat of its own, so if the checker stops, you hear about that too.
Test the alerts by breaking the job
An alert you have never seen fire is an assumption. Break each layer once, on purpose, and note how long each alert takes to reach you:
- A failed dump. Set
DB=appdb_missingand run the script by hand: expect exit status 1 and aFAILEDline in the log. Leave it broken for one scheduled run: expect cron's email, then the heartbeat alert when the grace time runs out. - A shrunken dump. Write a number far above your real dump size into
/var/backups/db/last-sizeand run the script. The size check must stop it before the upload. - A full disk. Point the dump at
/dev/full, where every write fails withNo space left on device, and confirm nothing is uploaded and no ping is sent. - A missed night. Comment out the line in /etc/cron.d/backup-db. The heartbeat should alert the next morning, and the freshness check soon after.
- An empty or stale bucket. Run the freshness check with
PREFIX=nothing-here/, then withMAX_AGE_HOURS=1. Both must exit 1.
Undo each change and check that the next run pings again. Repeat the drill whenever you change mail settings, the monitor or the script. Your grace times also set how much data you can lose before anyone knows; compare them with your RPO.
Alerts tell you a backup is missing. Only a restore tells you a backup works, so schedule a restore test as well.
Frequently asked questions
- Why does my cron backup succeed when the file is empty?
- Usually a pipe.
pg_dump | gzipreturns gzip's exit status, so a dump that fails still exits 0 and leaves a 20-byte gzip file. Addset -o pipefailand check the file's size. - What is a dead man's switch for backups?
- A monitor that expects a ping from the job on a schedule and alerts when the ping does not arrive. Because it alerts on silence, it catches jobs that never ran, hung, or ran on a server that is down.
- Does cron email me when a job fails?
- Only if the job prints something and the server can send mail. Cron mails output, not exit status, to MAILTO or the crontab's owner; without a mail transfer agent the output is discarded.
- How do I find the age of the newest file in an S3 bucket?
- List the prefix with
aws s3api list-objects-v2 --output json, pick the newestLastModifiedwith jq'smax_by, and subtract it from the current time withdate -d. The script in this guide does that and checks the size.
How this was checked
Commands, limits and prices were checked against these official pages, on October 4, 2026:
- Debian bash(1) manual (bash 5.2): set, pipelines, trap
- Debian crontab(5) manual (MAILTO)
- Debian cron(8) manual (output mailing)
- Debian chronic(1) manual (moreutils)
- Debian find(1) manual (-mtime)
- curl man page
- AWS CLI Command Reference: s3api list-objects-v2
- AWS CLI Command Reference: s3 cp
- AWS CLI User Guide: Filtering output
- jq 1.7 manual
- DigitalOcean: Why is SMTP blocked?