How to back up PostgreSQL, MySQL and MongoDB to S3 automatically
Pipe the dump through gzip into aws s3 cp -, which uploads whatever arrives on standard input: pg_dump mydb | gzip | aws s3 cp - s3://my-bucket/mydb.sql.gz. Nothing touches the server's disk, the same pipe works for mysqldump and mongodump --archive, and --endpoint-url points it at R2, B2, Spaces or Wasabi. Unattended, it also needs set -o pipefail, a key that cannot delete, a lifecycle rule, and a way to spot the failed dumps that still land in the bucket looking fine.
How the pipe works
aws s3 cp reads - as standard input. Past 8 MiB it switches to a multipart upload and sends 8 MiB parts while the dump is still running, so the server needs no free disk space for the backup. Install AWS CLI v2 for all users, so cron and the postgres user find it in /usr/local/bin:
curl -fsSL https://awscli.amazonaws.com/v2/install.sh | sudo bash -s -- --systemRun the backups as postgres for PostgreSQL, and as root with the option files from the mysqldump and mongodump guides for MySQL and MongoDB:
pg_dump --no-password mydb | gzip | aws s3 cp - s3://acme-db-backups/db-01/postgresql/mydb-2026-10-04.sql.gz --profile db-backupmysqldump --defaults-extra-file=/etc/mysql/backup.cnf --single-transaction --routines --events mydb | gzip | aws s3 cp - s3://acme-db-backups/db-01/mysql/mydb-2026-10-04.sql.gz --profile db-backupmongodump --config=/root/.mongodump.yaml --archive --gzip | aws s3 cp - s3://acme-db-backups/db-01/mongodb/mongodb-2026-10-04.archive.gz --profile db-backup--no-passwordmakes pg_dump fail at once instead of waiting for a password nobody will type.mongodump --archivewith no file name writes to standard output, and--gzipcompresses it, so there is no gzip stage.
We ran the dump half on PostgreSQL 16.15: a 295 MB database with 2 million rows became a 103,473,624-byte gzipped stream (99 MB) in 21 seconds. Dump options and locking are in the pg_dump guide.
Give the server a key that can upload but not delete
Whoever gets root on the database server gets every key on it. A key that can delete lets them erase every backup, as attackers who encrypt your data try to. Give the server a key that can add and read backups, and leave deleting to a lifecycle rule it cannot change:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ListOwnPrefix",
"Effect": "Allow",
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::acme-db-backups",
"Condition": { "StringLike": { "s3:prefix": ["db-01/*"] } }
},
{
"Sid": "UploadAndReadNoDelete",
"Effect": "Allow",
"Action": ["s3:PutObject", "s3:GetObject", "s3:AbortMultipartUpload"],
"Resource": "arn:aws:s3:::acme-db-backups/db-01/*"
}
]
}s3:PutObjectcovers every step of a multipart upload.s3:AbortMultipartUploadlets the CLI clean up a failed one; it cannot touch finished objects.s3:GetObjectlets the script check the stored size and lets you restore from the server.s3:ListBucketlets it list its own prefix.- No
s3:DeleteObject.s3:PutObjectcan still overwrite a key, so turn on versioning: an overwritten backup survives as an older version, and removing versions needss3:DeleteObjectVersion, which this key lacks.
Create the IAM user, attach the policy and make its access key as in the S3 bucket guide. Wasabi uses AWS-style policies, and its guide shows one without delete rights. Cloudflare R2 and DigitalOcean Spaces have no key level that uploads without deleting; protect those with R2 bucket locks or a second copy, and see the B2 guide for Backblaze's key types. Then store the key for the user that runs the backups:
sudo -iu postgres aws configure --profile db-backupEnter the key ID, the secret and the bucket's region (auto for R2; Backblaze says to leave it blank). More on this in protecting backups from ransomware.
Use R2, B2, Spaces or Wasabi
--endpoint-url sends the same command to another S3-compatible service, at the bucket's regional endpoint:
| Provider | --endpoint-url |
|---|---|
| Cloudflare R2 | https://<ACCOUNT_ID>.r2.cloudflarestorage.com, with the region set to auto |
| Backblaze B2 | https://s3.<region>.backblazeb2.com, such as s3.us-west-004.backblazeb2.com |
| DigitalOcean Spaces | https://<region>.digitaloceanspaces.com, such as fra1.digitaloceanspaces.com |
| Wasabi | https://s3.<region>.wasabisys.com, such as s3.us-east-2.wasabisys.com |
pg_dump --no-password mydb | gzip | aws s3 cp - s3://acme-db-backups/db-01/postgresql/mydb-2026-10-04.sql.gz --profile db-backup --endpoint-url https://s3.us-west-004.backblazeb2.comTo keep the script the same on every provider, add the endpoint to the profile in the backup user's ~/.aws/config instead; a flag on the command line still overrides it:
[profile db-backup]
endpoint_url = https://s3.us-west-004.backblazeb2.comCurrent AWS CLI versions add a checksum to every upload by default, which some S3-compatible services reject. If uploads fail there, add request_checksum_calculation = WHEN_REQUIRED and response_checksum_validation = WHEN_REQUIRED to the profile, as in the Akamai guide.
A failed dump still uploads
If pg_dump dies halfway, gzip still closes its stream properly, the AWS CLI sees the end of its input, completes the upload and exits 0. A pipeline reports the exit status of its last command, so without help the run looks like a success. We killed pg_dump's connection while it was copying the 2-million-row table:
pg_dump: error: Dumping the contents of table "orders" failed: PQgetResult() failed.
pg_dump: detail: Error message from server: FATAL: terminating connection due to administrator command
pg_dump: detail: Command was: COPY public.orders (id, customer_id, total, note, created_at) TO stdout;- Without
set -o pipefailthe pipeline exited 0. With it, it exited 1, so cron and your alerts see the failure. - Either way the stream that came out of the pipe was a complete gzip file of 6 to 7 MB instead of 103 MB, and
gzip -tpassed it. - Restoring that file with
psql -1 -v ON_ERROR_STOP=1also exited 0. It loaded 112,317 of 2,000,000 orders and none of the primary keys or indexes, which pg_dump writes after the data.
So pipefail reports the failure but cannot stop the upload: by the time the shell knows, the object exists. The script below writes a small .done object next to each dump only after the dump, upload and size check all succeed. A dump without one is from a failed run.
The backup script
This runs the pipe as postgres, checks what arrived and writes the marker:
#!/usr/bin/env bash
# Dump one PostgreSQL database, gzip it and stream it to S3 with no local copy.
set -euo pipefail
DB="mydb"
BUCKET="acme-db-backups"
PREFIX="db-01/postgresql"
export AWS_PROFILE="db-backup"
KEY="$PREFIX/$DB-$(date -u +%Y-%m-%dT%H%M%SZ).sql.gz"
TMP=$(mktemp -d)
trap 'rm -rf "$TMP"' EXIT
# Count the bytes as they stream past, to compare with what the bucket stores.
mkfifo "$TMP/tap"
wc -c < "$TMP/tap" > "$TMP/bytes" &
COUNTER=$!
pg_dump --no-password "$DB" \
| gzip \
| tee "$TMP/tap" \
| aws s3 cp - "s3://$BUCKET/$KEY" --only-show-errors
wait "$COUNTER"
SENT=$(cat "$TMP/bytes")
STORED=$(aws s3api head-object --bucket "$BUCKET" --key "$KEY" \
--query ContentLength --output text)
if [ "$SENT" != "$STORED" ]; then
echo "size mismatch for $KEY: sent $SENT bytes, bucket has $STORED" >&2
exit 1
fi
# Written last: a dump with no .done file next to it is from a failed run.
printf '%s\n' "$SENT" | aws s3 cp - "s3://$BUCKET/$KEY.done" --only-show-errors
echo "$(date -u +%FT%TZ) uploaded s3://$BUCKET/$KEY ($SENT bytes)"set -euo pipefailstops at the first failed command, unset variable or failed stage of a pipe. The UTC timestamp in the key means runs never overwrite each other.- The named pipe (
mkfifo) letswc -ccount the bytes on their way to the CLI without a copy on disk.head-objectreturns the storedContentLength, and the two must match.
We ran the script with a local stand-in for the AWS CLI. A normal run stored 103,473,624 bytes and the marker. With pg_dump killed as above, it exited 1 and left a 5.4 MB .sql.gz with no .done beside it. Make it executable and give it a log file:
sudo chmod 755 /usr/local/bin/pg-s3-backupsudo install -o postgres -g postgres -m 640 /dev/null /var/log/pg-s3-backup.logPATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
30 2 * * * postgres flock -n /run/lock/pg-s3-backup.lock /usr/local/bin/pg-s3-backup >> /var/log/pg-s3-backup.log 2>&1The PATH line lets cron find /usr/local/bin/aws, and flock -n skips a run while the last one is still uploading. Wire the exit status to an alert as in backup failure alerts and the cron guide.
For MySQL or MariaDB, run the script as root and replace the pg_dump line with this one:
mysqldump --defaults-extra-file=/etc/mysql/backup.cnf --single-transaction --routines --events --triggers "$DB" \For MongoDB, run it as root, end the key in .archive.gz, and replace both the pg_dump and gzip lines with this one, since mongodump compresses its own archive:
mongodump --config=/root/.mongodump.yaml --archive --gzip \Dumps over 50 GB need --expected-size
With no size to go on, the CLI uses 8 MiB parts, and S3 allows 10,000 parts per upload, so a stream fails at about 84 GB. AWS says to pass --expected-size, in bytes, for streams over 50 GB. It only sets the part size, so overestimate. Size on disk is a safe ceiling for a gzipped dump (ours: 295 MB on disk, 99 MB dumped). Add this before the pipe:
EXPECTED=$(psql -XAtc "SELECT pg_database_size('$DB')")Then add --expected-size "$EXPECTED" to the aws s3 cp line. One AWS re:Post user hit the part limit anyway by taking the size from du, which reports kilobytes; du -b gives bytes. Or raise the part size for every upload: aws configure set s3.multipart_chunksize 64MB --profile db-backup lifts the ceiling to about 670 GB, at the cost of more memory.
Retention: a lifecycle rule, not the script
The server's key cannot delete, so the bucket removes old dumps itself. Apply this once from an admin machine, with admin credentials:
{
"Rules": [
{
"ID": "expire-db-01-dumps",
"Filter": { "Prefix": "db-01/" },
"Status": "Enabled",
"Expiration": { "Days": 30 },
"NoncurrentVersionExpiration": { "NoncurrentDays": 7 },
"AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 2 }
}
]
}aws s3api put-bucket-lifecycle-configuration --bucket acme-db-backups --lifecycle-configuration file://lifecycle.json --profile adminExpirationremoves dumps and their.donemarkers 30 days after upload: about 30 nightly backups.NoncurrentVersionExpirationkeeps an expired or overwritten dump as an old version for 7 more days, in a versioned bucket.AbortIncompleteMultipartUploadclears the parts of uploads that died mid-stream, after a reboot or a killed script. You pay for those parts until they go.
The command replaces the bucket's whole lifecycle configuration, so the file must hold every rule you want. And a rule deletes by age, not by count: if the backups stop, it keeps deleting, and 30 days later the last good dump is gone. Alert on a failed or missing run, not only on errors in the log. Choosing the window is covered in backup retention policy.
Encrypt the dump before it leaves the server
Server-side encryption does not protect the dump from anyone holding a key to the bucket. To encrypt with a public key the server cannot decrypt with, add age to the pipe; the recipients file holds your age1… key:
pg_dump --no-password mydb | gzip | age -R /etc/backup/age-recipients.txt | aws s3 cp - s3://acme-db-backups/db-01/postgresql/mydb-2026-10-04.sql.gz.age --profile db-backupThe private key stays off the server, so restores run where it is: aws s3 cp s3://… - | age -d -i key.txt | gunzip | psql ….
Check a backup and restore it
List the finished dumps, newest last; each .done marker names one:
aws s3 ls s3://acme-db-backups/db-01/postgresql/ --profile db-backup | awk '/\.done$/ {print $4}' | sort | tail -n 3head-object shows a dump's size and upload time without downloading it:
aws s3api head-object --bucket acme-db-backups --key db-01/postgresql/mydb-2026-10-04T023001Z.sql.gz --profile db-backupTo restore PostgreSQL, create an empty database and stream the dump into it, as the postgres user. Nothing touches the disk on the way:
createdb mydb_restoredaws s3 cp s3://acme-db-backups/db-01/postgresql/mydb-2026-10-04T023001Z.sql.gz - --profile db-backup | gunzip | psql -X -v ON_ERROR_STOP=1 -d mydb_restoredFrom a local copy of our dump, gunzip | psql took 25 seconds and brought back all 2,000,000 rows. Since a truncated dump also restores without an error, check the ending too. pg_dump 16.10 and later put a \unrestrict line after the completion comment, so read the last few lines. This prints 1 for a complete dump; our truncated one gave 0:
aws s3 cp s3://acme-db-backups/db-01/postgresql/mydb-2026-10-04T023001Z.sql.gz - --profile db-backup | gunzip | tail -n 5 | grep -c "PostgreSQL database dump complete"A complete mysqldump file ends with -- Dump completed on; restore it with gunzip | mysql -u root -p mydb_restored. mongorestore reads an archive from standard input when --archive has no file name: aws s3 cp s3://….archive.gz - | mongorestore --archive --gzip. Compare row counts afterwards, as in testing a restore; the PostgreSQL restore guide explains psql's errors.
rclone instead of the AWS CLI
rclone streams the same way with rclone rcat remote:acme-db-backups/db-01/postgresql/mydb.sql.gz, which copies standard input to one object. Two limits from its docs: a streamed upload cannot be retried as a whole, and with the default 5 MiB S3 chunk size a stream tops out at 48 GiB unless you raise --s3-chunk-size.
Common errors
| Error | Cause and fix |
|---|---|
An error occurred (AccessDenied) when calling the PutObject operation (CreateMultipartUpload for streams over 8 MiB) | The policy does not cover this key; on AWS the message adds because no identity-based policy allows the s3:PutObject action. Check the bucket and prefix in Resource, with its trailing /*, and the profile. |
An error occurred (InvalidArgument) when calling the UploadPart operation: Part number must be an integer between 1 and 10000, inclusive | The stream outgrew 10,000 parts. Pass --expected-size in bytes, generously, or raise multipart_chunksize. |
An error occurred (EntityTooLarge) | The object would exceed the maximum size: 48.8 TiB on S3, often less elsewhere. Split the backup, for example one dump per database. |
An error occurred (RequestTimeTooSkewed) | The server's clock is too far off. timedatectl shows whether it is synchronized; sudo timedatectl set-ntp true turns sync on. |
An error occurred (SignatureDoesNotMatch) | Usually a wrong or mistyped secret key. On S3-compatible storage, also try the WHEN_REQUIRED checksum settings above. |
An error occurred (403) when calling the HeadObject operation: Forbidden | HEAD responses carry no message. The key lacks s3:GetObject, or the object is missing and the key lacks s3:ListBucket. |
| A dump restores without errors but tables are short or indexes are missing | It came from a failed run. Restore a dump that has a .done marker, and check for PostgreSQL database dump complete. |
Frequently asked questions
- Can I back up a database to S3 without saving the dump to disk first?
- Yes. aws s3 cp - s3://bucket/key uploads standard input as a multipart upload while the dump runs, so pg_dump, mysqldump or mongodump --archive can pipe straight into it.
- Does aws s3 cp work with Cloudflare R2, Backblaze B2, DigitalOcean Spaces and Wasabi?
- Yes. Add --endpoint-url with the provider's regional endpoint, or set endpoint_url in the AWS CLI profile.
- Why does my S3 backup script need set -o pipefail?
- Without it, a pipeline reports only the last command's exit status. A failed pg_dump then looks like a successful upload. Even with it, the truncated object is already in the bucket, so mark good dumps separately.
- How do I delete old database backups from S3 automatically?
- Use a lifecycle rule with an Expiration in days on the backup prefix, set with admin credentials. The server's own key then never needs delete rights.
- When do I need --expected-size?
- For streams over 50 GB, per AWS's documentation. Give the size in bytes; an overestimate only makes the parts bigger.
How this was checked
Commands, limits and prices were checked against these official pages, on October 4, 2026:
- AWS CLI Command Reference: s3 cp (streams and --expected-size)
- AWS CLI Command Reference: s3 ls
- AWS CLI: S3 configuration (multipart_chunksize, multipart_threshold)
- AWS CLI Command Reference: s3api put-bucket-lifecycle-configuration
- AWS CLI User Guide: Installing or updating the AWS CLI
- Amazon S3 User Guide: Multipart upload limits
- Amazon S3 User Guide: Multipart upload API and permissions
- Amazon S3 API Reference: HeadObject
- Amazon S3 API Reference: Error codes
- Amazon S3 User Guide: Troubleshoot access denied (403 Forbidden) errors
- AWS SDKs and Tools Reference Guide: Data integrity protections for Amazon S3
- AWS SDKs and Tools Reference Guide: Service-specific endpoints
- AWS re:Post: How to copy a large file on EC2 from stdin to S3
- AWS CLI source: S3 transfer defaults
- s3transfer source: part size and the 10,000-part limit
- Cloudflare R2: Use the AWS CLI
- Backblaze B2: Call the S3-compatible API
- DigitalOcean Spaces: Use the AWS SDKs
- Wasabi: Service URLs for Wasabi's storage regions
- PostgreSQL documentation: pg_dump
- MySQL 8.4 Reference Manual: mysqldump
- MongoDB Database Tools: mongodump
- MongoDB Database Tools: mongorestore
- age(1) manual page
- rclone rcat
- rclone: Amazon S3 storage providers (chunk size for streams)
- flock(1) manual