VPS Snaps

PostgreSQL incremental backups with pg_basebackup

From PostgreSQL 17, pg_basebackup --incremental=<previous>/backup_manifest copies only the blocks that changed since an earlier backup, once summarize_wal = on. An incremental backup cannot be started on its own: pg_combinebackup merges the full backup and every incremental after it, oldest first, into an ordinary data directory.

9 min readUpdated Checked against official documentation

How incremental backups work

Our test server runs PostgreSQL 16, so the commands here were checked against the PostgreSQL 18 documentation and source, not run. A background process, the WAL summarizer, reads the write-ahead log and records which blocks of which files changed, in summary files under pg_wal/summaries. For an incremental backup, pg_basebackup uploads the earlier backup's manifest. The server checks that summaries cover the WAL since that backup started, then sends other files whole and, for table and index files, INCREMENTAL. files holding only the changed 8 kB blocks.

Each incremental depends on the backup it was based on, back to a full backup: a chain. Lose one link and every later backup is useless, and PostgreSQL does not track chains for you. The payoff comes with a large database that changes in a small part. An example, with assumed figures rather than measurements:

200 GB cluster, 4 GB of blocks changed per dayStorage for 14 daily restore points
A full backup every day14 × 200 = 2,800 GB
Weekly full plus daily incrementals, two chains kept2 × (200 + 6 × 4) = 448 GB

One changed row sends its whole 8 kB block, so updates scattered across a table shrink the savings. The docs say incrementals are barely smaller than fulls when all of a large database is heavily modified, and that a small database is simpler to back up in full.

Turn on WAL summarization

/etc/postgresql/18/main/conf.d/summarize.conf
summarize_wal = on
wal_summary_keep_time = '10d'
  • summarize_wal is off by default and needs only a reload. It works on a primary or a standby, but the server will not start with it while wal_level = minimal.
  • wal_summary_keep_time is how long summaries are kept: 10 days by default, minutes when no unit is given, never deleted at 0. Keep it comfortably longer than the gap between a backup and the next incremental built on it.
Terminal
sudo systemctl reload postgresql@18-main
Terminal
sudo -u postgres psql -c "SELECT * FROM pg_get_wal_summarizer_state()"

summarizer_pid should be a process ID and summarized_lsn should move forward as WAL is written. pg_available_wal_summaries() lists each summary file and the WAL range it covers.

Summaries must cover every WAL position from the start of the earlier backup to the start of the new one. Turn summarization on before taking the full backup a chain will start from.

Take a full backup, then incrementals

Terminal
sudo install -d -o postgres -g postgres -m 700 /var/backups/pg
Terminal
sudo -u postgres pg_basebackup -D /var/backups/pg/2026-10-04_0215-full -c fast

This is a normal base backup: -D is a new directory, -c fast checkpoints at once, and the defaults are plain format plus -X stream, which saves the WAL written during the backup. Replication access and the other basics are in the point-in-time recovery guide. The backup_manifest it writes lists every file with its size and checksum. Keep plain format for chains: pg_combinebackup reads backup directories, so tar backups would have to be unpacked first. The next day:

Terminal
sudo -u postgres pg_basebackup -D /var/backups/pg/2026-10-05_0215-incr -c fast --incremental=/var/backups/pg/2026-10-04_0215-full/backup_manifest

--incremental (-i) names the manifest of any earlier backup of the same server. Point each one at the newest backup's manifest to keep every backup small, at the cost of a longer chain to restore. Or keep pointing at the full backup: each incremental then grows through the week, but a restore needs only two backups.

Automate a weekly chain

This script takes a full backup on Sundays (or when no chain exists) and an incremental on the newest backup on other days. It verifies each one before giving it its final name, then deletes whole chains older than the newest two full backups:

/usr/local/bin/pg-incremental-backup
#!/usr/bin/env bash
# Weekly full base backup (Sundays), incremental backups on the other days.
# Keeps the newest KEEP_CHAINS chains. Run as postgres. FULL=1 starts a new chain.
set -euo pipefail

BASE="/var/backups/pg"
BIN="/usr/lib/postgresql/18/bin"
KEEP_CHAINS=2

backups() {
  find "$BASE" -mindepth 1 -maxdepth 1 -type d \( -name '*-full' -o -name '*-incr' \) | sort
}

LAST=$(backups | tail -n 1)
STAMP=$(date +%Y-%m-%d_%H%M)

if [ "${FULL:-0}" = 1 ] || [ "$(date +%u)" = 7 ] || [ -z "$LAST" ]; then
  DEST="$BASE/$STAMP-full"
  ARGS=(-D "$DEST.partial" -c fast)
else
  DEST="$BASE/$STAMP-incr"
  ARGS=(-D "$DEST.partial" -c fast --incremental="$LAST/backup_manifest")
fi
trap 'rm -rf "$DEST.partial"' EXIT

"$BIN/pg_basebackup" "${ARGS[@]}"
"$BIN/pg_verifybackup" -q "$DEST.partial"
mv "$DEST.partial" "$DEST"

# Delete everything older than the oldest full backup we keep.
OLDEST_KEPT=$(backups | grep -- '-full$' | tail -n "$KEEP_CHAINS" | head -n 1)
backups | while read -r dir; do
  if [[ "$dir" < "$OLDEST_KEPT" ]]; then
    rm -rf "$dir"
  fi
done
/etc/cron.d/pg-incremental-backup
15 2 * * * postgres /usr/local/bin/pg-incremental-backup >> /var/backups/pg/backup.log 2>&1

Run with stand-in commands over 22 simulated days, it chained each incremental to the day before, started a chain every Sunday and left two chains: between 8 and 14 daily restore points. A failed run left no directory behind. If an incremental fails because summaries are missing, run it once with FULL=1 to start a new chain. Copy the directory off the server with rclone after each run, and delete whole chains there too.

Restore: combine the chain

Copy the full backup and every incremental after it, up to the one you want, to the machine you restore on, each in its own directory. This prints the newest chain, oldest first:

Terminal
CHAIN=$(find /var/backups/pg -mindepth 1 -maxdepth 1 -type d \( -name '*-full' -o -name '*-incr' \) | sort | awk '/-full$/ { chain = "" } { chain = chain " " $0 } END { print chain }'); echo $CHAIN

With the server stopped, move the old data directory aside and rebuild into an empty one:

Terminal
sudo systemctl stop postgresql@18-main
Terminal
sudo mv /var/lib/postgresql/18/main /var/lib/postgresql/18/main.old
Terminal
sudo -u postgres /usr/lib/postgresql/18/bin/pg_combinebackup -o /var/lib/postgresql/18/main $CHAIN
  • -o is the output directory. It must not exist or must be empty.
  • The backups follow, oldest first. pg_combinebackup checks that each starts where the next expects and stops if one is missing or out of order. It does not check that each backup is intact; pg_verifybackup does.
  • Use the pg_combinebackup of the cluster's major version: it rejects another version's control file.
  • -n (--dry-run) with -d (--debug) shows what it would do without writing anything.
Copy modeWhat it doesCaveat
--copyCopies every file. The default.Needs the full space and time.
--copy-file-rangeUses the copy_file_range system call (Linux, FreeBSD).Shares disk blocks on some file systems, copies on others.
--cloneReflinks: near-instant copies.Linux 4.5+ on Btrfs or reflink-enabled XFS, or macOS APFS. On other Linux file systems it stops with error while cloning file.
--link (-k)Hard links instead of copies. PostgreSQL 18+.Same file system only. Starting the server on the output can change the input backups, so use it only on copies you will delete.

Check the result before starting it, using its new manifest:

Terminal
sudo -u postgres /usr/lib/postgresql/18/bin/pg_verifybackup /var/lib/postgresql/18/main

The output is an ordinary full base backup. Start it as is and it recovers to the end of the last backup, using the WAL that backup streamed. To go further, set restore_command, a recovery target and recovery.signal as in the point-in-time recovery guide, then start the server.

Restrictions

  • Server and pg_basebackup 17 or newer. The base backup must come from the same cluster on 17+, so start a new chain after a major upgrade.
  • Physical backups restore the whole cluster, to the same major version only. Keep pg_dump dumps for single databases and upgrades.
  • On a standby, an incremental can fail if little has happened since the previous backup, because no new restartpoint exists.
  • If you turn data checksums on or off with pg_checksums, take a new full backup: pg_combinebackup does not recompute page checksums.
  • Plain backups write tablespaces to their original paths unless you pass --tablespace-mapping, so on the same host they fail when tablespaces exist.
  • Never delete a backup a later incremental needs; delete whole chains, as the script does.

Incremental backups versus WAL archiving

Incremental chain aloneBase backup + WAL archiveBoth
Restore pointsThe end of each backupAny moment since the base backupAny moment since the full backup
WAL replayed on restoreOnly what the last backup streamedEverything since the base backupOnly WAL since the newest incremental
Data lost with the serverUp to a day, with daily backupsAbout archive_timeout plus copy delaySame as the archive

Incrementals do not replace a WAL archive; they shorten the replay. With a weekly base backup and, say, 20 GB of WAL a day, a Saturday restore replays six days, 120 GB, of WAL. With daily incrementals in between, it replays at most one day's after pg_combinebackup. The archive setup is in the point-in-time recovery guide.

Test a restore

pg_verifybackup checks files against the manifest and that the WAL each backup needs is present and parses, but its docs say it cannot catch every problem. Rebuild the newest chain on a spare machine with the same major version, start it, query the tables that matter, and time the whole run: combining a large chain takes real time. Make it a routine with test restores and set recovery targets with RPO and RTO.

Common errors

ErrorCause and fix
incremental backups cannot be taken unless WAL summarization is enabledSet summarize_wal = on and reload.
WAL summaries are required on timeline 1 from … to …, but no summaries for that timeline and LSN range existSummaries for that stretch were removed or never written. Take a new full backup; raise wal_summary_keep_time if backups can be further apart.
server does not support incremental backupThe server is older than PostgreSQL 17.
backup manifest version 1 does not support incremental backupThe base backup came from PostgreSQL 16 or older. Take a new full backup.
system identifier in backup manifest is …, but database system identifier is …The manifest belongs to another cluster.
backup at "…" starts at LSN …, but expected …A backup in the chain is missing or out of order.
backup at "…" is an incremental backup, but the first backup should be a full backupList the full backup first.
directory "…" exists but is not emptyPoint -o at a new or empty directory.
unexpected control file versionpg_combinebackup from a different major version than the backups.

Frequently asked questions

Which PostgreSQL version supports incremental backups?
PostgreSQL 17 and newer, with pg_basebackup --incremental and pg_combinebackup. PostgreSQL 18 added pg_combinebackup --link and tar-format checks in pg_verifybackup.
Can I start PostgreSQL directly on an incremental backup?
No. Combine it with the full backup and every incremental before it using pg_combinebackup, then start the server on the result.
Do incremental backups replace WAL archiving?
No. They give restore points at each backup; WAL archiving gives any moment in between. Used together, a restore replays less WAL.
Can I back up one database incrementally?
No. pg_basebackup always copies the whole cluster. For a single database, use pg_dump.

How this was checked

Commands, limits and prices were checked against these official pages, on October 4, 2026: