VPS Snaps

RPO and RTO explained for small teams running servers

RPO (recovery point objective) is how much recent data you can afford to lose, measured in time: back up once a night and you can lose up to 24 hours of changes. RTO (recovery time objective) is how long a system can be down while you bring it back. Backup frequency sets the RPO you really get. Noticing the outage, getting a server, downloading the backup, restoring it and switching DNS add up to the RTO you really get, and the only way to know that number is to time a restore.

10 min readUpdated Checked against official documentation

The definitions

NIST's glossary takes both terms from its contingency planning guide, SP 800-34 Rev. 1. In plain words:

  • Recovery point objective (RPO): the point in time your data must be recoverable to after an outage. It is stated as a duration: an RPO of 4 hours means losing at most the last 4 hours of writes.
  • Recovery time objective (RTO): how long a system can stay in recovery before the outage starts to hurt the work it supports. An RTO of 2 hours means the service must be usable again within 2 hours.
  • Maximum tolerable downtime (MTD): how long the business process itself can be disrupted before the harm is significant. SP 800-34 says the RTO must normally be shorter than the MTD, since reprocessing data can follow the restore, and that RPO is not part of the MTD.

RPO looks backward from the failure; RTO looks forward. They are separate decisions: a brochure site can accept a week-old copy but should be back in an hour, while an accounting system might survive a day offline but not one lost invoice. An objective is a target; a restore gives you a measurement. The gap between them is what your plan has to close.

Worked example: what a nightly backup really gives you

A PostgreSQL database is dumped every night at 02:00 and uploaded to object storage at another provider. The server's disk fails at 01:50 the following night.

EventTimeNotes
Last dump starts02:00, day 1The dump holds the database as it was at this moment
Upload finishes02:25, day 1From here, the off-site copy exists
Disk fails01:50, day 223 hours 50 minutes after the dump began
Data lost23 h 50 minEvery write since 02:00 on day 1

So the worst case for a nightly schedule is just under 24 hours, and a failure at a random time loses about 12 on average. Three things make the real number worse than the schedule:

  • Missed runs. If last night's backup failed unnoticed, the newest good copy is two days old. Alert on missing backups, not only failed ones.
  • The consistency point. PostgreSQL's docs say a pg_dump file is a snapshot of the database at the time pg_dump began running, so a 40-minute dump is already 40 minutes old when it finishes.
  • Upload lag. If the disaster takes out the server's provider, only the off-site copy counts. A failure at 02:10 on day 2 finds that night's dump not yet uploaded: 24 hours 10 minutes lost.

How backup frequency, snapshots and replication change RPO

Your RPO is roughly the gap between copies, but which copies count depends on what went wrong. A replica protects you from a dead disk, not from a bad migration.

MethodWorst-case RPOSurvives a dropped table?Survives losing the provider account?
Nightly dump to another providerAbout 24 hYesYes
Dump every 6 hours to another providerAbout 6 hYesYes
Daily provider snapshotAbout 24 hYesNo: same account
Hourly provider snapshotAbout 1 hYesNo: same account
WAL or binlog archiving (point-in-time recovery)MinutesYes: restore to just before itYes, if archived off-provider
Streaming replicaSecondsNo: the drop replicates tooOnly if it runs elsewhere

Snapshots have frequency limits: Google Cloud allows at most 6 snapshots of a disk every 60 minutes, for example. Continuous archiving goes further. With PostgreSQL, the RPO is the age of the newest WAL segment that reached the archive; the docs say an archive_timeout of about a minute is usually reasonable, and point to streaming replication if you need data copied off the server faster than that. See PostgreSQL point-in-time recovery and MySQL binlog recovery.

Write down an RPO per kind of disaster, not just per system. Hardware failure, a human mistake and losing the provider account usually give three different answers, and the last is often the worst, because only off-provider copies count. See snapshots versus backups.

What RTO is made of

RTO is the sum of every step between the outage and a working service. Take an app with a 40 GB PostgreSQL database whose compressed custom-format dump is 8 GB, restored onto a new server. The times are an example of what you might write down after a drill, not a benchmark:

PhaseExampleWhat sets itHow to shorten it
Notice and decide15 minMonitoring, and who may declare a disasterUptime alerts; a rule for when to stop debugging
Get a server5 minProvisioning and setupA prepared image or setup script; a warm standby
Download the backup11 minFile size and bandwidth: 8 GB at 100 Mbit/s (12.5 MB/s) is 640 seconds; at 1 Gbit/s it is about 64Restore near the storage; smaller, compressed dumps
Restore48 minDatabase size, indexes, CPU and diskParallel restore with pg_restore -j; faster disks
Check it works10 minYour checklistA scripted smoke test
Switch traffic5 minDNS TTLLower the TTL in advance
Total94 min

Downloads from object storage often run below the line rate, so measure yours. DNS is the step people forget. Cloudflare's docs describe TTL as how long a record is cached, and so how long changes take to reach users: with a 3600-second TTL, some visitors keep reaching the dead server for up to an hour after you change the record. Cloudflare's Auto TTL is 300 seconds, and proxied records are fixed at 300. If your RTO is shorter than your TTL, the TTL wins. Keep a backup of your DNS zone too.

Measure RTO for real: time a restore

Objectives come from the business; achieved values come from drills. Time each phase separately, so you know which one to fix. In bash, put time in front of a command. When it finishes, bash prints real, user and sys; real is the wall-clock time, the number you want.

Terminal
time aws s3 cp s3://acme-server-backups/web-01/appdb-2026-10-03.dump /var/backups/
Terminal
sudo -u postgres createdb -T template0 appdb_restore
Terminal
time sudo -u postgres pg_restore --exit-on-error -j 4 -d appdb_restore /var/backups/appdb-2026-10-03.dump
  • -T template0 makes the new database from the empty template, as PostgreSQL's docs recommend before a restore.
  • -j 4 runs the slowest steps, loading data and creating indexes and constraints, in 4 parallel sessions; the docs suggest starting at the number of CPU cores. It needs a custom- or directory-format archive read from a file, not standard input, so the postgres user must be able to read the dump.
  • --exit-on-error stops at the first error, so a bad dump fails fast instead of finishing with errors ignored.

Record every drill in the same file. One row per phase makes trends easy to spot. The rows below are examples matching the table above:

restore-log.csv
date,system,backup_file,phase,seconds,result,notes
2026-10-03,appdb,appdb-2026-10-03.dump,download,660,ok,8 GB from off-site bucket
2026-10-03,appdb,appdb-2026-10-03.dump,restore,2880,ok,pg_restore -j 4 on 4 vCPU
2026-10-03,appdb,appdb-2026-10-03.dump,smoke-test,600,ok,login and order page

The full drill, with row counts and a script, is in how to test a backup restore. Time it again after the database grows a lot or the hardware changes: last year's number is not this year's.

Tier your systems

Not everything needs the same targets. Three tiers cover most small teams:

TierExamplesRPO targetRTO targetTypical setupRestore test
CriticalOrders or payments database, the customer-facing app15 min or less1 to 4 hOff-provider point-in-time recovery and nightly dumps; warm standbyMonthly
ImportantInternal tools, CMS, configuration24 h1 business dayNightly dumps and file backups off-provider; daily snapshotsQuarterly
StandardStaging, build runners, sites rebuilt from git1 week, or rebuildSeveral daysWeekly snapshot, or rebuild from codeYearly

Recover in tier order, and write down dependencies: the app is not up until its database, its secrets and its DNS are.

Cost trade-offs

  • Shorter RPO costs storage and load. A 2 GB compressed dump taken nightly and kept 7 days is 14 GB. Taken hourly and kept 7 days, it is 168 copies and 336 GB: about $2.34 a month instead of $0.10 at Backblaze B2's $6.95 per TB per month (as of October 2026, ignoring the free 10 GB), plus the load of 24 dumps a day. WAL or binlog archiving ships only changes, so it can cost less than frequent full dumps.
  • Shorter RTO costs idle capacity. Restoring onto a new server is cheap but slow. A warm standby, a second server kept ready with recent data, removes most of the provisioning and restore time, and adds a second server's cost.
  • Testing costs time. A monthly drill of a critical system is an hour or two of someone's time, and the only thing that turns an RTO from a guess into a number.

A quick test: estimate what one hour of downtime and one hour of lost data cost each system. If an option costs less per month than one bad hour, it is usually worth it.

A worksheet to copy

Fill in one block per system, then keep the sheet with your disaster recovery plan.

rpo-rto.md
# RPO / RTO worksheet
Reviewed: ____________   Owner: ____________

## System: ____________________   Tier: critical / important / standard
What it does for the business: ____________________
Cost of 1 hour down: ________   Cost of 1 hour of lost data: ________
Depends on: ____________________ (database, secrets, DNS, other services)

### Targets
RPO target: ______   RTO target: ______   Maximum tolerable downtime: ______

### Copies we have
| Copy                   | How often | Where (provider, account) | Kept for |
|------------------------|-----------|---------------------------|----------|
| e.g. pg_dump           | 6 h       | other provider, own login | 30 days  |
| e.g. provider snapshot | daily     | same account as server    | 7 days   |

### Achieved RPO
Hardware failure:         ______ (newest copy on another disk)
Mistake (dropped table):  ______ (newest copy from before the mistake)
Provider account lost:    ______ (newest off-provider copy)

### Measured RTO (last drill: ____________)
| Phase              | Minutes |
|--------------------|---------|
| Notice and decide  |         |
| Get a server       |         |
| Download backup    |         |
| Restore            |         |
| Check it works     |         |
| Switch DNS (TTL)   |         |
| Total              |         |

Gap: target ______ vs measured ______
Fix: ____________________   Owner: ________   Due: ________

Frequently asked questions

What is the difference between RPO and RTO?
RPO is how much data you can lose, measured as time back from the failure. RTO is how long recovery can take, measured forward from it. Backup frequency drives RPO; restore speed, provisioning and DNS drive RTO.
What are good RPO and RTO targets for a small business?
There is no standard number. Start from what an hour of downtime and an hour of lost data cost each system. A nightly backup gives an RPO of about 24 hours, which suits many internal tools but rarely a database that takes orders or payments.
Can RPO be zero?
Close to zero for hardware failure, with a replica that confirms each write before it commits. A mistake or an attack replicates too, so for those your RPO is set by your backups or point-in-time recovery.
How do I calculate RTO?
Add up the phases: noticing the problem, getting a server, downloading the backup, restoring it, checking it and switching DNS. Then time a real restore, because the restore phase is the hardest to estimate.

How this was checked

Commands, limits and prices were checked against these official pages, on October 3, 2026: