Disaster recovery
Plan for the bad day: business continuity vs disaster recovery, RPO and RTO, retention, ransomware, DNS failover, rebuilding a server, a DR plan template, the 3-2-1 rule, testing restores, and recovering a hacked server or deleted files.
The 3-2-1 backup rule for servers: what it means for a VPS
What three copies, two kinds of storage and one off-site copy mean for a cloud server, why the off-site copy belongs at a different provider, and three setups that meet the rule at different budgets.
9 min readHow to test a backup restore: files, PostgreSQL and a full server
Restore drills you can run today: verify a tar backup with sha256sum, restore a PostgreSQL dump into a scratch database and count its rows, boot a server from a snapshot, and record what each drill proves.
10 min read · TestedSnapshot vs backup: what is the difference for a cloud server?
What a cloud snapshot captures, how it differs from a file or database backup, crash-consistent vs application-consistent, the same-account risk, and the combination that covers a typical VPS.
8 min read · TestedRPO and RTO explained for small teams running servers
What recovery point and recovery time objectives mean for a server, how backup frequency and restore speed set them, how to measure both with a timed restore, and a worksheet for setting targets per system.
10 min readHow long to keep backups: writing a backup retention policy
Choose how long to keep server backups with grandfather-father-son rotation, work out the storage cost, avoid minimum-duration charges, plan deletion, and copy a sample retention policy.
9 min readDisaster recovery plan template for a small business server
What a disaster recovery plan for a few servers must contain, a complete fill-in-the-blanks template, how to test it with a tabletop exercise and a live restore, and how often to review it.
10 min readHow to protect backups from ransomware and account compromise
Stop an attacker with root on your server, or with your cloud login, from deleting your backups: write-only keys, Object Lock, append-only repositories, a copy at another provider, MFA, and alerts on deletions.
9 min readHow to back up and restore Cloudflare DNS records
Export Cloudflare DNS records as a zone file and as JSON, automate it with a read-only API token and cron, and restore by putting back only what changed.
10 min readHow to restore a server from backup when the old one is gone
Rebuild a lost server from its backups in the right order: matching OS and versions, users, /etc file by file, databases from dumps, files, services, TLS and DNS, with checks before traffic moves and a runbook to copy.
11 min readHow to fail over a website to a standby server with DNS
Move a website to a standby server when the primary fails by changing its DNS record: pick a TTL, keep the standby in sync, run a health check that switches a Cloudflare record once, fail back, and know when a movable IP is the better tool.
11 min readWhat to do when your server is hacked: recovering from backups
Isolate a compromised Linux server without destroying evidence, work out when the attack began, rebuild from a backup taken before it, rotate every credential, check your backups were not touched, and meet the GDPR's 72-hour notice rule.
10 min readHow to recover deleted files on Linux
Get a deleted file back on a Linux server: copy it from /proc while a program still has it open, check whether TRIM already erased it, what debugfs, extundelete and PhotoRec can and cannot do on ext4, XFS and Btrfs, and how to restore one file from a backup.
10 min read · TestedBusiness continuity vs disaster recovery: what's the difference?
Business continuity keeps the business running through a disruption; disaster recovery brings its IT systems back. Definitions from NIST and ISO, how a business impact analysis sets RTO and RPO, a worked example, how to test each, and an outline to copy.
10 min read