VPS Snaps

Disaster recovery plan template for a small business server

A disaster recovery plan is a short document that says who decides a disaster has happened, which systems come back first, exactly how each one is restored, and how long that should take. For a small business with a few servers, it fits in a few pages. The template below covers scope, people, declaring a disaster, communication, priorities with RPO and RTO, recovery steps, providers, testing and sign-off. Copy it, fill in the blanks, then test it.

10 min readUpdated Checked against official documentation

What a DR plan is for

NIST describes a disaster recovery plan as a written plan for recovering information systems at an alternate facility after a major hardware or software failure, or the destruction of facilities. For a team running cloud servers, the alternate facility is usually a new server, often at another provider. The plan is what lets someone other than the person who built everything bring it back at 3 a.m.

It is not a backup policy, and not a business continuity plan. Backups are an input: the plan says which backup to restore and how. A business continuity plan covers the rest of the business, such as people, premises and suppliers. Customers and auditors often ask for a DR plan.

NIST's contingency planning guide, SP 800-34 Rev. 1, organizes a plan into five parts: supporting information, activation and notification, recovery, reconstitution (returning to normal operations) and appendices. The template follows the same shape, scaled down for a small team: sections 1 and 2 are supporting information, 3 and 4 activation and notification, 5 to 7 recovery, 8 reconstitution, and 9 to 11 the appendices.

A DR plan restores the systems; a business continuity plan keeps the business running while that happens. See business continuity vs disaster recovery.

What the plan must contain

SectionWhat goes in itWhy it matters
ScopeWhich systems and providers the plan covers, and what it does notStops the plan quietly missing DNS or the mail relay
People and contactsWho leads, who restores, who talks to customers; phone numbers; a deputy for each roleThe person who knows everything may be the one who is unreachable
Declaring a disasterWho may declare, the triggers, and how long to troubleshoot before restoringHours get lost debugging a server that should have been rebuilt
CommunicationStatus page, customer email, internal channel, and what to use if your own email is downSilence creates support tickets and costs trust
Systems and prioritiesEach system with its tier, RPO, RTO, dependencies and restore orderThe database comes back before the app
Recovery stepsExact commands and checks for each system, starting from a clean serverSomeone else has to be able to follow them
Providers and escalationAccount IDs, support links and plans, who holds each login and MFA deviceYou need them when a dashboard is down or an account is locked
TestingTabletop and live-restore schedule, with the last resultsAn untested plan is a guess
Review and sign-offVersion, owner, approver, date, change logShows the plan is current

Before you write it

  • An inventory. Every server, database, bucket, domain and outside service the business depends on, where it runs, and who owns the account.
  • Targets per system. An RPO and RTO for each, from the RPO and RTO worksheet.
  • Where the backups are. Bucket names, provider and account, where the encryption key is kept, and how old the newest copy normally is. If any of this is missing, fix the backups first; the 3-2-1 rule is the minimum.
  • Access that survives the disaster. Logins and keys stored outside the systems they recover, in a password manager the recovery lead can open from any laptop, with MFA that is not tied to one person's phone.

The template

Copy this into a document your team can reach without your servers, such as a shared drive at another company, a printed copy, or both. Replace every ____, and copy section 6 once for each system.

disaster-recovery-plan.md
# Disaster recovery plan: [company name]
Version: ____   Owner: ____   Approved by: ____   Approved on: ____
Next review: ____   Copies kept at: ____ (at least one outside our own systems)

## 1. Scope
Covers: ____ (production servers, databases, storage, DNS, email)
Does not cover: ____ (office equipment, laptops)
Relies on: our backup and retention policies, version ____

## 2. People and contacts
| Role                   | Name | Phone | Deputy | Deputy's phone |
|------------------------|------|-------|--------|----------------|
| Recovery lead          |      |       |        |                |
| Technical recovery     |      |       |        |                |
| Customer communication |      |       |        |                |
| Approves spending      |      |       |        |                |
Fallback channel if our email or chat is down: ____

## 3. Declaring a disaster
Anyone in section 2 may declare a disaster when:
- a production system is down and not fixed within ____ minutes
- data is lost, corrupted, or encrypted by an attacker
- a provider account is suspended, compromised or unreachable
- a provider reports an outage expected to last longer than our RTO
When declaring: note the time, start the incident log (section 9), call the
recovery lead, and start section 6. Keep investigating in parallel; do not
wait for a root cause before restoring.

## 4. Communication
| Audience      | Channel           | Sent by | First update within | Then every |
|---------------|-------------------|---------|---------------------|------------|
| Staff         | ____              |         | 15 min              | 30 min     |
| Customers     | Status page: ____ |         | 30 min              | 1 h        |
| Key customers | Email or phone    |         | ____                | ____       |
If our own email domain is down, send from: ____
Message outline: what happened, what still works, when the next update is.

## 5. Systems, priorities and targets
| Order | System     | Tier | RPO | RTO | Depends on | Backup used | Steps |
|-------|------------|------|-----|-----|------------|-------------|-------|
| 1     | DNS zone   |      |     |     | -          |             | 6.1   |
| 2     | Database   |      |     |     | -          |             | 6.2   |
| 3     | App server |      |     |     | Database   |             | 6.3   |
| 4     | ____       |      |     |     |            |             |       |
Restore in this order unless the recovery lead decides otherwise.

## 6. Recovery steps (copy this block for each system)
### 6.__ System: ____
Backup location: ____ (provider, account, bucket, path)
Encryption key or passphrase kept in: ____
Logins needed: ____ (kept in: ____)
1. Create a server: ____ (provider, region, size, image)
2. Install software: ____ (exact commands, or script name and location)
3. Download the newest good backup: ____ (exact command)
4. Restore: ____ (exact command)
5. Check: ____ (row counts, time of newest record, test login)
6. Connect dependents: ____ (app config, DNS record, firewall)
7. Turn on backups for the new server and confirm the first one completes.
Expected time: ____   Measured in last drill: ____
If a step fails: ____

## 7. Providers and escalation
| Provider | What we use | Account ID | Login held by | Support link and plan | MFA device |
|----------|-------------|------------|---------------|-----------------------|------------|
|          |             |            |               |                       |            |
If an account is locked: ____ (who contacts the provider, what proof they need)

## 8. Returning to normal
- Every system in section 5 passes its checks.
- Backups run on every new server, and the first ones completed.
- Credentials that may have been exposed are rotated.
- Old servers stay stopped, not deleted, until the incident review is done
  if an attack is suspected.
- Customers are told the incident is over.

## 9. Incident log
| Time (UTC) | Who | What happened or was decided |
|------------|-----|------------------------------|

## 10. Testing
| Test                             | How often | Last run | Result | Actions |
|----------------------------------|-----------|----------|--------|---------|
| Tabletop exercise                | ____      |          |        |         |
| Database restore                 | ____      |          |        |         |
| File restore                     | ____      |          |        |         |
| Full rebuild at another provider | ____      |          |        |         |

## 11. Review and sign-off
Reviewed every ____ months, and after every test, incident, or change to
systems, providers or people.
| Version | Date | Changes | Approved by |
|---------|------|---------|-------------|

The parts people get wrong

  • The trigger. Set a time limit on troubleshooting before recovery starts, such as 30 minutes for a critical system. Without one, teams keep debugging long after a rebuild would have finished. Start the restore in parallel; you can always stop it.
  • Commands, not intentions. 'Restore the database' is not a step. The download command, the restore command and the query that proves it worked are. The restore drill guide has commands you can paste in.
  • Where the plan lives. A plan stored only on the server it describes, or behind a login that is down, is no plan. Keep a copy elsewhere, and on paper if the business can't run without it.
  • Losing the account. It is easy to plan only for a dead server at a working provider. Also write the opposite case: the account is locked and you rebuild at another provider from off-provider backups. The migration guide covers moving a server between providers.
  • Returning to normal. A new server is not finished until its own backups run. That is why step 7 is in every system block.

Test it: tabletop exercises and live restores

NIST SP 800-84 defines a tabletop exercise as a discussion: the people with roles in the plan meet, a facilitator presents a scenario and asks questions, and they talk through what they would do. SP 800-34 adds functional exercises, where people carry out their roles in a simulated environment, and for more important systems it says the exercise should include recovering a system from backup.

For a small team, that becomes two kinds of test:

TabletopLive restore
What happens45 to 60 minutes on a call or around a table, walking through one scenario with the plan openA real restore of one system from its backup onto a scratch server, following the plan's steps
What it findsMissing contacts, unclear authority, steps nobody owns, logins only one person hasWrong commands, missing files, slow restores, your real RTO
What it costsAn hour of people's timeAn hour or more, plus a temporary server
How oftenTwice a year, and when people changePer system by tier: monthly for critical, quarterly for important

Scenarios worth running:

  • The provider suspends your account at 9 a.m. on a Monday. Every server and snapshot is out of reach.
  • A migration drops the orders table, and nobody notices for 3 hours.
  • Files on the app server are encrypted and a ransom note is in /root.
  • The person who set everything up is on a 12-hour flight.
  • Someone takes over the DNS account and points your domain elsewhere.

After each test, write down what failed, fix the plan or the systems, and record it in section 10. A test that finds nothing was probably too easy.

How often to review it

SP 800-34 says to review the plan at a frequency you set and whenever significant changes occur, and notes that contact lists need reviewing more often. In practice:

  • A full read-through by the owner every 6 to 12 months.
  • An update whenever a system, provider, backup tool or key person changes.
  • An update after every test and every real incident.
  • Contact details and logins checked every quarter.

Each review ends with a new version number, date and approver in section 11. That record, with the test results, is what you show when a customer or auditor asks for your plan.

Frequently asked questions

What should a disaster recovery plan include?
Scope, people and contacts, how a disaster is declared, communication, each system with its priority and RPO and RTO, step-by-step recovery for each system, provider details and escalation, a testing schedule, and a review and sign-off record.
What is the difference between a disaster recovery plan and a business continuity plan?
A DR plan restores IT systems. A business continuity plan keeps the whole business running, including people, premises and suppliers, and usually points to the DR plan for the IT part.
How often should a disaster recovery plan be tested?
A tabletop exercise twice a year is a reasonable rhythm for a small team, plus live restores of each system on a schedule that matches its importance, such as monthly for critical systems and quarterly for the rest.
Who should own the disaster recovery plan?
One named person with a named deputy, usually the most senior technical lead. Owning it means keeping it current and running the tests, not doing every restore personally.

How this was checked

Commands, limits and prices were checked against these official pages, on October 3, 2026: