VPS Snaps

How to fail over a website to a standby server with DNS

DNS failover points your site's name at a standby server when the primary stops answering. It works across any providers, but it is not instant: a visitor moves only when their cached answer expires, so set a 60-second TTL in advance and expect most traffic to arrive within a few minutes, with a few clients later. Keep the standby in sync, switch after several failed checks, switch once, and fail back by hand.

11 min readUpdated Checked against official documentation

How DNS failover works

Browsers find your server by looking up its name. The answer comes from a resolver (your ISP's, a public one like 1.1.1.1, or a company's), which caches it for the record's TTL. To fail over, you change the A record from the primary's IP to the standby's. New lookups get the standby; cached answers keep sending people to the dead primary until they expire.

You need a standby with recent data, a health check on a third machine, an API call that changes the record, and a way back. Only the record moves, so the standby can sit in another region or at another provider.

What DNS failover cannot do

  • Beat every cache. Resolver operators can override your TTL: with Unbound's cache-min-ttl, "the data is cached for longer than the domain owner intended", in its manual's words.
  • Reach applications with their own cache. Java caches lookups for as long as networkaddress.cache.ttl says, not your record; -1 means forever, until a restart. Cloudflare notes that local DNS caches can also delay a change.
  • Move open connections. Keep-alive HTTP, WebSockets and database connections stay on the old server until they close.
  • Bring data the standby never received. Writes after the last sync are missing.

So plan for a tail: most visitors follow within TTL plus detection time, a few much later.

Choose a TTL

Cloudflare's docs put the trade-off simply: longer TTLs speed up lookups because more answers come from cache, but changes take longer to take effect. DNS-only records accept 60 seconds to 1 day (30 seconds on Enterprise). Proxied records always use Auto, 300 seconds, and cannot be edited.

Worked example, using the script below: cron starts it within 60 seconds of the outage. Three failed checks, 20 seconds apart with a 10-second timeout each, plus the Cloudflare and standby checks, take up to about 90 seconds more. So the record changes at most about 2.5 minutes after the primary dies. Then cached answers have to expire:

RecordRecord changed afterVisitors whose resolver follows the TTL arrive by
DNS-only, TTL 60about 2.5 minutesabout 3.5 minutes
DNS-only, TTL 300about 2.5 minutesabout 7.5 minutes
DNS-only, TTL 3600about 2.5 minutesabout 62.5 minutes
Proxied (orange cloud)about 2.5 minutesabout 2.5 minutes: resolvers cache Cloudflare's addresses, which do not change

Set the short TTL days before you need it. A resolver that cached the record with a TTL of 3600 keeps it for up to an hour, whatever you change during the outage.

Keep the standby in sync

DataHow to copy itWhat a failover loses
PostgreSQLStreaming replicationCommits not yet sent; PostgreSQL's docs say the delay is typically under a second
MySQL or MariaDBBuilt-in replicationChanges the replica had not received
Any databaseRestore the newest dump on the standby on a scheduleUp to one interval plus the dump's age
Uploaded filesrsync every few minutes, or object storage both servers readChanges since the last run
Code, config, certificatesDeploy to both servers every timeNothing, if no deploy skips the standby

For PostgreSQL, create a role with REPLICATION and LOGIN on the primary, allow the standby's IP for the replication database in pg_hba.conf, and put the password in the standby's ~/.pgpass. Then, with PostgreSQL stopped on the standby and its data directory empty, clone the primary:

Terminal
sudo -u postgres pg_basebackup -h 203.0.113.10 -U replicator -D /var/lib/postgresql/17/main -R -P

-R writes standby.signal and the connection settings, so the copy starts as a standby; -P shows progress. Use a private network or TLS: the stream carries all your data. On failover, promote the standby's database before traffic arrives:

Terminal
sudo -u postgres psql -c "SELECT pg_promote();"

A standby rebuilt from backups is simpler and loses more; RPO and RTO helps you decide what loss is acceptable, and restoring a server from backup covers the restore.

Health checks, flapping and split brain

The switch should happen once, for a real outage, and only to a standby that works:

  • Check from a third machine. Neither server can judge itself.
  • Check a page that proves the app works: a /healthz route that runs a database query.
  • Require consecutive failures. One failed request is often a blip; the script wants three in a row, 20 seconds apart.
  • Check the checker. If it cannot reach the internet, the fault may be its own.
  • Check the standby before switching.
  • Never switch back automatically. A flaky primary would bounce traffic, and writes, between two servers. That is flapping.

Split brain is the worst case: both servers accept writes. PostgreSQL's failover docs call for a way to tell the old primary it is no longer primary, known as STONITH, "to avoid situations where both systems think they are the primary, which will lead to confusion and ultimately data loss." With DNS failover the old primary may still serve visitors whose resolver holds its IP. So once you have switched, power it off through your provider's API:

Terminal
doctl compute droplet-action power-off <droplet-id>
Terminal
hcloud server poweroff <server>

DigitalOcean's power_off is a hard shutdown, and a powered-off Droplet is still billed.

Switch a Cloudflare record with a script

In the Cloudflare dashboard, go to My Profile > API Tokens > Create Token and use the Edit zone DNS template. Limit Zone Resources to this one zone and, under Client IP Address Filtering, to the checker's IP. The token gets DNS Write, which the update endpoint PATCH /zones/{zone_id}/dns_records/{dns_record_id} requires. Store it on the checker as a header file only root can read, as in backing up Cloudflare DNS:

Terminal
sudo install -d -m 700 /etc/cloudflare
Terminal
sudo install -m 600 /dev/null /etc/cloudflare/dns-edit.header
Terminal
read -rsp 'Cloudflare token: ' T && printf 'Authorization: Bearer %s\n' "$T" | sudo tee /etc/cloudflare/dns-edit.header > /dev/null; unset T

Find the record's ID. The zone ID is on the domain's Overview page. This prints ID, name, content, proxy status and TTL for each A record:

Terminal
sudo curl -sS --fail-with-body -H @/etc/cloudflare/dns-edit.header "https://api.cloudflare.com/client/v4/zones/$ZONE_ID/dns_records?type=A&per_page=5000" | jq -r '.result[] | [.id, .name, .content, .proxied, .ttl] | @tsv'

Save the script, fill in the variables, and serve /healthz on both servers. It needs curl 7.76 or newer and jq.

/usr/local/bin/dns-failover.sh
#!/usr/bin/env bash
# Point one Cloudflare A record at a standby server when the primary is down.
# Runs every minute from a third machine. It switches once, then waits for a person.
set -euo pipefail

NAME="www.example.com"
PRIMARY_IP="203.0.113.10"
STANDBY_IP="198.51.100.20"
ZONE_ID="your-zone-id"
RECORD_ID="your-record-id"
PROXIED=false   # true for an orange-cloud record
TTL=60          # 1 (Auto) for an orange-cloud record
AUTH="/etc/cloudflare/dns-edit.header"
STATE="/var/lib/dns-failover/failed-over"
CF="https://api.cloudflare.com/client/v4"

log() { echo "$(date -Is) $*"; }

healthy() {
  # Ask one server directly, with the real host name, so TLS and virtual hosts match.
  curl -fsS -o /dev/null --max-time 10 --resolve "$NAME:443:$1" "https://$NAME/healthz"
}

# Already switched: do nothing until a person deletes the state file.
[ -e "$STATE" ] && exit 0

# The primary answers: nothing to do.
healthy "$PRIMARY_IP" && exit 0

# If this machine cannot reach Cloudflare, the fault may be here, not there.
if ! curl -fsS -o /dev/null --max-time 10 -H @"$AUTH" "$CF/user/tokens/verify"; then
  log "cannot reach the Cloudflare API; not switching"
  exit 1
fi

# Two more checks, 20 seconds apart. One success cancels the failover.
for _ in 1 2; do
  sleep 20
  healthy "$PRIMARY_IP" && exit 0
done

# Never switch to a standby that is failing too.
if ! healthy "$STANDBY_IP"; then
  log "primary and standby both failing; not switching"
  exit 1
fi

BODY=$(jq -nc --arg name "$NAME" --arg ip "$STANDBY_IP" \
  --argjson proxied "$PROXIED" --argjson ttl "$TTL" \
  '{type: "A", name: $name, content: $ip, ttl: $ttl, proxied: $proxied}')

if ! RESP=$(curl -sS --fail-with-body -X PATCH -H @"$AUTH" \
    -H "Content-Type: application/json" --data "$BODY" \
    "$CF/zones/$ZONE_ID/dns_records/$RECORD_ID") \
  || ! jq -e '.success' <<< "$RESP" > /dev/null; then
  log "Cloudflare did not apply the change: $RESP"
  exit 1
fi

mkdir -p "$(dirname "$STATE")"
date -Is > "$STATE"
log "$NAME now points to $STANDBY_IP"
  • --resolve sends the request to one server's IP under the real host name, so TLS and virtual hosts behave as they do for visitors. -f fails on HTTP errors; --max-time 10 fails a server that hangs.
  • The /user/tokens/verify call proves the checker can reach Cloudflare, and the token works, before anything changes.
  • An HTTP error or "success": false stops the script, and the log shows Cloudflare's reply. The state file makes the switch one-way.
Terminal
sudo chmod 700 /usr/local/bin/dns-failover.sh
/etc/cron.d/dns-failover
* * * * * root flock -n /run/dns-failover.lock /usr/local/bin/dns-failover.sh >> /var/log/dns-failover.log 2>&1

A failing run lasts over a minute; flock -n makes the next one exit instead of overlapping. We ran the script against local stand-ins for the API and both servers: with the primary down it checked three times and sent one PATCH, and the next run exited at once; with a failing standby, an unreachable API, a rejected token or a refused update, it exited with code 1 and changed nothing.

Fail over an AAAA record on the same name too, run one copy per record (the apex and www are separate), let the checker through the origin's firewall, and add an alert after the last log line so a person knows.

Proxied or DNS-only: what changes

Proxied (orange cloud)DNS-only (grey cloud)
Lookups returnCloudflare's anycast addressesYour server's IP
TTLAuto (300 seconds), not editable60 seconds to 1 day (30 seconds on Enterprise)
A switch reaches visitorsWhen Cloudflare applies the edit; resolver caches hold Cloudflare's addresses, which stay the sameAs each cached answer expires
Traffic coveredHTTP and HTTPS through CloudflareAny port and protocol
Your servers' IPsHiddenPublic

Cloudflare's Load Balancing docs say proxying "offers faster failover and more accurate routing, which can otherwise be affected by DNS caching." For a proxied record, set PROXIED=true and TTL=1. The standby must already hold a certificate that satisfies your SSL/TLS mode; copy /etc/letsencrypt as in moving a server to a new provider. Keep names for SSH, mail or databases DNS-only. Cloudflare's Load Balancing product can also run the health checks and switch traffic for you.

Fail back

  1. Leave the standby serving while you find out why the primary failed.
  2. Copy everything written since the switch back to the primary. For PostgreSQL, rebuild it as a standby of the new primary with pg_basebackup (or pg_rewind on a large cluster); rsync uploads back.
  3. In a quiet window, turn on maintenance mode, let replication catch up, and promote the primary's database.
  4. Send the same PATCH with the primary's IP as content.
  5. Make the other server a standby again and re-arm the script:
Terminal
sudo rm /var/lib/dns-failover/failed-over

PostgreSQL's docs note that switching regularly is useful: it "also serves as a test of the failover mechanism". Run a planned switchover a few times a year, time it, and record the result in your recovery plan.

The alternative: a movable IP

A movable IP stays the same in DNS; you move it between servers through the provider's API, so no resolver cache is involved and every port moves with it. But it moves only within one provider and one location, so it cannot save you from a region or provider going down. And as DigitalOcean's docs put it, "a reserved IP alone does not automatically provide high availability": you still need the health check. Prices as of October 2026:

ProviderNameMoves withinCostNotes
DigitalOceanReserved IPOne datacenterFree while assigned; unassigned IPv4 $5.00/monthRemaps instantly. Outbound traffic uses the anchor IP by default.
Hetzner CloudFloating IPOne network zone (eu-central: Falkenstein, Helsinki, Nuremberg)IPv4 €3.00/month, IPv6 €1.00/month, excl. VATMust be configured inside each server's OS.
AWS EC2Elastic IPOne RegionCharged like every public IPv4 address, in use or idleAssociating it with a new instance detaches it from the old one.
Terminal
doctl compute reserved-ip-action assign 203.0.113.25 386734086
Terminal
hcloud floating-ip assign <floating-ip> <server>
Terminal
aws ec2 associate-address --allocation-id <eipalloc-id> --instance-id <instance-id>

On Hetzner, add the floating IP to both servers' network configuration in advance (Hetzner's guide uses /etc/netplan/60-floating-ip.yaml on Ubuntu), so the standby answers the moment it arrives; if assign is refused because the IP is still assigned, run hcloud floating-ip unassign <floating-ip> first. On AWS, --instance-id needs an instance with exactly one network interface. The rest is the same as DNS failover: check from elsewhere, require several failures, fence the old server, switch once.

Frequently asked questions

How fast is DNS failover?
Detection time plus the record's TTL for most visitors. With checks every minute and a 60-second TTL, about 3.5 minutes at worst. Some resolvers and applications hold old answers longer. With a proxied Cloudflare record, resolver caches do not delay the switch.
What TTL should I use for DNS failover?
60 seconds, the lowest Cloudflare allows below Enterprise, set days before any outage. Answers cached under an older, longer TTL keep it until they expire.
Should failover switch back automatically?
No. A flaky primary then bounces traffic and writes between two servers. Fail back by hand after copying the standby's new data to the primary.
Is a floating IP better than DNS failover?
It is faster and covers every port, but it only moves within one datacenter, zone or Region at one provider. DNS failover is slower but works across regions and providers.

How this was checked

Commands, limits and prices were checked against these official pages, on October 4, 2026: