How to fail over a website to a standby server with DNS
DNS failover points your site's name at a standby server when the primary stops answering. It works across any providers, but it is not instant: a visitor moves only when their cached answer expires, so set a 60-second TTL in advance and expect most traffic to arrive within a few minutes, with a few clients later. Keep the standby in sync, switch after several failed checks, switch once, and fail back by hand.
How DNS failover works
Browsers find your server by looking up its name. The answer comes from a resolver (your ISP's, a public one like 1.1.1.1, or a company's), which caches it for the record's TTL. To fail over, you change the A record from the primary's IP to the standby's. New lookups get the standby; cached answers keep sending people to the dead primary until they expire.
You need a standby with recent data, a health check on a third machine, an API call that changes the record, and a way back. Only the record moves, so the standby can sit in another region or at another provider.
What DNS failover cannot do
- Beat every cache. Resolver operators can override your TTL: with Unbound's
cache-min-ttl, "the data is cached for longer than the domain owner intended", in its manual's words. - Reach applications with their own cache. Java caches lookups for as long as
networkaddress.cache.ttlsays, not your record;-1means forever, until a restart. Cloudflare notes that local DNS caches can also delay a change. - Move open connections. Keep-alive HTTP, WebSockets and database connections stay on the old server until they close.
- Bring data the standby never received. Writes after the last sync are missing.
So plan for a tail: most visitors follow within TTL plus detection time, a few much later.
Choose a TTL
Cloudflare's docs put the trade-off simply: longer TTLs speed up lookups because more answers come from cache, but changes take longer to take effect. DNS-only records accept 60 seconds to 1 day (30 seconds on Enterprise). Proxied records always use Auto, 300 seconds, and cannot be edited.
Worked example, using the script below: cron starts it within 60 seconds of the outage. Three failed checks, 20 seconds apart with a 10-second timeout each, plus the Cloudflare and standby checks, take up to about 90 seconds more. So the record changes at most about 2.5 minutes after the primary dies. Then cached answers have to expire:
| Record | Record changed after | Visitors whose resolver follows the TTL arrive by |
|---|---|---|
| DNS-only, TTL 60 | about 2.5 minutes | about 3.5 minutes |
| DNS-only, TTL 300 | about 2.5 minutes | about 7.5 minutes |
| DNS-only, TTL 3600 | about 2.5 minutes | about 62.5 minutes |
| Proxied (orange cloud) | about 2.5 minutes | about 2.5 minutes: resolvers cache Cloudflare's addresses, which do not change |
Set the short TTL days before you need it. A resolver that cached the record with a TTL of 3600 keeps it for up to an hour, whatever you change during the outage.
Keep the standby in sync
| Data | How to copy it | What a failover loses |
|---|---|---|
| PostgreSQL | Streaming replication | Commits not yet sent; PostgreSQL's docs say the delay is typically under a second |
| MySQL or MariaDB | Built-in replication | Changes the replica had not received |
| Any database | Restore the newest dump on the standby on a schedule | Up to one interval plus the dump's age |
| Uploaded files | rsync every few minutes, or object storage both servers read | Changes since the last run |
| Code, config, certificates | Deploy to both servers every time | Nothing, if no deploy skips the standby |
For PostgreSQL, create a role with REPLICATION and LOGIN on the primary, allow the standby's IP for the replication database in pg_hba.conf, and put the password in the standby's ~/.pgpass. Then, with PostgreSQL stopped on the standby and its data directory empty, clone the primary:
sudo -u postgres pg_basebackup -h 203.0.113.10 -U replicator -D /var/lib/postgresql/17/main -R -P-R writes standby.signal and the connection settings, so the copy starts as a standby; -P shows progress. Use a private network or TLS: the stream carries all your data. On failover, promote the standby's database before traffic arrives:
sudo -u postgres psql -c "SELECT pg_promote();"A standby rebuilt from backups is simpler and loses more; RPO and RTO helps you decide what loss is acceptable, and restoring a server from backup covers the restore.
Health checks, flapping and split brain
The switch should happen once, for a real outage, and only to a standby that works:
- Check from a third machine. Neither server can judge itself.
- Check a page that proves the app works: a
/healthzroute that runs a database query. - Require consecutive failures. One failed request is often a blip; the script wants three in a row, 20 seconds apart.
- Check the checker. If it cannot reach the internet, the fault may be its own.
- Check the standby before switching.
- Never switch back automatically. A flaky primary would bounce traffic, and writes, between two servers. That is flapping.
Split brain is the worst case: both servers accept writes. PostgreSQL's failover docs call for a way to tell the old primary it is no longer primary, known as STONITH, "to avoid situations where both systems think they are the primary, which will lead to confusion and ultimately data loss." With DNS failover the old primary may still serve visitors whose resolver holds its IP. So once you have switched, power it off through your provider's API:
doctl compute droplet-action power-off <droplet-id>hcloud server poweroff <server>DigitalOcean's power_off is a hard shutdown, and a powered-off Droplet is still billed.
Switch a Cloudflare record with a script
In the Cloudflare dashboard, go to My Profile > API Tokens > Create Token and use the Edit zone DNS template. Limit Zone Resources to this one zone and, under Client IP Address Filtering, to the checker's IP. The token gets DNS Write, which the update endpoint PATCH /zones/{zone_id}/dns_records/{dns_record_id} requires. Store it on the checker as a header file only root can read, as in backing up Cloudflare DNS:
sudo install -d -m 700 /etc/cloudflaresudo install -m 600 /dev/null /etc/cloudflare/dns-edit.headerread -rsp 'Cloudflare token: ' T && printf 'Authorization: Bearer %s\n' "$T" | sudo tee /etc/cloudflare/dns-edit.header > /dev/null; unset TFind the record's ID. The zone ID is on the domain's Overview page. This prints ID, name, content, proxy status and TTL for each A record:
sudo curl -sS --fail-with-body -H @/etc/cloudflare/dns-edit.header "https://api.cloudflare.com/client/v4/zones/$ZONE_ID/dns_records?type=A&per_page=5000" | jq -r '.result[] | [.id, .name, .content, .proxied, .ttl] | @tsv'Save the script, fill in the variables, and serve /healthz on both servers. It needs curl 7.76 or newer and jq.
#!/usr/bin/env bash
# Point one Cloudflare A record at a standby server when the primary is down.
# Runs every minute from a third machine. It switches once, then waits for a person.
set -euo pipefail
NAME="www.example.com"
PRIMARY_IP="203.0.113.10"
STANDBY_IP="198.51.100.20"
ZONE_ID="your-zone-id"
RECORD_ID="your-record-id"
PROXIED=false # true for an orange-cloud record
TTL=60 # 1 (Auto) for an orange-cloud record
AUTH="/etc/cloudflare/dns-edit.header"
STATE="/var/lib/dns-failover/failed-over"
CF="https://api.cloudflare.com/client/v4"
log() { echo "$(date -Is) $*"; }
healthy() {
# Ask one server directly, with the real host name, so TLS and virtual hosts match.
curl -fsS -o /dev/null --max-time 10 --resolve "$NAME:443:$1" "https://$NAME/healthz"
}
# Already switched: do nothing until a person deletes the state file.
[ -e "$STATE" ] && exit 0
# The primary answers: nothing to do.
healthy "$PRIMARY_IP" && exit 0
# If this machine cannot reach Cloudflare, the fault may be here, not there.
if ! curl -fsS -o /dev/null --max-time 10 -H @"$AUTH" "$CF/user/tokens/verify"; then
log "cannot reach the Cloudflare API; not switching"
exit 1
fi
# Two more checks, 20 seconds apart. One success cancels the failover.
for _ in 1 2; do
sleep 20
healthy "$PRIMARY_IP" && exit 0
done
# Never switch to a standby that is failing too.
if ! healthy "$STANDBY_IP"; then
log "primary and standby both failing; not switching"
exit 1
fi
BODY=$(jq -nc --arg name "$NAME" --arg ip "$STANDBY_IP" \
--argjson proxied "$PROXIED" --argjson ttl "$TTL" \
'{type: "A", name: $name, content: $ip, ttl: $ttl, proxied: $proxied}')
if ! RESP=$(curl -sS --fail-with-body -X PATCH -H @"$AUTH" \
-H "Content-Type: application/json" --data "$BODY" \
"$CF/zones/$ZONE_ID/dns_records/$RECORD_ID") \
|| ! jq -e '.success' <<< "$RESP" > /dev/null; then
log "Cloudflare did not apply the change: $RESP"
exit 1
fi
mkdir -p "$(dirname "$STATE")"
date -Is > "$STATE"
log "$NAME now points to $STANDBY_IP"--resolvesends the request to one server's IP under the real host name, so TLS and virtual hosts behave as they do for visitors.-ffails on HTTP errors;--max-time 10fails a server that hangs.- The
/user/tokens/verifycall proves the checker can reach Cloudflare, and the token works, before anything changes. - An HTTP error or
"success": falsestops the script, and the log shows Cloudflare's reply. The state file makes the switch one-way.
sudo chmod 700 /usr/local/bin/dns-failover.sh* * * * * root flock -n /run/dns-failover.lock /usr/local/bin/dns-failover.sh >> /var/log/dns-failover.log 2>&1A failing run lasts over a minute; flock -n makes the next one exit instead of overlapping. We ran the script against local stand-ins for the API and both servers: with the primary down it checked three times and sent one PATCH, and the next run exited at once; with a failing standby, an unreachable API, a rejected token or a refused update, it exited with code 1 and changed nothing.
Fail over an AAAA record on the same name too, run one copy per record (the apex and www are separate), let the checker through the origin's firewall, and add an alert after the last log line so a person knows.
Proxied or DNS-only: what changes
| Proxied (orange cloud) | DNS-only (grey cloud) | |
|---|---|---|
| Lookups return | Cloudflare's anycast addresses | Your server's IP |
| TTL | Auto (300 seconds), not editable | 60 seconds to 1 day (30 seconds on Enterprise) |
| A switch reaches visitors | When Cloudflare applies the edit; resolver caches hold Cloudflare's addresses, which stay the same | As each cached answer expires |
| Traffic covered | HTTP and HTTPS through Cloudflare | Any port and protocol |
| Your servers' IPs | Hidden | Public |
Cloudflare's Load Balancing docs say proxying "offers faster failover and more accurate routing, which can otherwise be affected by DNS caching." For a proxied record, set PROXIED=true and TTL=1. The standby must already hold a certificate that satisfies your SSL/TLS mode; copy /etc/letsencrypt as in moving a server to a new provider. Keep names for SSH, mail or databases DNS-only. Cloudflare's Load Balancing product can also run the health checks and switch traffic for you.
Fail back
- Leave the standby serving while you find out why the primary failed.
- Copy everything written since the switch back to the primary. For PostgreSQL, rebuild it as a standby of the new primary with
pg_basebackup(orpg_rewindon a large cluster); rsync uploads back. - In a quiet window, turn on maintenance mode, let replication catch up, and promote the primary's database.
- Send the same
PATCHwith the primary's IP ascontent. - Make the other server a standby again and re-arm the script:
sudo rm /var/lib/dns-failover/failed-overPostgreSQL's docs note that switching regularly is useful: it "also serves as a test of the failover mechanism". Run a planned switchover a few times a year, time it, and record the result in your recovery plan.
The alternative: a movable IP
A movable IP stays the same in DNS; you move it between servers through the provider's API, so no resolver cache is involved and every port moves with it. But it moves only within one provider and one location, so it cannot save you from a region or provider going down. And as DigitalOcean's docs put it, "a reserved IP alone does not automatically provide high availability": you still need the health check. Prices as of October 2026:
| Provider | Name | Moves within | Cost | Notes |
|---|---|---|---|---|
| DigitalOcean | Reserved IP | One datacenter | Free while assigned; unassigned IPv4 $5.00/month | Remaps instantly. Outbound traffic uses the anchor IP by default. |
| Hetzner Cloud | Floating IP | One network zone (eu-central: Falkenstein, Helsinki, Nuremberg) | IPv4 €3.00/month, IPv6 €1.00/month, excl. VAT | Must be configured inside each server's OS. |
| AWS EC2 | Elastic IP | One Region | Charged like every public IPv4 address, in use or idle | Associating it with a new instance detaches it from the old one. |
doctl compute reserved-ip-action assign 203.0.113.25 386734086hcloud floating-ip assign <floating-ip> <server>aws ec2 associate-address --allocation-id <eipalloc-id> --instance-id <instance-id>On Hetzner, add the floating IP to both servers' network configuration in advance (Hetzner's guide uses /etc/netplan/60-floating-ip.yaml on Ubuntu), so the standby answers the moment it arrives; if assign is refused because the IP is still assigned, run hcloud floating-ip unassign <floating-ip> first. On AWS, --instance-id needs an instance with exactly one network interface. The rest is the same as DNS failover: check from elsewhere, require several failures, fence the old server, switch once.
Frequently asked questions
- How fast is DNS failover?
- Detection time plus the record's TTL for most visitors. With checks every minute and a 60-second TTL, about 3.5 minutes at worst. Some resolvers and applications hold old answers longer. With a proxied Cloudflare record, resolver caches do not delay the switch.
- What TTL should I use for DNS failover?
- 60 seconds, the lowest Cloudflare allows below Enterprise, set days before any outage. Answers cached under an older, longer TTL keep it until they expire.
- Should failover switch back automatically?
- No. A flaky primary then bounces traffic and writes between two servers. Fail back by hand after copying the standby's new data to the primary.
- Is a floating IP better than DNS failover?
- It is faster and covers every port, but it only moves within one datacenter, zone or Region at one provider. DNS failover is slower but works across regions and providers.
How this was checked
Commands, limits and prices were checked against these official pages, on October 4, 2026:
- Cloudflare API: Update DNS Record
- Cloudflare API: List DNS Records
- Cloudflare docs: Create API token
- Cloudflare docs: API token permissions
- Cloudflare DNS docs: Time to Live (TTL)
- Cloudflare DNS docs: Proxy status
- Cloudflare Load Balancing docs: Proxy modes
- Java SE 21 API: InetAddress (InetAddress caching)
- Unbound documentation: unbound.conf (cache-min-ttl)
- PostgreSQL documentation: Failover
- PostgreSQL documentation: Log-Shipping Standby Servers
- PostgreSQL documentation: pg_basebackup
- PostgreSQL documentation: System Administration Functions (pg_promote)
- DigitalOcean: Reserved IP features
- DigitalOcean: Reserved IP limits
- DigitalOcean: Reserved IP pricing
- DigitalOcean: doctl compute reserved-ip-action assign
- DigitalOcean: doctl compute droplet-action power-off
- Hetzner Docs: Floating IPs overview
- Hetzner Docs: Floating IP FAQ
- Hetzner Docs: Floating IP persistent configuration
- Hetzner Docs: Adding a Floating IP
- hcloud CLI reference: floating-ip assign
- hcloud CLI reference: server poweroff
- Amazon EC2 User Guide: Elastic IP addresses
- AWS CLI reference: ec2 associate-address
- curl(1) manual (Ubuntu 24.04)
- flock(1) manual