VPS Snaps

How to monitor website and server uptime

To monitor uptime, request your site every one to five minutes from a machine outside your own network, check the status code and a piece of text that only a working page shows, and alert a named person once the failure repeats two or three times. Add certificate expiry, port and heartbeat checks for what a page check cannot see. This guide covers each check and what the interval costs you in detection time, and ends with a checker built from curl and cron.

11 min readUpdated Tested on Ubuntu 24.04 LTS, curl 8.5.0, OpenSSL 3.0.13, bash 5.2

What to check

CheckWhat it catchesWhat it misses
HTTP statusA server that is off, a crashed app (502, 503), DNS and TLS failuresA 200 page that shows an error
Text on the pageError pages served with a 200, a dead database behind a cached pageWrong data in the right layout
Certificate expiryA renewal failing quietly, weeks before browsers refuse the siteEverything except one date
TCP portA service that stopped listening (SSH, mail), or a database port that should be closed but is openWhether it answers correctly
DNSChanged or deleted records, silent name servers, a lapsed domainAnything after the name resolves

For the text check, pick words that appear only when the page was built from live data, such as a product name from the database. A page cache can keep serving the site's header after the app has died. A /health endpoint works if it runs one cheap database query and returns 200 only when that succeeds.

Check a page with curl

curl shows you exactly what a checker sees:

Terminal
curl -fsS -o /dev/null -w '%{http_code} %{time_total}\n' --max-time 10 https://example.com
Output
200 0.075790
  • -f exits with code 22 when the response is 400 or above; without it, a 500 page counts as success. The curl manual warns that 401 and 407 can slip through.
  • -sS hides the progress meter but keeps error messages.
  • -o /dev/null discards the page, and -w prints the status code and total seconds, even when the transfer fails.
  • --max-time 10 gives up after 10 seconds instead of hanging.

The exit code says what went wrong. When no response arrived at all, the printed code is 000:

Exit codeMeaning
6Could not resolve the host name: a DNS problem
7Failed to connect: nothing listening, or refused
22HTTP status 400 or above (only with -f)
28Timed out: --max-time was reached
35The TLS handshake failed
60Certificate not trusted. Here that covered expired, self-signed and wrong-name certificates.

To check for text, search the page:

Terminal
curl -fsS --max-time 10 https://example.com/ | grep -qF 'Example Domain' && echo "keyword found"
Output
keyword found

grep -qF matches plain text and only sets its exit code. curl follows redirects only with -L, and a redirect is not a failure, so check the final https:// address.

Check when the certificate expires

An expired certificate takes a site down as surely as a crash, and lifetimes are shrinking. Under the CA/Browser Forum's Baseline Requirements, a public TLS certificate issued from March 15, 2026 is valid for at most 200 days, from March 15, 2027 at most 100, and from March 15, 2029 at most 47. At those lifetimes renewal has to be automated, and automated renewal can fail quietly.

Terminal
openssl s_client -connect example.com:443 -servername example.com </dev/null 2>/dev/null | openssl x509 -noout -enddate
Output
notAfter=Dec 25 22:56:35 2026 GMT

-servername sends the name in SNI, so a server hosting many sites returns the right certificate, and </dev/null makes s_client hang up after the handshake. s_client prints only the server's own certificate, and x509 -noout -enddate reads its expiry date. For scripts, -checkend exits with 1 if the certificate expires within the given seconds; 14 days is 1,209,600:

Terminal
openssl s_client -connect example.com:443 -servername example.com </dev/null 2>/dev/null | openssl x509 -noout -checkend 1209600
Output
Certificate will not expire

Pick a window longer than it takes you to fix a broken renewal.

Check intervals and detection time

The interval sets how late you can learn about an outage. With three tries 20 seconds apart before alerting:

IntervalWorst case to alertChecks a month (30 days)
5 minutes5 min 40 s8,640
2 minutes2 min 40 s21,600
1 minute1 min 40 s43,200

With a 10-second timeout, a site that hangs rather than refusing adds up to 30 seconds, as each try waits out its timeout. Outages under 40 seconds never alert, which is the point. Set that against your target:

Uptime targetDowntime allowed in a 30-day month
99%7 h 12 min
99.9%43 min 12 s
99.95%21 min 36 s
99.99%4 min 19 s

At 99.9% with 5-minute checks, eight outages spend the month's allowance on detection alone. A 99.99% month allows less downtime than one 5-minute window, so a 5-minute checker cannot even measure it. One request a minute costs a site almost nothing; check anything customers pay for every minute. Detection time is part of your RTO.

Confirm an outage before you alert

Single failed checks are common and usually mean nothing: a dropped packet, a deploy restarting the app. Alerting on each one teaches people to ignore alerts.

  • Retry first. Three tries, 20 to 30 seconds apart, filter out blips shorter than about 40 seconds.
  • Check from more than one place. With three locations, alert when two agree. A failure seen from one may be a bad route near that checker.
  • Check the checker. Before declaring an outage, fetch a large site you do not run. If that fails too, the checker is offline.
  • Recover on two passes, so a flapping site sends one alert, not ten.
  • Let the checker through. A rate limit or bot filter answering a monitor with 403 or 429 looks exactly like an outage.

What makes a good alert

  • It goes to a named person, with a backup if they have not responded in 15 minutes. A shared inbox is read by no one at night.
  • Its channel matches the urgency. Email works if someone watches it or a phone sounds for that sender. A night-time outage needs something that wakes a person.
  • Quiet hours apply to warnings only. A disk at 80% can wait until morning; an outage cannot.
  • It says what failed: the error (404, curl: (28) Connection timed out, the missing text), when the first failed check ran, with a timezone, and a runbook link.
  • One message per incident, not one per failed check.
  • A recovery message says how long it was down. It stops people debugging a fixed problem and gives the status page its duration.

Heartbeat monitoring for cron jobs

A nightly backup that stopped running or a dead queue worker leaves nothing down, so nothing alerts. A heartbeat check turns it around: the job requests a unique URL each time it finishes, and the monitor alerts when a request is late.

crontab
15 2 * * * /usr/local/bin/backup-www.sh && curl -fsS --max-time 10 --retry 3 -o /dev/null https://<heartbeat-url>

&& pings only when the script exits with 0, so a failed run looks like a missing one; the script must exit non-zero when any step fails, as the cron guide shows. --retry 3 retries timeouts and responses such as 429, 500 and 503.

Set the expected period to the schedule plus a grace period. A backup that runs from 02:15 to about 02:35 expects a ping every 24 hours; with an hour's grace, a missed night alerts at 03:35. For backups, backup failure alerts goes further.

Monitor from outside, and from inside

A check running on the server it watches cannot report that the server is off, and one inside the same network cannot see a broken route, a firewall change or a DNS mistake. Run uptime checks from another provider, or at least another region. Inside checks still matter for disk, memory, processes and services with no public port. Have each one request a heartbeat URL when it runs, so a server that dies goes silent and the outside monitor notices.

A do-it-yourself checker with curl and cron

Run this on a machine at another provider. It tries three times, 20 seconds apart, checks its own connection before calling an outage, and prints one line when the site goes down and one when it comes back. cron emails whatever a job prints, so those lines are your alerts.

/usr/local/bin/check-site.sh
#!/usr/bin/env bash
# check-site.sh: check a URL from outside and print one line when it goes
# down and one when it comes back. cron emails whatever it prints.
# Usage: check-site.sh <url> <text that must be on the page> <reference url>
set -u

url="$1"
text="$2"
reference="$3"
state_dir="$HOME/.check-site"
state_file="$state_dir/$(printf '%s' "$url" | sha256sum | cut -c1-16)"
page=$(mktemp)
trap 'rm -f "$page"' EXIT
mkdir -p "$state_dir"

check() {
  local err
  err=$(curl -fsS --max-time 10 -o "$page" "$url" 2>&1) || { echo "$err"; return 1; }
  grep -qF -- "$text" "$page" || { echo "the page does not contain \"$text\""; return 1; }
}

# Three tries, 20 seconds apart. Any pass counts as up.
started=$(date +%s)
reason=""
for attempt in 1 2 3; do
  reason=$(check) && break
  [ "$attempt" -lt 3 ] && sleep 20
done

# Before calling it an outage, make sure this machine can reach the internet.
if [ -n "$reason" ] && ! curl -fsS --max-time 10 -o /dev/null "$reference" 2>/dev/null; then
  exit 0
fi

old="up"
since="$started"
[ -f "$state_file" ] && read -r old since < "$state_file"

if [ -n "$reason" ] && [ "$old" = "up" ]; then
  echo "DOWN: $url since $(date -u -d "@$started" '+%F %T') UTC: $reason"
  echo "down $started" > "$state_file"
elif [ -z "$reason" ] && [ "$old" = "down" ]; then
  echo "UP: $url is back after $(( ($(date +%s) - since + 59) / 60 )) minutes"
  echo "up $(date +%s)" > "$state_file"
fi

The state file under ~/.check-site records whether the site was up and since when, so you get one email per change and the outage's length. Make the script executable with chmod +x and add it to a crontab:

crontab
[email protected]
*/2 * * * * /usr/local/bin/check-site.sh https://www.example.com/ '<text on the page>' https://<reference-site>/

MAILTO sets the recipient. It runs every 2 minutes because a failing run can take about 80 seconds; to run it every minute, wrap it in flock as the cron guide shows. In a crontab line, escape any % as \%, or cron turns it into a newline.

cron sends mail through the machine's mail transfer agent. With none installed, Debian's and Ubuntu's cron logs No MTA installed, discarding output and the alert is lost. Test once: run the script by hand with text that is not on the page and wait for the email.

Against example.com, with missing text, then the right text, then a missing page, it printed:

Output
DOWN: https://example.com/ since 2026-10-04 02:11:04 UTC: the page does not contain "Database is up"
UP: https://example.com/ is back after 2 minutes
DOWN: https://example.com/no-such-page since 2026-10-04 02:13:19 UTC: curl: (22) The requested URL returned error: 404

A second script warns about certificates, run once a day:

/usr/local/bin/check-cert.sh
#!/usr/bin/env bash
# check-cert.sh: print a warning when a site's certificate expires soon.
# Usage: check-cert.sh <host> [days]
set -u

host="$1"
days="${2:-14}"
cert=$(timeout 15 openssl s_client -connect "$host:443" -servername "$host" </dev/null 2>/dev/null)

if ! end=$(printf '%s\n' "$cert" | openssl x509 -noout -enddate 2>/dev/null); then
  echo "Could not read the certificate for $host"
  exit 1
fi
if ! printf '%s\n' "$cert" | openssl x509 -noout -checkend $(( days * 86400 )) >/dev/null; then
  echo "The certificate for $host has expired or expires within $days days ($end)"
fi
crontab
0 8 * * * /usr/local/bin/check-cert.sh www.example.com 14

s_client has no connection timeout for TLS over TCP (-timeout is for DTLS), hence timeout 15. Neither script checks from several places, waits for two passes, escalates or keeps history: a fair list of what a monitoring service adds.

What monitoring cannot tell you

  • That the content is right. A 200 with yesterday's prices passes every check above.
  • That everyone can reach you. One location sees one network route.
  • How fast the site feels to real users. Your access logs are closer.
  • That logins, checkout, workers and email work, unless each has its own check or heartbeat.
  • That your backups restore. Only a test restore proves that.
  • Why it broke. Logs, metrics and the last deploy tell you that.

Frequently asked questions

How often should I check my website's uptime?
Every minute for anything customers pay for, every 5 minutes for the rest. With three tries 20 seconds apart, that alerts within about 1 minute 40 seconds or 5 minutes 40 seconds.
Why does my uptime monitor say the site is down when it isn't?
Usually it alerted on one failed request, its timeout was too short, its own network failed, or a firewall or rate limit answered it with 403 or 429.
What is heartbeat monitoring?
The job requests a unique URL each time it finishes, and the monitor alerts when the request is late. It catches jobs that silently stop running.
How far ahead should I alert on certificate expiry?
Far enough to fix a failed renewal calmly. 14 days is common; openssl x509 -checkend 1209600 tests exactly that.

How this was checked

The commands were run on Ubuntu 24.04 LTS, curl 8.5.0, OpenSSL 3.0.13, bash 5.2 on October 4, 2026. Any that need something this test server does not have, such as a second server, a cloud account or another database engine, were checked against the official pages below instead.

Sources, on October 4, 2026: