Uptime monitoring: monitors, heartbeats and alerts
Watch your sites, ports and scheduled jobs, be told once when something is down, and find the way back in the alert.
Monitoring is in the sidebar on the Starter plan and above. Add monitor, choose what to watch, and Check it now runs one check before you save, showing exactly what the checker sees.
Three kinds of monitor
- A web page: up when it answers with a status code you accept (200 to 399 unless you say otherwise) and, if you give one, when a word is on the page, or is not, for an error message. Redirects are followed, up to five. Every https page's certificate is watched too.
- A port: up when it accepts a connection. For a database, SSH or a mail server.
- A heartbeat, for a scheduled job: it gets a URL to call when it finishes, and is down when no call arrives within the job's period and the grace you allow.
When a monitor counts as down
One failed check alerts nobody. It is checked again 20 seconds later, and again, and only the third failure in a row is an outage, dated from the first. Before a failure counts, the checker makes sure it can reach the internet itself, and if many sites fail at once it treats the fault as its own and holds every alert until it clears. A monitor is back up after two successful checks in a row.
Heartbeats for cron jobs and scripts
Open the heartbeat's page for its URL. Call it at the end of the job, and only when the job succeeded, so a failure is a missed call:
0 3 * * * /usr/local/bin/backup.sh && curl -fsS -m 10 --retry 3 https://vpssnaps.com/api/heartbeat/<your-token> > /dev/nullAny method works: GET, POST or HEAD, with curl or wget. The answer is always an empty 204, whether or not the URL exists, so check the monitor's page to see its last ping. Anyone with the URL can report the job ran, so replace it from the monitor's page if it has been shared.
Alerts
Each outage sends one alert when it is confirmed and another when it is over, saying how long it lasted. They go where your other alerts go: in Settings → Notifications, add destinations for "A monitored site or server is down", "…is back up" and "…certificate expires soon" (14 and 3 days before). With none set, they go to the workspace owner's email. Each monitor can have its alerts switched off.
Linking a monitor to a server
Choose a server when you add a monitor, or open a server and press Monitor this server, which adds its SSH port and, if its recovery plan has a health URL, its site. An outage of a linked monitor names the server's newest backups and how old they are, and on its page offers Recover now when the server has a recovery plan.
Allowing our checks through a firewall
Checks come from 165.227.255.101, with the User-Agent "VPS Snaps Uptime (+https://vpssnaps.com)". Allow that address if your firewall or a security plugin blocks visitors it doesn't know, or every check will fail.
Plans
Starter: 10 monitors, checked every 5 minutes, with 30 days of history. Standard: 30, every minute, 90 days. Pro: 100, every minute, a year. Agency: 300, every minute, a year. Every kind of monitor counts toward the number. The free plan has no monitoring; a workspace that moves to it keeps its monitors, paused.
Checks come from one place, New York, so a network fault between there and your server can look like an outage. There are no text messages or phone calls, no checks of pages behind a login, and no ping or DNS checks yet.