Monitoring & alerts
Know when a site is down or a backup has stopped: uptime and certificate checks, heartbeats for scheduled jobs, backup failure alerts and status pages.
How to monitor website and server uptime
What to check (status code, page text, certificate expiry, ports, DNS), how often, how to confirm an outage before alerting, heartbeat checks for cron jobs, and a curl-and-cron checker you can run yourself.
11 min read · TestedHow to set up a status page for your service
What a status page is for, what to put on it, where to host it so it stays up when your service does not, how to write incident updates (with templates to copy), and the mistakes that make people stop trusting it.
10 min readHow to know when a server backup fails
Catch failed, missing and empty backups: strict bash settings and pipefail, a heartbeat that alerts when success pings stop, a freshness and size check on the bucket, cron mail, and how to prove each alert fires.
11 min read