VPS Snaps

How to set up a status page for your service

A status page is a public page that says whether each part of your service is working, what you are doing about it when it is not, and what went wrong before. Host it away from your own infrastructure, usually on a subdomain such as status.example.com that points elsewhere with a CNAME record, list components the way customers name them, and post short, timed updates through every incident. This guide covers each part and includes update templates you can copy.

10 min readUpdated Checked against official documentation

What a status page is for

When something breaks, the first question users ask is whether the problem is theirs or yours. A status page answers it before they open a ticket, and keeps answering while you work on the fix. Done well, it does three jobs:

  • It cuts support load during an incident. One update reaches everyone, instead of a hundred tickets asking the same question.
  • It builds trust. People forgive outages more readily than silence. A page that admits problems quickly is believed when it says everything works.
  • It keeps a record. The incident history shows customers, and you, how reliable the service really is.

It is not a monitoring dashboard. Server names, internal metrics and per-host graphs belong in your own tools. The status page is written for customers, in their words.

What to put on it

  • Components. The parts of your service as customers know them: Website, Dashboard, API, Email notifications, Scheduled jobs. Five to ten is plenty. Split them by region only if customers choose a region.
  • The current state of each component, from a short fixed list (below), with a banner at the top when anything is not operational.
  • Active incidents. A title, the affected components, and every update with its time, newest first.
  • Scheduled maintenance, posted days ahead, with the start time, expected length and what users will notice.
  • History. Past incidents and maintenance, and each component's uptime over the last 90 days. Keep old incidents; a spotless history reads as an untrustworthy one.
StateUse it when
OperationalWorking normally
Degraded performanceWorking, but slow or failing for some requests
Partial outageDown for some users, regions or features
Major outageDown for most or all users
Under maintenanceUnavailable during announced, planned work

Use the same words every time. Once customers learn that a partial outage means some of them are affected, the label does half the explaining.

Host it where your outage cannot reach it

A status page that goes down with your service is worse than none: people find it broken at the moment they need it. List everything the page depends on and make sure none of it is shared with your service:

  • Servers and provider. Host the page with a different provider from your app, or at least in another region and another account. A static page in object storage at another provider is enough.
  • Load balancer and CDN. If your app sits behind one, keep the status page off it. The two should not share a front door.
  • Sign-in. Never put the page behind your own login or single sign-on. If authentication is what broke, nobody can read it.
  • Deploys. Do not ship it with your app's release pipeline. A bad release or a broken pipeline should not be able to take it down or stop you updating it.
  • DNS. Your DNS provider resolves the status page's name too. If that provider fails, status.example.com fails with everything else.

The usual setup is a subdomain with a CNAME record pointing at wherever the page is hosted:

Zone file
status.example.com.  300  IN  CNAME  <status-page-host>.

The trailing dots make both names absolute. A CNAME cannot share its name with any other record, which is why it goes on a subdomain: the bare example.com already holds the zone's SOA and NS records. The short TTL (300 seconds) lets you repoint the name quickly if the page's host has trouble of its own. Make sure that host serves a valid certificate for status.example.com, or visitors get a browser warning instead of your update.

To survive a DNS outage as well, put a second name for the page on a separate domain at a different DNS provider, such as status.example.net. List that address in your help center, support email signatures and social profiles, so people can still find it when your main domain does not resolve. Keep a copy of your DNS records, too; DNS failover covers moving traffic when a server fails.

Write incident updates people can use

Most incidents pass through four stages. Name them the same way every time:

StageMeansPost it
InvestigatingSomething is wrong; you don't know why yetWithin 10 minutes of the alert, even with no detail
IdentifiedYou know the cause and are working on a fixAs soon as you know
MonitoringA fix is in and you are watching that it holdsWhen the fix is live
ResolvedBack to normalOnce it has held, with start and end times
  • Say what users see, not what you see. “Uploads fail with an error” beats “elevated 5xx on the ingest tier.”
  • Give times with a timezone. Write 14:32 UTC, not “a few minutes ago”; the update will be read hours later.
  • Promise the next update, not the fix. “Next update by 15:00 UTC” is a promise you can keep. A fix time usually is not.
  • Update on time even with nothing new. During a major outage, at least every 30 minutes. Silence reads as nobody working on it.
  • Say what users should do: wait, retry, use a workaround, or nothing, because nothing was lost.
  • Close it properly. The resolved update gives start and end times, the duration, a one-line cause, and when a fuller report will appear.

Copy these and fill in the brackets. Delete lines that do not apply:

incident-update-template.txt
TITLE: [Component]: [what users notice]
       e.g. "Dashboard: logins failing"

INVESTIGATING  [HH:MM UTC]
We are investigating reports that [what users see] on [component].
[Who is affected: all users / users in region X / some API requests.]
Next update by [HH:MM UTC].

IDENTIFIED  [HH:MM UTC]
We have found the cause: [one plain sentence].
We are [what you are doing]. [Workaround, if there is one.]
[No data has been lost. / Requests sent since HH:MM UTC need to be retried.]
Next update by [HH:MM UTC].

MONITORING  [HH:MM UTC]
A fix is in place and [component] is working again. We are watching
to make sure it holds. [Anything users must do, e.g. reload the page.]
Next update by [HH:MM UTC].

RESOLVED  [HH:MM UTC]
This incident is resolved. [Component] was [unavailable / slow] from
[HH:MM] to [HH:MM UTC] ([duration]). Cause: [one sentence].
We will publish a full report by [date].
maintenance-notice-template.txt
TITLE: Scheduled maintenance: [component], [date]

SCHEDULED  (post at least [3] days ahead)
On [date], from [HH:MM] to [HH:MM UTC], we will [what you are doing, in
one plain sentence]. During this window [what users will notice: the
dashboard is read-only / the API returns errors / no impact expected].
[What users should do beforehand, if anything.]

IN PROGRESS  [HH:MM UTC]
Maintenance has started. [Component] is [read-only / unavailable].
Expected to finish by [HH:MM UTC].

COMPLETED  [HH:MM UTC]
Maintenance is complete and [component] is working normally.

Worked example: one incident, start to finish

A database disk fills at 14:02 UTC and the dashboard starts failing. With 1-minute checks and three tries 20 seconds apart, the alert arrives at about 14:04. Here is the incident as the status page shows it, newest update first:

example-incident.txt
Dashboard: logins and saves failing

RESOLVED  15:21 UTC
This incident is resolved. The dashboard was unavailable from 14:02 to
14:58 UTC (56 minutes). Cause: the database server ran out of disk
space. We will publish a full report by Friday.

MONITORING  14:58 UTC
A fix is in place and the dashboard is working again. Changes made
between 14:02 and 14:58 UTC were not saved; please make them again.
Next update by 15:30 UTC.

IDENTIFIED  14:31 UTC
We have found the cause: the database server ran out of disk space.
We are moving the data to a larger disk. Nothing saved before 14:02 UTC
has been lost.
Next update by 15:00 UTC.

INVESTIGATING  14:09 UTC
We are investigating reports that logins and saves are failing in the
dashboard. The API and scheduled jobs are working.
Next update by 14:30 UTC.

Every entry is timed and each one promises the next. The first post went up five minutes after the alert and said almost nothing, which is fine: “we know” is the update people are waiting for. The dashboard component went to Major outage at 14:09 and back to Operational at 14:58, and those 56 minutes stay in its 90-day history. The incident stayed in Monitoring for 23 minutes before anyone called it resolved.

  • Your app. A footer link, and a banner inside the app while an incident is open. Load the status in the browser with a short timeout and show nothing if it fails, so the status page can never slow down or break your app.
  • Error pages. Your 502 and 503 pages are what people see while the app is down; put the status page address on them. Serve those pages from the web server or load balancer as static files, not from the app.
  • Help center and support. Link it at the top of the help center and in support auto-replies during an incident.
  • API docs. Developers check the status page first when their requests fail.
  • Social profiles. Many people look there first. Post a short note linking to the status page instead of running the incident in two places.

Automatic or manual updates

Component states can be set by monitors, by people, or by both:

Set by monitorsSet by hand
SpeedWithin a minute or two of a confirmed failureWhen someone gets to it, often 10 to 30 minutes in
AccuracyShows false alarms in public, and misses what the checks don't coverAs good as the person's view of the incident
WordsUp or down, nothing moreWhat users see and what to do
RiskFlaps when checks flap; can call a recovery too earlyForgotten: the page shows green during an outage

A mix works best. Let confirmed monitor failures set a component's state and open a draft incident, and have a person write and publish every word. Map each component to the checks behind it, so one flaky internal check cannot mark a customer-facing component down. When a monitor reports recovery, leave the incident in Monitoring until a person resolves it.

Common mistakes

  • Hosting it on the infrastructure it reports on. The first real outage takes it down too.
  • Green during an outage. Nothing loses trust faster than “All systems operational” while customers cannot log in. If you are not sure yet, say you are investigating.
  • Internal names. “db-primary-2 degraded” means nothing to a customer. Name what they use.
  • Vague updates. “Some users may be experiencing issues” says nothing. Say which users and which issues.
  • No next-update time, or a missed one.
  • Times without a timezone.
  • Deleting or rewriting old incidents. Keep the history, embarrassing ones included.
  • Forgetting to resolve. An incident left open for a week makes the page useless the next time it matters.
  • Never checking it. Open the page from your phone, off your own network, now and then and after any change to its hosting or DNS.

Frequently asked questions

Should a status page be on a subdomain or a separate domain?
Use a subdomain such as status.example.com with a CNAME to a host outside your infrastructure. To stay reachable when your DNS provider fails as well, add a second name on a separate domain at another DNS provider.
How often should I update a status page during an incident?
Post within about 10 minutes of the alert, then at least every 30 minutes during a major outage, and always before the next-update time you gave.
Should a status page update automatically from monitoring?
Let confirmed monitor failures change a component's state, but have a person write the incident text and decide when it is resolved.
Does a small project need a status page?
If people depend on your service and contact you when it is down, yes. One page with three components and honest updates is enough.
What should a status page say about scheduled maintenance?
The start and end times with a timezone, what users will notice, and anything they should do first, posted days ahead. Update it when the work starts and when it ends.

How this was checked

Commands, limits and prices were checked against these official pages, on October 4, 2026: