What is broken, and since when
The status page is fed by the same probes that wake the on-call engineer. No editorial step sits between a failing check and a red panel, which is deliberate and occasionally embarrassing.
90 days · Latency figures are measured medians from our own probes, not vendor brochures. They move; the page updates when they do.
Journal
An engineering diary. Incidents get a post-mortem with real timestamps, hardware decisions get the numbers that made them, and policy changes get the reasoning even when the reasoning makes us look slow. Nothing here was written by a marketing department, which becomes obvious inside a paragraph.
What the page shows
Probes run every 30 seconds from outside our own network. History stays up for 90 days.
Per site, not per fleet
All 34 sites are listed separately. A problem in one city does not colour the others, because an aggregate availability figure hides exactly the outage you came to the page about.
Six components per site
Network, hypervisor, storage, provisioning, panel and API, each checked independently. A provisioning queue backing up therefore does not claim that your running instances are down.
Measured from outside
Probes sit on networks we do not operate, in several regions, and a check has to fail from more than one before anything changes colour. Monitoring a network from inside it measures very little.
The same uptime figure as the SLA
99.993% across the last 12 rolling months, calculated from these probes rather than from a friendlier definition. Downtime counts from the first failed check, not from the moment somebody acknowledged it.
Post-incident notes
Anything graded P1 or P2 gets a written follow-up within 5 working days: what broke, what was done, what changed afterwards. They are dull on purpose and they are never quietly deleted.
How incidents are graded
Four levels, assigned by the on-call engineer within minutes and revised upward without hesitation when the picture changes. Nothing is graded down afterwards to make a report read better.
| Level | Meaning | First update | Then |
|---|---|---|---|
| P1 | Customer instances down or unreachable at a site, or a fault touching more than one site. | Within 15 minutes | Every 30 minutes until resolved |
| P2 | Severe degradation. Packet loss, storage latency, or a component down while instances keep serving. | Within 30 minutes | Hourly |
| P3 | Control-plane fault. Provisioning, the panel, the API or billing unavailable while running instances are unaffected. | Within 2 hours | Twice a day |
| P4 | Cosmetic or single-instance. A stuck queue item, a wrong graph, one node degraded with a hot spare already carrying it. | Next working day | On resolution |
A single customer instance failing is a P4 on this page and a ticket with an 11-minute median response in the panel. The grade describes blast radius rather than how much it matters to you.
Maintenance
Planned maintenance is announced at least 7 days ahead, by mail to the affected accounts and on the status page. Resellers get 14 days. Every notice names the site, the window, the expected impact and whether a reboot is involved.
Windows run from 01:00 to 05:00 local time at the site, which is the only reasonable choice when the fleet spans 29 countries and somebody is always awake. Most work finishes inside the first hour.
Where a workload can be live-migrated it is, and the visible effect is a few seconds of additional latency rather than a restart. Firmware, kernel and hypervisor work cannot be done that way, and the notice says so plainly instead of hiding a reboot behind the word “brief”.
Emergency maintenance
Announced as soon as it is decided, sometimes with under an hour of notice. Reserved for a security fix under active exploitation, or hardware that is going to fail either way.
Deferrals
One per instance per quarter, by ticket, up to 14 days. Beyond that the node has to be done, and we will help you plan around the date rather than move it again.
What never happens in a window
Price changes, policy changes, and anything touching your data. Maintenance covers hardware and software we operate, never the contents of your disks.
Being told about it
Four ways to subscribe. None of them requires an account, and none of them will ever be used to send you anything other than incidents.
Atom feed
One feed for everything, or one per site. It is the option that keeps working when mail does not, which during a network incident is precisely the situation you are in.
Email per site
Subscribe an address to the cities you use. Customers are subscribed automatically to their own sites and can switch it off, though we would rather you did not.
Webhooks
A signed POST on every state change, for anyone routing incidents into their own tooling. Same payload shape as the API, retried with backoff for 24 hours.
The status API
A public read-only endpoint returning current state and open incidents as JSON. No token needed, rate-limited politely, and unchanged since 2024.
That is the whole use of the address. There is no newsletter, no product announcement mailing, and no marketing list that a status subscription quietly enrols you into.
Status questions
Because it tracks sites and components rather than individual instances. One node or one instance is a ticket, not a public incident, and the ticket is by far the faster route to a fix.
Deliberately not on the infrastructure it watches. It is served from a separate site with separate transit, so a network fault cannot take down the page describing that fault.
Announced windows do not count against the SLA figure. Everything else does, including the outages that were our fault and the ones that were somebody else’s.
Ninety days on the page itself, and indefinitely for P1 and P2 write-ups. Old incidents are not removed once they stop being flattering.
Each site page carries probe round-trip times from the four reference hubs, updated continuously. The figures printed on the location pages are medians of that same data.
Subscribe before you need it
The feed and the per-site mail take one address and about ten seconds. Doing it during an incident is possible and considerably less pleasant.