Kept since October 2019

Journal

An engineering diary. Incidents get a post-mortem with real timestamps, hardware decisions get the numbers that made them, and policy changes get the reasoning even when the reasoning makes us look slow. Nothing here was written by a marketing department, which becomes obvious inside a paragraph.

01

What ends up in here

We were writing these notes anyway. Somebody has to explain to the next engineer why the two halves of a mirror come from different delivery boxes, and once that is written down there is no good argument for keeping it behind a login.

Five categories. Product announcements dressed as insight are not one of them.

You will not find a tutorial here. Guides belong in [the guides section](/guides) and reference material belongs in [the knowledge base](/docs), and mixing the three produces a blog that nobody can search. Current fleet state is on [the status page](/status), which updates without anyone writing prose about it.

01

Incident

Something broke and customers noticed. Published within ten working days, timestamps in UTC, cause named in the first paragraph. When the cause was us, the post says so before it says anything else.

02

Hardware

What we bought, what we measured, and what we turned down. Generation migrations, drive families, and the occasional line that stays a generation behind on purpose because the newer part does not help the workload.

03

Network

Transit, peering, filtering, addressing. This is also where the numbers other hosts keep vague get printed, starting with what unmetered means before anyone emails you about it.

04

Platform

The panel, the provisioning queue, the API, new sites. Capacity notes live here too, including the unglamorous ones explaining why a particular city is rationed rather than expanded.

05

Policy

Changes to what we sell, what we keep, and what we refuse to do. Deleting the cheap tier is still the most-read post on this site, five years after it was written.

02

How an incident gets written

A house style is not an aesthetic choice. It stops us quietly reorganising a bad week into a better-looking story, because the section that would have to be deleted is conspicuous by its absence.

Same five headings, same order, every time.

  1. 01

    Summary

    What broke, for how long, and who was affected, in under sixty words. Read nothing else and you should still know whether it touched you.

  2. 02

    Timeline

    UTC timestamps from first symptom to full recovery, including the minutes spent looking in the wrong place. Those minutes are usually the interesting part of the document.

  3. 03

    Root cause

    The mechanism, not the category. Human error is not a root cause. A classification rule written in 2023 that let a schema change skip the canary stage is a root cause, and it has an owner.

  4. 04

    What we changed

    Concrete, dated, checkable. Every item is something a customer could ask us to demonstrate on a call, and twice somebody has.

  5. 05

    What we did not change, and why

    The section most companies leave out. Sometimes the obvious fix costs more privacy than the fault cost you, and we would rather argue about that in public than settle it quietly.

It is three in the morning here and your status page says the site is fine. The site is not fine.
A ticket, opened during the October 2020 outage

Post-mortems get published even when nobody asked and even when the affected customer never noticed. Two of them describe credits we paid that the SLA did not oblige us to pay, which is the sort of thing you only write down once.

03

Reading the archive

Sixteen posts, six and a half years, one conspicuous gap.

The gap runs from late 2019 to October 2020. We were racking, not writing, and the thing that finally produced a post was an attack that took Amsterdam off the network for eleven minutes. That is a common pattern in this trade: the documentation habit starts the day something expensive goes wrong.

Old posts are never edited in place. Where a claim turned out to be wrong, the correction is appended with its own date and the original sentence stays exactly where it was. A post that quietly improves itself over time is worth nothing to the person reading it.

PeriodWhat we were doingWhat is written down
2019 – 2020Four sites lit in eight weeks, then six more across a pandemic yearTwo posts. One founding note and one apology.
2021 – 2022Cheap tier deleted, IPv6 by default, Zen 4 ordered, Reykjavík openedFour posts, including the one we get quoted on most.
2023 – 2024Canary, GPUs, Zen 5, and a drive batch that failed togetherFive posts, one of them an apology with a refund attached.
2025 – 2026Provisioning rewritten, Zurich rationed, the API opened, Johannesburg rackedFive posts and one more post-mortem.

Platform

05

Network

03

Incident

03

Hardware

03

Policy

02
04

Questions about the journal itself

An Atom feed, linked at the foot of this page, with no tracking pixel in it and no accompanying email list. If you want a newsletter you will have to want one somewhere else.

No. Everything here was written by somebody who can be paged, which is the entire point of it existing.

Because deleting them would be dishonest. Corrections are appended and dated underneath the original text, so you can see both what we thought and when we stopped thinking it.

Yes, including the parts that make us look bad, and especially those. Attribution and a link is enough; no permission is required and none will be granted more formally than this sentence.

If more than one customer was affected, it is already here or it is being drafted. Single-instance faults are answered in the ticket, because a post-mortem about one dead drive helps nobody.

Ready when you are

The hardware is more interesting than the writing

Every claim in these posts is testable on a server you can have in under a minute. Pick a city, pick a size, pay in coin, and check the numbers yourself.