Platform

What is actually in the rack

Two CPU generations, no third. Memory that corrects its own errors on the line where that matters, enterprise NVMe in mirrored pairs with a spare sitting idle beside it, and three days of burn-in before a node is allowed to meet a customer.

01

Two generations, and nothing older

Most hosts run three or four generations at once and sell them at the same price, which is why two servers with identical specifications can differ by forty percent on the same workload. We keep the fleet narrow instead.

Nothing in production is more than one generation behind current. That rule has held since 2022.

01

Ryzen 9 9950X, Zen 5

16 cores, 32 threads, boosting to 5.7 GHz. This is the fastest single thread we can buy, and single-thread speed is still what decides how quickly a request finishes on most real applications.

02

EPYC 9354P, Zen 4

32 cores and 64 threads on one socket, twelve memory channels, registered ECC. Where core count, memory capacity and predictable NUMA behaviour matter more than peak clock, this is the part that wins.

03

One socket wherever possible

Cross-socket memory access is a latency cliff you cannot tune your way out of. Only the largest bare-metal configuration is dual socket, and it is dual socket because the customer explicitly asked for 64 cores.

04

Cores are allocated once

A core sold to you is scheduled for you. No burst credits, no oversubscription ratio, no steal time appearing under load and disappearing before you can screenshot it.

05

Why not the newest server silicon

The Zen 5 server parts are in the lab and not in the fleet. When they replace the 9354P it will appear on the changelog with migration dates, rather than as a phrase on the front page six months before anyone can buy one.

02

Memory, and where ECC actually lives

Memory errors are not rare; they are merely invisible until something important becomes subtly wrong. On the server line every one of them is corrected and recorded.

EPYC nodes carry twelve DDR5-4800 registered ECC modules, one per channel, all populated. Populating every channel is not about capacity, it is about bandwidth: a half-filled memory bus on a 32-core part is a bottleneck you will feel on anything that touches RAM in a hurry.

Ryzen nodes run DDR5-5600 without ECC, because the consumer platform does not offer it in a form we would rely on. That is stated plainly here rather than buried: if your workload needs error correction, the EPYC line starts at eighty-nine euro and exists for exactly that reason.

01

Correctable errors are counted

A rank producing more than a handful of corrections in a day is flagged, and the module is swapped at the next maintenance window whether or not anything has misbehaved.

02

Uncorrectable errors drain the node

Instances migrate off, the node leaves the pool, and the DIMM is replaced before it comes back. You may see a live migration; you should not see a crash.

03

No ballooning, no page merging

Memory is allocated once and not counted twice. Same-page merging is off across the fleet, because it is a side channel between tenants and a performance claim that evaporates under real load.

04

Capacity per node is generous on purpose

An EPYC node ships with enough memory that the largest plan it hosts still leaves headroom. Nodes packed to the last gigabyte have nowhere to put a migration when a neighbour has to move.

03

NVMe, endurance and the spare that does nothing

Every drive is an enterprise Gen4 part with power-loss protection, which is a capacitor bank that finishes writing what it acknowledged when the power goes away mid-flush.

A hot spare per node, idle, waiting. It is the cheapest insurance in the building.

01

Mirrored pairs, always

Reads are spread across both halves; writes are acknowledged from protected cache and land on both. A single drive failure changes the performance profile slightly and changes your availability not at all.

02

A hot spare per node

One drive per node sits idle and empty. When a mirror loses a half, the spare takes over automatically and the rebuild starts before anyone has read the alert.

03

Endurance headroom, not endurance limits

General nodes use 1 DWPD class drives; storage nodes use 3 DWPD for the write tier. Drives are replaced at eighty percent of rated endurance rather than being run to the number on the datasheet.

04

No consumer drives anywhere

Not in the cache tier, not in the bulk tier, not in a node someone built quickly in 2021 and forgot about. Consumer NVMe is fast for eleven seconds and then it is not, which is a fine trade in a laptop.

05

Capacity is allocated, not promised

Storage is not thin-provisioned across tenants. If the plan says 800 GB, 800 GB exists on that node with your name on it, and nobody else can grow into it.

04

What is in each class of node

Five node classes carry the five product lines. Bulk storage on the S class is enterprise SATA behind an NVMe write tier, which is the only spinning media anywhere in the fleet.

ClassCPUMemoryNVMeNetworkPower
R — Ryzen NVMeRyzen 9 9950X, 16C/32T128 GB DDR5-56002 × 3.84 TB mirrored, 1 spare2 × 10 GbE, LACP2 × 800 W, A and B
E — EPYC High-MemoryEPYC 9354P, 32C/64T12 × DDR5-4800 ECC RDIMM4 × 3.84 TB mirrored, 1 spare2 × 25 GbE, LACP2 × 1200 W, A and B
G — GPUEPYC 9354P, 32C/64T256 GB ECC DDR5-48004 × 3.84 TB mirrored, 1 spare2 × 40 GbE, LACP2 × 2000 W, A and B
S — StorageEPYC 9354P, 32C/64T128 GB ECC DDR5-48002 × 3.84 TB write tier, SATA bulk behind it2 × 25 GbE, LACP2 × 1200 W, A and B
BM — Bare MetalRyzen 9950X, or one or two EPYC 9354P128 GB to 1 TBUp to 8 × 7.68 TB10 to 40 GbE2 × 1200 W, A and B

GPU nodes pass whole physical cards through on PCIe 4.0 x16 with no time-slicing, which is why the power figure is what it is. Two cards in one instance sit on the same NUMA node, because splitting them across sockets would cost more than the second card gains.

05

Power feeds, and management that is not on the internet

Two of everything that can be doubled, and a management network that has never had a route to the outside world.

01

A and B feeds, separate paths

Two supplies per node on independent distribution, each sized to carry the whole machine alone. Losing a feed is a logged event and nothing else, and we prove that on every node during burn-in.

02

Battery, then generator

Battery covers the transfer; generators cover the rest. Transfer is tested on a schedule at every site, and a site that fails a test does not receive new nodes until it passes one.

03

No tier number quoted

Ratings belong to buildings and to the operators who commissioned them, not to a tenant repeating a number in marketing copy. What we will tell you, per site, is the feed topology and the last transfer test result.

04

Out-of-band on its own VLAN

Management interfaces have no public route and never have had one. Access goes through an authenticated tunnel, is logged, and is available to you as well as to us.

05

Console included, at no cost

Serial and VNC to your instance through that tunnel. It works when your network configuration does not, which is the only moment anybody ever wants it.

06

IPMI on metal, credentials per term

Bare-metal customers get IPMI over a VPN with power control, boot order and virtual media. Credentials are rotated when a term ends, and the controller firmware is patched on our schedule rather than yours.

06

Burn-in: three days before it meets a customer

About one node in nine fails something on the first pass. It is nearly always a single DIMM.

  1. 01

    Firmware to a known baseline

    System, controller, network and drive firmware are pinned to the versions the fleet runs before any test is allowed to start. A node tested on the wrong firmware has tested nothing useful.

  2. 02

    Memory, at length

    Multiple full passes across every populated module, hours rather than minutes. This stage catches most of what burn-in catches, and it is the reason the schedule is measured in days.

  3. 03

    Thermal soak

    Every core at full load for six hours with the inlet at the warm end of its range. We watch sustained boost clocks decay, not just temperature, because a node that thermally throttles at hour five would do it to you at month three.

  4. 04

    Full-surface storage verification

    Each drive written end to end, read back and compared. SMART counters at the end of this stage become the baseline the drive is judged against for the rest of its life.

  5. 05

    Line rate, both directions at once

    Every port driven to capacity in both directions for an hour, with switch-side counters checked for errors the host would never notice.

  6. 06

    Pull a feed, then the other

    A feed removed, restored, then the same on the other side. Any node that so much as blinks does not enter service.

  7. 07

    A quiet day in the fleet

    Twenty-four hours in monitoring carrying no instances at all. Then it starts taking customers.

Total elapsed time is around 72 hours. It is also why stock at a new site appears in batches rather than trickling in: a rack arrives, and three days later all of it is available at once.

07

When a drive starts to fail

Drives do not usually die suddenly. They announce it for days in counters nobody reads, which is why ours are read every five minutes.

  1. 01

    Thresholds, not eulogies

    Media errors, reallocated sectors, wear level and latency outliers are sampled continuously. A drive crossing any threshold is marked for replacement while it is still working perfectly well.

  2. 02

    The spare takes over first

    The hot spare is brought into the mirror before the marked drive is removed, so the pair is never running as a single copy while anyone waits for a technician.

  3. 03

    Rebuild, throttled on purpose

    A 3.84 TB mirror could rebuild in about forty minutes at full speed. We run it slower, over roughly two hours, because rebuilding at full speed is indistinguishable from a performance incident to the instances on that node.

  4. 04

    You are told, and nothing else changes

    A note appears on the instance page. No reboot, no migration, no maintenance window, and no ticket you have to answer.

  5. 05

    The drive leaves in pieces

    Failed and retired drives are destroyed on site rather than returned for warranty credit. The credit is worth less than the certainty.

A mirror is redundancy, not a backup. It protects you from a drive, not from a bad migration or an enthusiastic delete. Snapshots cost five euro for five slots and off-node backup nine euro per 500 GB, written to a different site from the instance.

08

Hardware questions

No. Ryzen nodes run non-ECC DDR5-5600, and we would rather say that here than let you find out from a datasheet later. Anything where silent memory corruption is unacceptable belongs on EPYC.

Yes, and it is checkable. Run a sustained single-core benchmark at three in the morning and again at peak; the numbers should not move. If they do, open a ticket, because something is wrong and we would want to know.

Not by name, but you can ask for anti-affinity: two instances of yours placed on separate nodes, separate switches and separate power feeds. It costs nothing and takes one ticket.

Where the work allows it, instances live-migrate and you see a few hundred milliseconds of pause. Otherwise you get at least five days of notice, with a window you can move once if the timing is wrong for you.

No. Cores, memory and NVMe are allocated once each. The only shared resource is the uplink, which is why the fair-use figure is published rather than implied.

Every node in production runs Zen 4 or Zen 5. The last of the previous generation left the fleet in 2022, and nothing has been sold on older silicon since.

Ready when you are

Same hardware, thirty-four cities

The node specification does not change with the flag on the building. Pick the jurisdiction and the latency you want, then pick the size.