Kept since October 2019

Moving the whole fleet to Zen 4

Every production node moves to current-generation EPYC or Ryzen silicon over the next nine months, one node at a time, with announced windows and no quiet reboots.

Starting in January, every node in the fleet moves to current-generation silicon: EPYC 9004 series on the server and storage lines, Ryzen 7000 series on the general line. The plan runs about nine months. Nothing older than the current generation will be carrying customer workloads by the end of it.

This is the largest thing we have ever done to the running fleet and it is worth explaining before it starts rather than afterwards.

Why now

The existing fleet is a mixture of two previous generations, which was defensible in 2020 and stopped being defensible some time last year. Three things changed at once: DDR5 with ECC became available in the capacities the server line needs, memory bandwidth per socket went up enough to matter for the memory-bound customers, and the per-core numbers on the older parts stopped being something we wanted to print next to a price.

There is also a duller reason. A fleet holding three generations needs three sets of spares, three qualification procedures and three sets of firmware to track. Consolidating to one removes an entire category of two-in-the-morning surprise.

The order

Amsterdam first, because everything ships to AMS-01 first and always has. Frankfurt, London and New York follow, then the remaining European sites, then Asia-Pacific, then the Americas. Roughly fourteen nodes a week once the process settles, less at the start while we find out what we got wrong.

Sites opening during the migration open directly on the new generation and never see the old one.

What you will experience

Live migration where the workload supports it. Most instances move between hosts with a pause measured in low hundreds of milliseconds. You will see it in a graph if you are looking and not otherwise.

A scheduled reboot where it does not. Instances using GPU passthrough, custom kernels with pinned devices, or anything holding hardware state cannot be moved live. Those get a window.

Five working days notice, by email, with a specific window. Not a month of vagueness. If we miss an announced window, the SLA credit applies automatically without anybody having to open a ticket about it.

No price change. The new generation costs the same as the old one did. We do not sell a generation surcharge and we do not keep the old hardware around at a discount for customers who did not read the email. It leaves the building.

The one thing to check

The CPU feature set changes, and AVX-512 is present on the new server parts where it was not before. Compile with feature detection and nothing happens. If you have pinned a build to a specific flag set, or you run numerical code that dispatches on detected features at start-up, test before your window rather than after it.

We will list the exact flag differences in the migration notice for your site. About one customer in two hundred needs to do anything at all, and the ones who need to already know who they are.

What has already gone wrong

A batch of registered ECC DDR5 modules for the first four server nodes trained at a lower speed than specified and it took nine days to work out that the modules, not the boards, were at fault. Those nine days are the reason this announcement is in December rather than October.

Supply is the risk for the whole plan. Server memory in these capacities is not something you can source at short notice, and if the schedule slips it will slip because of memory rather than because of anything we control. We will say so here if it does.

What we are not doing

We are not offering an opt-out. A fleet with a long tail of customers who declined to move is a fleet with an old generation in it forever, and the tail becomes the thing nobody wants to maintain. Everything moves.

We are also not taking the opportunity to change plan specifications. The same cores, the same memory, the same storage, on faster silicon, at the same price. Any change to what a plan contains will be a separate announcement with its own reasoning attached, because bundling a specification change into a hardware migration is how customers stop trusting migration notices.

Ready when you are

Pick a city. Pick a size. Pay in coin.

No forms about who you are, no wait for a human to approve you, no phone call to verify anything. The invoice clears and the credentials land in your inbox.