Kept since October 2019

Post-mortem: AMS-01, eleven minutes, 14 October 2020

A 340 Gbit/s attack aimed at one customer took the entire Amsterdam site off the network for eleven minutes. The filtering we have run ever since exists because of it.

Summary

On 14 October 2020, between 02:41 and 02:52 UTC, every instance at AMS-01 was unreachable. A volumetric attack aimed at a single customer address peaked at 340 Gbit/s and saturated the site’s uplinks. We resolved it by asking upstream to discard all traffic to the target address, which restored the site and left that one customer offline for a further fifty minutes. No data was lost and no instance was harmed. The site had no scrubbing capacity at the time, and that was a purchasing decision rather than an accident.

Timeline

All times UTC, 14 October 2020.

TimeEvent
02:41:04Traffic to a single customer address rises from about 200 Mbit/s to 90 Gbit/s in under twenty seconds.
02:41:30Both site uplinks saturate. Packet loss becomes total across every prefix at the site, not only the target.
02:42Automated alerting fires on reachability from three external probes.
02:44On-call engineer online. Attack is visible on the port graphs and nowhere else, because our flow collection was sampled at a rate that could not keep up.
02:46Peak measured at 340 Gbit/s. Composition is reflected UDP, mostly DNS and NTP, from tens of thousands of sources.
02:47Decision made to request an upstream discard of all traffic to the target address. This was the only tool available to us.
02:49Discard propagates at the first upstream.
02:51Second upstream applies it. Uplink utilisation falls below capacity.
02:52:10Site fully reachable. Total customer-visible outage, eleven minutes and six seconds.
03:20Affected customer contacted and offered a new address.
03:42Customer moved to a new address; their service returns.

Root cause

The proximate cause was an attack we could not absorb. What actually caused it is that we sold hosting in a city where attacks of this size were routine, with two hundred gigabits of uplink and no filtering whatsoever, and treated an upstream discard agreement as a plan.

That is not bad luck. It was a decision, made in 2019 in favour of spending the money on hardware, and it was wrong.

Two secondary faults made it worse. Our flow collection sampled at a rate that gave us no useful detail during a real event, so the first four minutes were spent reading interface counters. And the discard was manual, requiring a human to be awake, authenticated and confident, which is three requirements too many at three in the morning.

What this cost the customer

Being blunt about it: we fixed our problem by turning their service off. A discard route is a decision that one customer will be unreachable so that everybody else is not. It was the correct call with the tools we had, it was still their outage rather than ours, and they had bought no product from us that promised otherwise.

They stayed. We did not charge them for October.

What we changed

  1. Always-on scrubbing at the edge, purchased within three weeks, starting at 1.2 Tbit/s of capacity. Filtering sits permanently in the path, so there is no detection delay and no button for anyone to press at 02:47. This is now standard at every site, up to 12 Tbit/s at the largest.
  2. Uplink at AMS-01 raised, first to 200 Gbit/s of additional capacity and later to the 400 Gbit/s the site runs today.
  3. Unsampled flow telemetry on every site edge, retained for operational purposes for seven days. This is traffic metadata for our own ports; it is not instance traffic and it is not the contents of anything.
  4. The discard route became automatic, with a documented threshold, an announced policy, and a mail to the affected customer within sixty seconds rather than thirty-nine minutes.
  5. Attacks are published on the status page, with size and duration, whether or not anybody noticed.

What we did not change, and why

We did not start charging for filtering. Base scrubbing is included on every plan at every site and always has been since the day we bought it. An attack is not a service the victim asked for. Layer-7 filtering with custom rules exists as an add-on because it needs a human on our side, and that is a different thing from a volumetric flood.

We did not remove the discard route. Twelve terabits per second is a number, not infinity, and pretending we will never again need the blunt instrument would be dishonest. What changed is that it is now the documented last resort rather than the only step.

We did not impose per-customer traffic ceilings. Rate-limiting every customer to a safe fraction of the uplink would have prevented this and would also throttle every legitimate spike. The filtering belongs at the edge, on the attack, not on the customer.

We did not move the customer to a different product. They were running the thing they had paid for, on the plan that suited them, and it was not their fault that somebody aimed a botnet at them.

Ready when you are

Pick a city. Pick a size. Pay in coin.

No forms about who you are, no wait for a human to approve you, no phone call to verify anything. The invoice clears and the credentials land in your inbox.