A card, not a slice of one
A shared GPU is somebody else’s job finishing before yours starts. Every card in this line is passed through to one instance over a full sixteen-lane link and stays there until you cancel, idle or not.
A shared GPU is someone else’s job finishing before yours starts. Every card in this line is passed straight through to a single instance and stays there until you cancel.
| Plan | Dedicated cores | Memory | NVMe storage | Graphics | Port speed | Price | |
|---|---|---|---|---|---|---|---|
| G-ADA Inference and light training | 8 | 64 GB | 1 TB | 1 × RTX 4000 Ada · 20 GB | 40 Gbit/s | €239.00 /mo | Configure |
| G-L40SMost taken 48 GB of VRAM, undivided | 16 | 128 GB | 2 TB | 1 × L40S · 48 GB | 40 Gbit/s | €749.00 /mo | Configure |
| G-L40S×2 Two cards, one NUMA node | 32 | 256 GB | 4 TB | 2 × L40S · 96 GB | 40 Gbit/s | €1420.00 /mo | Configure |
Prices exclude VAT where applicable. Traffic: Unmetered, fair use.
EPYC 9354P host · PCIe 4.0 x16 passthrough · no vGPU time-slicing
Thirty-four cities, twenty-nine countries
Latency figures are measured medians from our own probes, not vendor brochures. They move; the page updates when they do.
Passthrough, and why it matters
Most cheap GPU capacity is a fraction of a card handed round a queue. The number on the product page is the card’s memory, not yours, and the throughput you measure depends on who else woke up this morning.
No vGPU, no time-slicing, no scheduler between you and the silicon.
Here the card is bound to your instance at the hypervisor level over a sixteen-lane PCIe 4.0 link. The whole device is yours: all of the memory, all of the compute units, the full copy engine bandwidth. Nothing is partitioned, nothing is queued, and no other tenant appears in your device list.
The practical consequence is that your benchmark this evening matches your benchmark last Tuesday. Fine-tuning runs that take two days finish in the time your first hour predicted. Interactive work stays interactive, because there is no other tenant’s batch job in front of yours.
All the VRAM, all the time
Forty-eight gigabytes on an L40S, twenty on an RTX 4000 Ada, and none of it reserved for a hypervisor partition. What the card has is what your process can address.
Idle costs you nothing extra
The card is yours by the month, not by the second. Leave a model resident for three weeks between experiments and nobody reclaims it while you are asleep.
Two cards, one host
The dual-card tier puts both devices on the same NUMA domain of the same host. They talk over PCIe rather than a card-to-card link, which is worth knowing before you plan a training topology around them.
Your own stack
We hand you a machine with a card in it and stay out of the way. Drivers, container runtime, framework versions and kernel are yours. There is no managed layer that has opinions about your Python.
What these cards suit
Both cards are workstation-class rather than the top of a training cluster, and the honest positioning follows from that.
Inference, fine-tuning, rendering and batch numerical work. In that order.
If your model fits in forty-eight gigabytes, an L40S will serve it competently for as long as you keep paying for it. If it does not fit, no amount of enthusiasm makes it fit, and the next honest step is quantisation, sharding across the dual-card tier, or a different class of machine entirely.
| Workload | Which card | The real constraint |
|---|---|---|
| Serving a quantised model to real users | RTX 4000 Ada | Memory, then request concurrency |
| Serving a mid-size model at full precision | L40S | Forty-eight gigabytes is the ceiling that decides it |
| Fine-tuning with adapters | L40S | VRAM during the backward pass, not raw throughput |
| Full fine-tune of a large model | Two L40S, or somewhere else | Card-to-card bandwidth over PCIe becomes the limit |
| Batch rendering and video encode | Either | Disk and network far more often than the card |
| Numerical simulation in double precision | Neither, usually | These are not double-precision parts. Take EPYC cores |
Who should not buy this
GPU capacity attracts optimistic sizing more than any other product we sell. Three groups of people should walk past this line.
The most useful thing we can tell some customers is that we are the wrong shop.
You are training something enormous from scratch
Multi-node training with high-bandwidth interconnect between hosts is not what this is. Two cards over PCIe on one host is our ceiling, and pretending otherwise would waste weeks of your time.
You want per-second billing
Cards are rented by the month. If your usage is genuinely twenty minutes a week, a metered platform will cost you less than we will, and we would rather say so now.
Your model does not fit
Ninety-six gigabytes across two cards is the maximum addressable here. Work out the memory arithmetic for weights, activations and optimiser state before ordering rather than after.
You need double precision
These are single-precision and mixed-precision parts. Scientific codes that insist on FP64 throughput will run faster on forty-eight EPYC cores than on either card.
You want us to manage the stack
We do not maintain your driver, your CUDA version or your framework matrix. The hardening add-on covers the operating system baseline and stops there.
The host around the card
A starved GPU is an expensive space heater. The host matters as much as the card, and it is where most cheap GPU offerings quietly cut the cost.
EPYC 9354P, ECC memory, NVMe local to the machine, forty-gigabit port.
Cores that can feed it
Eight to thirty-two dedicated EPYC cores depending on tier, pinned in the same NUMA domain as the card. Data loading and preprocessing are the usual reason a training loop stalls, and they are CPU work.
ECC memory on the host
Registered DDR5-4800, sixty-four to two hundred and fifty-six gigabytes. A forty-eight-hour job is exactly the kind of thing that should not be undone by one flipped bit in a dataloader buffer.
NVMe that keeps up
One to four terabytes of local Gen4 NVMe, mirrored. Datasets stream from the machine itself rather than across a network volume that someone else is also reading.
A forty-gigabit port
Pulling a multi-terabyte dataset in should take an evening. Fair use on a forty-gigabit port is two hundred and fifty terabytes a month, published on the pricing page.
Out-of-band console
Serial and VNC over an authenticated tunnel, included. When a driver install goes wrong at midnight you can still reach the machine, which is the entire reason it exists.
The card is passed through, so the driver is installed inside your instance like any other machine. Nothing on the host touches your framework, and a reboot does not renegotiate anything.
Where the cards are
Amsterdam, Frankfurt, London, New York, Dallas, Los Angeles, Singapore and Tokyo. Those are the floors with the power density and the cooling to run cards at full load continuously rather than in polite bursts.
Eight sites in six countries. GPUs need power and cooling, not postcodes.
Frankfurt is the busiest GPU floor we operate and gets new cards first. Dallas is the easiest place to get capacity at short notice, because power there is cheap and the floor is large. Los Angeles carries cards but runs low often, so order early if the westbound route into Asia is what you are buying.
Stock genuinely moves on this line. A card that shows as available at midday may not be there in the evening, and we would rather show you an accurate sold-out label than take an order we cannot fill this week.
| Site | Notes | Typical use |
|---|---|---|
| Amsterdam | Densest site in the fleet, everything ships here first | European inference with low latency to most of the continent |
| Frankfurt | Busiest GPU floor, first to receive new cards | Training and fine-tuning where capacity matters more than the last millisecond |
| London | Lowest transatlantic latency in Europe | Serving users on both sides of the Atlantic from one place |
| New York | East-coast anchor, full product range | North American inference |
| Dallas | Cheap power, large floor, easiest capacity | Long batch jobs and anything price-sensitive |
| Los Angeles | Best westbound routes we have outside Asia | Serving the Pacific rim from North America |
| Singapore | Default choice for South-East Asia | Regional inference |
| Tokyo | Most stable round-trip time in the region | Latency-sensitive Japanese and Korean traffic |
Long jobs, and how to not lose one
A forty-eight-hour fine-tune is a bet on nothing going wrong for two days. Structure the job so that losing the bet costs you an hour rather than the weekend.
Checkpoint. We will still say it after you have heard it a hundred times.
- 01
Checkpoint to local NVMe, often
Every few hundred steps, to the local disk, because it is fast enough that the cost is negligible. A job that cannot resume is a job you will eventually run twice.
- 02
Copy checkpoints off the box
Local NVMe is mirrored, not immortal. Off-node backup writes to a different city nightly, and it is the cheapest insurance on this page.
- 03
Snapshot before you touch the driver
Driver and framework upgrades are the leading cause of a working environment becoming a non-working one. A snapshot takes seconds to make and seconds to restore.
- 04
Watch temperature and clocks, not just loss
Cards throttle when the room is having a bad day. If your throughput drops without your code changing, look at the clock figures before you rewrite the dataloader.
- 05
Ask before you scale sideways
If the next step is four or eight cards, tell support what you are trying to do. Sometimes the answer is a bare-metal machine, and occasionally the answer is that we are not the right supplier.
When the hardware fails
GPUs usually degrade rather than die: falling clocks, corrected memory errors on the card, a job that suddenly runs a third slower for no reason you can find in your code.
Cards fail differently from disks. Slower, and with more warning.
Our monitoring watches card temperature, clock behaviour and error counters on every host. When a card starts misbehaving we contact you before it stops working, and we schedule the swap around your job rather than through the middle of it. Where the card fails outright, your instance is rebuilt on another host in the same site with its volumes and addresses intact.
Recovery on this line is slower than elsewhere in the fleet, and it is worth being plain about why: a terabyte of dataset and a resident model take real minutes to move and reload. Plan for an hour rather than fifteen minutes, and keep checkpoints somewhere other than the machine that is on fire.
Warning before failure, where possible
Error counters and clock telemetry are read continuously. Most card replacements happen in a window we agreed with the customer in advance.
Volumes and addresses survive
The rebuild carries your NVMe contents and IP addresses across. Nothing in your DNS or your certificates needs to change.
Credits are automatic
The same 99.99% contractual uptime applies here as everywhere else, and the credit lands without a claim form.
Questions about this line
The physical device is bound to your instance over a full sixteen-lane link. There is no vGPU partitioning, no time-slicing and no other tenant on it. Your device list shows one card because there is one card.
No. Cards are monthly, and the shortest commitment is one month. If your workload is genuinely bursty, a metered provider will be cheaper, and there is no version of this conversation where we pretend otherwise.
None, unless you ask. Images ship clean and the driver goes on inside your instance, because half of our GPU customers have strong views about versions and the other half have a container that carries its own.
No. They sit on the same host and the same NUMA domain, and they communicate over PCIe. For data-parallel work with modest gradient exchange this is fine. For anything that assumes a high-bandwidth card-to-card fabric, it is not, and you should size your expectations accordingly.
No. Passthrough requires the host to be built for it, with the PCIe topology and the power budget in place. Moving to this line means a new instance and a data copy.
The volume is wiped and the card goes back into the pool. Take your checkpoints off before the term ends, because a cancelled instance is genuinely gone rather than parked somewhere waiting for you to change your mind.
Take a whole card
Check the memory arithmetic first, pick the site with the capacity, and pay in coin. Root access arrives in under a minute, the driver decision stays yours, and the first seven days are refundable without a reason.