RUEN
To articles
Free Migration

Renting out your GPU on Vast.ai: the pre-launch checklist

Vast Sale editorial, 8 min

A quick checklist worth going through before you migrate. It saves time and lowers the risk of downtime in your first days on Vast.ai. We run six of our own machines with ten cards and maintain client rigs, and almost every item below got here not from the marketplace docs but from us tripping over it. The order matters too: the first three items are checked before you create a host account, the last two on switchover day. The numbers for each of our machines are open on the "Fleet" page of vast.sale: our live fleet page

The short version is five items. Each has its own section below with concrete thresholds, commands and our own measurements.

  • Up-to-date NVIDIA drivers and CUDA version — and pinned, so a nightly update does not take them away
  • Stable internet channel and open ports — by measurement, not by the number in your ISP contract
  • Adequate cooling and power — a dedicated 16 A line, a UPS and auto-restart after a failure
  • Backup access to the server — a second way in, tested before you need it
  • A planned transition window from Clore — do not switch over in a single day
Vast.ai official minimums for hardware, network and OS — every item below is checked against this
Vast.ai official minimums for hardware, network and OS — every item below is checked against this

1. Drivers and CUDA: pinning matters more than the version

Our OS is always Ubuntu 22.04 Server with the HWE 6.8 kernel: newer releases we consider raw for a host. The driver must show up in nvidia-smi, all cards in nvidia-smi -L, and the count there must match the physical one. But checking the version is not enough. The classic mistake is leaving everything upgradable: a routine apt upgrade swaps the kernel, the NVIDIA module fails to build against it, and the machine goes offline at night. And offline costs you not an hour of revenue but the reliability score the marketplace uses to decide whether to show you to clients at all.

So the kernel and the whole NVIDIA stack on our machines sit on apt-mark hold — 49 packages on the client rig, and satellite packages have to be listed by explicit name, because a glob pattern does not work there. Fifty-series cards add their own rules: the -open driver variant, plus nvidia_drm.modeset=0 pcie_aspm=off in GRUB. Installing the system from scratch up to this state: installing Ubuntu for Vast.ai

2. Bandwidth: calculate from VRAM, confirm by measurement

The marketplace requirement scales with the machine's total VRAM: roughly 2.6 Mbit/s per 1 GiB, but never below 100 and never above 500 Mbit/s, separately for download and upload. A single RTX 5090 (32 GB) works out to 83 Mbit by the formula — the floor kicks in, so such a node qualifies at 100/100. A 4×RTX 4090 machine is 96 GB, i.e. 250 Mbit. The 500 ceiling is only reached from roughly 192 GB. Symmetric 500/500 is headroom, not a mandatory minimum.

What you compare is not tariffs or NIC vendors but your own measurements. Across our fleet the rig on Realtek 2.5GbE delivers 706/625 Mbit while the rig on an Intel I225-V does 524/518: the intuition that "Intel beats Realtek" does not hold here. You can get a first read on your own link to the marketplace at the network check

A non-obvious point about measurements: the catalogue shows not your latest result but a moving average — a new measurement enters it with a weight of about 0.1, and the test itself fires from cron by lottery roughly once every nine hours. After a move or an ISP change there is no point in waiting: run the test manually, as a series. We ran the client machine ten times in a row, and only then did the catalogue figure reach 578/595 Mbit — before that it had shown the speed from the old address for months.

On ports: the host gets an allocated range that must be forwarded in full (a thousand ports per machine in our case). It is worth looking the other way too: an audit on one server found iperf3 listening to the world on port 5201 — installed for bandwidth tests and never closed. A look at connectivity under Russian conditions, including throttling by the VPS provider's autonomous system: what network a host needs

3. Power and cooling: and what happens after the lights go out

A 4×RTX 4090 machine draws 2.2–2.4 kW at peak and a steady 1.8–2.0 kW under load. A single 16 A line is 3.7 kW, i.e. exactly one such server: a second one needs a second line with its own breaker and 2.5 mm² copper or thicker. We keep card temperatures below 80–83 °C and check them on an hour-long run, not at idle. The UPS has to cover the router too, otherwise after a brief sag the machine comes up into a network that is not there. BIOS auto-restart after power loss is mandatory: otherwise one tripped breaker turns into a day offline. The power maths: how much power a server draws

Our own case from this item. On 19 August 2026 one machine lost power, and the CMOS reset along with it. The server came back up on its own, cards visible, ordinary rentals running — but clients who need virtual-machine mode would not start: the reset had turned hardware virtualization off. It is caught by a single command — ls /sys/kernel/iommu_groups: an empty directory means IOMMU is off. So after any power cut we check not "did it boot" but the BIOS settings against a list.

4. Backup access, and keys instead of a password

Since 26 August 2026 the marketplace flags password login as a configuration error every hour — formally a soft loss of trust. Moving the fleet to keys is a valid item, but the order inside it is strict: first confirm that key login works from a second machine, only then disable the password. On the client rig a small detail surfaced afterwards: the familiar PuTTY password login no longer gets in, you need a separate .ppk. On a calm day that is five minutes; mid-incident it is an hour lost.

You need a second way in precisely because the first one will fail some day. In our case it is a reverse tunnel through a foreign VPS: the same thing that gives access without a public IP and survives the home ISP changing your address. How it works and why it is not a hole: server access without a public IP

5. The transition window from Clore

The main rule: do not shut the old marketplace down on the same day you switch the new one on. Between "the machine is registered" and "the machine is verified and starting to appear in search" there is a gap, and in that window your income is non-zero only if you still have the old load. Note one more thing: changing a machine's IP address costs about 0.10 of the reliability score, and only on machines that have a rental running at that moment — so better not to combine a relocation with a switchover. A seven-point comparison of the marketplaces — our Clore vs Vast.ai comparison; the full step-by-step migration — the full migration guide

How to check each item

What to checkCommand or toolWhat we treat as a pass
Cards and drivernvidia-smi -Lall cards listed, count matches the physical one
Package pinningapt-mark showholdkernel and the whole NVIDIA stack (49 packages for us)
Virtualizationls /sys/kernel/iommu_groupsa non-empty list; empty = VT-d/IOMMU is off
Link bandwidthvast.sale/speedcheck, then your own test as a series≥ 2.6 Mbit per 1 GiB of VRAM, never below 100/100
Temperaturesnvidia-smi -q -d TEMPERATURE under load< 80–83 °C on an hour-long run, not at idle
Power linebreaker panel: rating and cable gaugea dedicated 16 A per server, 2.5 mm² copper or better
Recovery after a failurecut power at the main breakerthe machine returns online and into the catalogue by itself
System memoryswapping modules between slots, not memtest alonean hour under load with no dmesg errors
Backup accesskey login from a second machineit works before password login is disabled
Economicsvast.sale/calculatorthe numbers work at 50 % paid hours, not at 90 %

What broke on our fleet in 2026

This list is not meant to scare anyone but to calibrate: this is roughly the kind of work that lands on a host, and roughly how often. Not one of these is about a GPU dying.

  • A dead memory module on the four-4090 machine: memtest stayed silent, we proved it by swapping modules between slots and by swapping the module out
  • A stuck marketplace controller: the machine shows as inactive while everything works. On 25 August four rigs went that way; on 5 September all six in one day. Cured by resetting /var/lib/vastai_kaalia/controller_connection; the write-up — the stuck controller write-up
  • The marketplace's hourly cron wipes the VM flag when all cards are passed through: one machine sat without VM for ten days, from 19:22 on 11 September to 17:04 on 21 September
  • A tunnel loop at boot, before the DHCP address arrives: 28 thousand sockets filled the router's connection table, and on 20 September four rigs failed over to the backup VPS for no reason
  • Download throttling by the VPS provider's autonomous system: with two large providers the link was strangled, with a niche one it was clean. Tunnel settings do not help here, only moving does
  • The marketplace agent hanging on a blocking HTTP call with no timeout: the machine is alive, the last log line is about a call to the web API, and nothing goes out
  • The marketplace does not bill traffic on VM rentals: with $3 per terabyte set, a client moved 8.32 TB in 37 hours and the report showed $0.00

Where we got it wrong

Four mistakes that cost us income and that you will not find in other people's checklists, because those were not written from their own machines.

  • We trusted memtest. On a dead module it passed clean, and for months we looked for the cause in drivers and power. Our rule now: memtest can confirm a problem but cannot rule one out.
  • We misdiagnosed a card. The 4090 machine ran for more than a month without VM mode because we blamed the failure on an "incurable RTX 4090 reset bug". That card has no such bug — it belongs to Blackwell cards; the real cause was the card dropping into deep sleep when detached from the guest and never coming back, which device rules fix. All that time we simply were not taking VM rentals. The full procedure: VM mode on a multi-GPU host
  • We made failover to the backup VPS too twitchy. The threshold was three failed checks, and on 20 September ordinary noise on the home link moved four rigs onto the backup. Worse: the "is the second VPS alive" check was a fiction — the tunnel answered locally in 1–8 ms and we took that for the server's reply. The threshold is now nine, and the check is a real one, answering in 36–46 ms.
  • We started work on a client machine without measuring its link first. Afterwards we had to drag the catalogue's moving average back into place with a series of ten test runs. Measuring before the work starts is now the first item, not the last.
If any point raises questions, start with a free consultation — it is cheaper than losing days to downtime. Name your cards, your electricity tariff and your link, and we will tell you which of the pitfalls above apply to you: usually it is one or two items you can check yourself in an evening. Request form — our services

One last thing, about order. Running the numbers makes sense before you buy: an estimate for your cards and your tariff is at the income calculator — and use 50 % paid hours in it, not 90 %. Check your bandwidth at the network check. To compare the estimate with real money, see the "Fleet" page — our live fleet page: income for each of our machines refreshes once an hour, downtime and the failures listed above included. A month broken down card by card — our income breakdown card by card

Need help with the move?

Turnkey setup with a launch guarantee.

View services
© Vast Sale · Source: https://vast.sale/articles/a4
Copying, paraphrasing and AI processing are permitted only with an active link to https://vast.sale