PCI passthrough
Updated
A machine can take direct ownership of a host PCI device, whether it is a classic VM or an AppVM. Use it when a workload needs the real hardware: an accelerator, a storage controller, or a dedicated network card.
The unit is the IOMMU group, not the device
Section titled “The unit is the IOMMU group, not the device”Devices are assigned a whole IOMMU group at a time, because that is the boundary the hardware can actually isolate. Selecting one device brings its group with it: a graphics card normally travels with its own audio function, and the picker puts one checkbox on the group rather than one on each device. PCI bridges inside the group are the exception: you do not have to name them, and you cannot.
Once assigned, the host has given the group up. It stays assigned while the machine is stopped, and comes back to the host only when you remove the assignment or delete the machine.
Only a running machine holds a device
Section titled “Only a running machine holds a device”Virtainer Free tracks two separate facts about every device, and the difference between them is what lets you lend a card out without configuring anything twice.
| Fact | Who can hold it |
|---|---|
| Listed in a machine’s configuration | Any number of machines, running or stopped |
| Claimed for use | One machine that is running, or one an operation has stopped only so it can boot again |
Two stopped machines can both name the same device, and that is a valid configuration rather than a conflict waiting to be fixed. The decision is made at boot: the second machine to start is refused, and the message names the device and the machine using it. Stop that machine and this one can boot. Nothing has to be unassigned first, and a machine that stops keeps its configuration while losing its claim.
The second kind of holder is the one worth knowing about. Rolling an AppVM to a new image stops it and boots it again inside one operation, and that stop does not put the device up for grabs. Without this, every release of an AppVM with a device assigned would be a race for the device, lost at the last step with the system disk already swapped.
Assignments change while the machine is stopped
Section titled “Assignments change while the machine is stopped”Passthrough is cold configuration. Change it while the machine is created or stopped, and the new list takes effect at its next boot. A running machine refuses the change with a 400 that tells you to stop it first. A paused machine is refused too, with a 409, because pausing freezes a machine rather than ending it.
-
Stop the machine.
A running or paused machine cannot change its devices.
-
Edit the passthrough list.
The device picker is offered while the machine is created or stopped. The review lists the change as something the machine picks up when it next boots.
-
Boot the machine.
It comes up with the devices on the new list, or with none if you cleared it.
The stop is not a formality. A machine uses an assigned device through a handle it holds, and a handle cannot be taken back from outside. The one moment the host can be certain a machine is finished with a device is when the process using it has ended, so that is where the boundary sits. A hypervisor reporting a successful hot removal has not proved that, and the device may be about to be handed to a neighbour.
A machine that is created but never booted costs less, since it has never used the device: the hypervisor prepared behind the scenes is retired as part of the change, and the machine is left stopped with your new list.
A device keeps the slot it was first given, so adding or removing other devices does not move it. Names the guest already uses for a device it still owns stay valid across a change.
What you give up
Section titled “What you give up”| Constraint | Detail |
|---|---|
| No memory reclaim | A machine with a device assigned is given no balloon device, and setting a reclaim target on it is refused |
| No online memory growth | The device pins the guest’s memory, so a memory growth ceiling cannot be declared alongside it |
| vCPU growth still works | Only memory is pinned. A vCPU ceiling can still be declared, and vCPUs can still be added while the machine runs |
| No device state in a snapshot | Disks and configuration are captured; the device assignment is not. Restoring neither takes a device away nor hands one over. See Snapshots |
| An AppVM changes guest kernel | See the next section |
Passthrough on an AppVM
Section titled “Passthrough on an AppVM”Assigning devices to an AppVM also selects its guest kernel, because the drivers this path needs are out-of-tree modules and the default AppVM kernel cannot load modules at all. You are not choosing a kernel: the assignment chooses it, and an AppVM without devices keeps the kernel that cannot load code.
- There is no silent fallback. If the host does not carry the module-capable
kernel, creating the AppVM is refused with a 400 that names the missing file,
rather than producing a machine that lists a device nothing can drive. The
refusal is the remedy too: on a host whose
/usris read-only, place a copy of that kernel under/varand pointVIRTAINER_APPVM_PASSTHROUGH_KERNELat it, or create the machine without assigned devices. - The image has to ship a driver built for that kernel. Virtainer Free does not supply drivers.
- The alternate kernel adds module loading only, not drivers for other device classes. It has none for storage or USB controllers, so hand those to a classic VM, where the guest brings its own kernel.
- The console offers the device picker when you create the AppVM and not afterwards, so plan the choice as one made at creation.
Devices you cannot select
Section titled “Devices you cannot select”The host’s own boot display is deliberately not offered. On a typical machine that is the integrated GPU, and handing it over leaves the host with no console while the guest still gets a half-working device. The host also keeps the disk controller behind its own mounted filesystems and swap, any network interface it is currently using, and any device with no IOMMU group. The device list marks each of those unavailable and says why.
PCI bridges are the exception to that: they are not shown at all. A bridge neither passes through nor takes part in the group closure below, so the picker leaves it out rather than offering a row that could never be selected.
Availability is not a compatibility promise. Virtainer Free keeps no list of approved devices: the vfio binding and the guest’s own driver decide whether a device works. Two shapes are worth planning around:
- Some discrete GPUs depend on video BIOS that lives on the host rather than on the card, most commonly in laptops with switchable graphics. Those enumerate in the guest but the driver fails to initialise them. Desktop and server cards with their own ROM are not affected.
- An integrated GPU that is not acting as the host’s boot display is offered, and the result is poor on both sides: the guest’s driver gets no description of the display outputs, while the host keeps faulting on the memory region the graphics device still uses for itself.
A USB device cannot be handed over on its own. What moves is the whole controller, so every other device plugged into it goes to the guest too.
Do not treat GPU passthrough as something that generally works. On an AppVM the module-capable kernel is paired with NVIDIA’s current out-of-tree driver line, which covers Turing and newer GPUs. Older architectures have no verified driver for that kernel, and no release has yet carried a verified claim that the CUDA path works on supported hardware: treat it as unreleased rather than unproven. On a classic VM the driver comes from your image, so the guest’s own kernel and driver version decide. Test on your own hardware before a workload depends on it.
One machine can be given up to 16 devices, counted by address, so a group of five functions uses five of them. They have to be named as whole IOMMU groups: leave one function out and the request is refused, with the missing device named.