A Nix-Managed GPU VM Is Genuinely Great, and Here's the Proof
Table of Contents
Passing an RTX 3060 through to a Proxmox VM, lying to NVIDIA's driver about being virtualized, and making the whole thing reproducible from one command.
beta-cuda is a NixOS VM I run on a Proxmox host, with an RTX 3060 passed through via
PCIe, serving as the CI/build server for my homelab's Rust projects. Getting a consumer
NVIDIA GPU working inside a VM at all meant lying to the driver about where it's
running. Making that setup something I can destroy and recreate on demand, instead of a
pile of manual qm commands I'd never remember the order of, is the part that actually
makes the VM worth having.
The part where I lie to NVIDIA's driver
Consumer NVIDIA drivers have historically refused to initialize if they detect they're running inside a hypervisor — a market-segmentation check, not a technical limitation. Passing the physical GPU through via VFIO doesn't get around this by itself; QEMU still presents a CPU that announces itself as virtualized unless I tell it not to. Getting the driver to load meant clearing every signal QEMU normally leaves behind:
# infra/proxmox/vm-beta-cuda.tf
cpu {
type = "host" # expose real host CPU features
flags = ["+pcid"]
}
# applied out-of-band via SSH (Proxmox API tokens can't set these):
# -cpu 'host,+kvm_pv_unhalt,+kvm_pv_eoi,hv_vendor_id=NV43FIX,kvm=off'
# --cpu host,hidden=1,flags=+pcid
# --hostpci0 0000:01:00,pcie=1
kvm=off turns off the KVM CPUID leaf that would otherwise announce "you are running
under KVM." hidden=1 is the Proxmox-side flag for the same idea — hide the hypervisor
bit. hv_vendor_id=NV43FIX overwrites the Hyper-V vendor ID string QEMU exposes by
default with something that isn't a recognizable hypervisor signature. None of these are
NixOS-side settings; they're QEMU/Proxmox arguments the guest never sees as
configuration, which is the point. From inside the VM, nvidia-smi just works, the
same as it would on bare metal.
Terraform can't set the two options that matter
args (the raw KVM arguments above) and hostpci0 (the actual PCI passthrough
assignment) are both things the bpg/proxmox Terraform provider can express — and both
are things Proxmox's API rejected on me from any token-based auth, including a
root-owned token, with a hard check for authuser eq 'root@pam'. A token is never that
user, by design. So I let tofu apply create the VM shape (disks, network, memory, CPU
socket count) through the API as normal, and added a null_resource with a local-exec
provisioner to push the two token-hostile settings over SSH afterward, as a qm set
call gated on the same triggers Terraform tracks for everything else:
resource "null_resource" "beta_cuda_args" {
triggers = {
args = "-cpu 'host,+kvm_pv_unhalt,...,hv_vendor_id=NV43FIX,kvm=off'"
hostpci = "0000:01:00,pcie=1"
affinity = "0-7,16-23"
}
provisioner "local-exec" {
command = <<-EOT
ssh adlink-beta bash <<'REMOTE'
qm set ${self.triggers.vm_id} --args "${self.triggers.args}" --hostpci0 "${self.triggers.hostpci}" --affinity "${self.triggers.affinity}"
REMOTE
EOT
}
}
It's a workaround for a real API limitation, not something I chose for its own sake —
and it's still fully declarative. The values live in version control next to everything
else that defines this VM, re-applied automatically whenever a trigger changes, instead
of living only in whatever state I happened to leave on the hypervisor during some past
interactive qm set session.
CPU affinity: the bug that only showed up when two VMs competed
beta-cuda shares its physical host with beta-meerkat, a second VM I run a desktop
session on. Left alone, Proxmox floats both VMs' vCPUs across all 32 threads of the host
7950X. That's fine until both VMs are doing something CPU-heavy at once — a cargo build
on beta-cuda, a desktop session doing anything on beta-meerkat — and the two start
evicting each other from L3 cache, because the AMD 7950X's cache domains are split by
CCD (core complex die), and "any of 32 threads" doesn't respect that boundary at all.
I fixed it with affinity pinning, one CCD per VM: beta-cuda gets cores 0–7 plus their
SMT siblings 16–23 — exactly the physical cores sharing L3 domain 0, which I confirmed
against lscpu -e=CPU,CORE,CACHE rather than assumed from the core-count math — and
beta-meerkat gets the other CCD. Same qm-over-SSH mechanism as the GPU args, applied
via the same Terraform pattern, because Proxmox restricts CPU affinity to a direct
root@pam session too.
The thin-provisioning trap that hid behind a passing check
fstrim.timer runs weekly on every VM by default in NixOS — the guest-side half of
reclaiming space on a thin-provisioned virtual disk was never what I was missing. What I
was missing lived entirely on the Proxmox side: I'd attached the disk with discard = "ignore", which advertises discard support to the guest and then drops every UNMAP
request on the floor. fstrim ran every week, reported success every week, and none of
it did anything — the guest filesystem stayed at 44% used while the underlying thin LV
crept toward 99.85% allocated. A check that reports success isn't the same as a check
that verified the thing it claims to have checked; it verified that the command
succeeded, which told me nothing about whether the disk actually reclaimed space. The
fix was one word, on the hypervisor side of a boundary the guest config can't see
across: discard = "on" in the Terraform disk block.
Redeploying, and why that's the actual point
The reason I bother with any of the above instead of clicking through Proxmox's web UI
once and never touching it again: nix run .#deploy-vm -- beta-cuda builds my current
NixOS config and pushes it to the running VM, and the VM's shape — GPU passthrough,
CPU pinning, disk size, network — comes from the same tofu apply regardless of whether
this is the first time the VM has existed or the fifth time I've destroyed and rebuilt
it. Nothing about getting an RTX 3060 correctly recognized inside a VM again lives only
in my own memory of what I clicked. It's the difference between a GPU VM being a one-off
feat and being infrastructure.
Lessons
A workaround for a platform's API gap doesn't have to give up being declarative.
Proxmox's token-auth restriction on args/hostpci/cpu looked at first like a forced
retreat to manual configuration. Routing those specific fields through null_resource +
SSH, still triggered by the same state Terraform tracks for everything else, kept the
whole VM definition in one file instead of splitting it between version control and
hypervisor-side tribal knowledge.
A passing health check only proves the check ran, not that the thing it's checking
actually happened. fstrim.timer succeeding every week for a disk that never
reclaimed space is the same shape of trap as a program returning 0 having done nothing —
the exit code was never lying to me, it just wasn't checking what I assumed it was
checking.
Cache topology is invisible until two things compete for it. Neither VM was
individually slow or wrong; the interaction only showed up as unexplained variance under
simultaneous load, and I only found the fix by going looking for a hardware-topology
fact (lscpu -e) that no amount of staring at either VM's own config would have
surfaced.