A Nix-Managed GPU VM Is Genuinely Great, and Here's the Proof

beta-cuda is a NixOS VM I run on a Proxmox host, with an RTX 3060 passed through via PCIe, serving as the CI/build server for my homelab's Rust projects. Getting a consumer NVIDIA GPU working inside a VM at all meant lying to the driver about where it's running. Making that setup something I can destroy and recreate on demand, instead of a pile of manual qm commands I'd never remember the order of, is the part that actually makes the VM worth having.

The part where I lie to NVIDIA's driver

Consumer NVIDIA drivers have historically refused to initialize if they detect they're running inside a hypervisor — a market-segmentation check, not a technical limitation. Passing the physical GPU through via VFIO doesn't get around this by itself; QEMU still presents a CPU that announces itself as virtualized unless I tell it not to. Getting the driver to load meant clearing every signal QEMU normally leaves behind:

# infra/proxmox/vm-beta-cuda.tf
cpu {
  type = "host"       # expose real host CPU features
  flags = ["+pcid"]
}
# applied out-of-band via SSH (Proxmox API tokens can't set these):
#   -cpu 'host,+kvm_pv_unhalt,+kvm_pv_eoi,hv_vendor_id=NV43FIX,kvm=off'
#   --cpu host,hidden=1,flags=+pcid
#   --hostpci0 0000:01:00,pcie=1

kvm=off turns off the KVM CPUID leaf that would otherwise announce "you are running under KVM." hidden=1 is the Proxmox-side flag for the same idea — hide the hypervisor bit. hv_vendor_id=NV43FIX overwrites the Hyper-V vendor ID string QEMU exposes by default with something that isn't a recognizable hypervisor signature. None of these are NixOS-side settings; they're QEMU/Proxmox arguments the guest never sees as configuration, which is the point. From inside the VM, nvidia-smi just works, the same as it would on bare metal.

Terraform can't set the two options that matter

args (the raw KVM arguments above) and hostpci0 (the actual PCI passthrough assignment) are both things the bpg/proxmox Terraform provider can express — and both are things Proxmox's API rejected on me from any token-based auth, including a root-owned token, with a hard check for authuser eq 'root@pam'. A token is never that user, by design. So I let tofu apply create the VM shape (disks, network, memory, CPU socket count) through the API as normal, and added a null_resource with a local-exec provisioner to push the two token-hostile settings over SSH afterward, as a qm set call gated on the same triggers Terraform tracks for everything else:

resource "null_resource" "beta_cuda_args" {
  triggers = {
    args    = "-cpu 'host,+kvm_pv_unhalt,...,hv_vendor_id=NV43FIX,kvm=off'"
    hostpci = "0000:01:00,pcie=1"
    affinity = "0-7,16-23"
  }
  provisioner "local-exec" {
    command = <<-EOT
      ssh adlink-beta bash <<'REMOTE'
      qm set ${self.triggers.vm_id} --args "${self.triggers.args}" --hostpci0 "${self.triggers.hostpci}" --affinity "${self.triggers.affinity}"
      REMOTE
    EOT
  }
}

It's a workaround for a real API limitation, not something I chose for its own sake — and it's still fully declarative. The values live in version control next to everything else that defines this VM, re-applied automatically whenever a trigger changes, instead of living only in whatever state I happened to leave on the hypervisor during some past interactive qm set session.

CPU affinity: the bug that only showed up when two VMs competed

beta-cuda shares its physical host with beta-meerkat, a second VM I run a desktop session on. Left alone, Proxmox floats both VMs' vCPUs across all 32 threads of the host 7950X. That's fine until both VMs are doing something CPU-heavy at once — a cargo build on beta-cuda, a desktop session doing anything on beta-meerkat — and the two start evicting each other from L3 cache, because the AMD 7950X's cache domains are split by CCD (core complex die), and "any of 32 threads" doesn't respect that boundary at all.

I fixed it with affinity pinning, one CCD per VM: beta-cuda gets cores 0–7 plus their SMT siblings 16–23 — exactly the physical cores sharing L3 domain 0, which I confirmed against lscpu -e=CPU,CORE,CACHE rather than assumed from the core-count math — and beta-meerkat gets the other CCD. Same qm-over-SSH mechanism as the GPU args, applied via the same Terraform pattern, because Proxmox restricts CPU affinity to a direct root@pam session too.

The thin-provisioning trap that hid behind a passing check

fstrim.timer runs weekly on every VM by default in NixOS — the guest-side half of reclaiming space on a thin-provisioned virtual disk was never what I was missing. What I was missing lived entirely on the Proxmox side: I'd attached the disk with discard = "ignore", which advertises discard support to the guest and then drops every UNMAP request on the floor. fstrim ran every week, reported success every week, and none of it did anything — the guest filesystem stayed at 44% used while the underlying thin LV crept toward 99.85% allocated. A check that reports success isn't the same as a check that verified the thing it claims to have checked; it verified that the command succeeded, which told me nothing about whether the disk actually reclaimed space. The fix was one word, on the hypervisor side of a boundary the guest config can't see across: discard = "on" in the Terraform disk block.

Redeploying, and why that's the actual point

The reason I bother with any of the above instead of clicking through Proxmox's web UI once and never touching it again: nix run .#deploy-vm -- beta-cuda builds my current NixOS config and pushes it to the running VM, and the VM's shape — GPU passthrough, CPU pinning, disk size, network — comes from the same tofu apply regardless of whether this is the first time the VM has existed or the fifth time I've destroyed and rebuilt it. Nothing about getting an RTX 3060 correctly recognized inside a VM again lives only in my own memory of what I clicked. It's the difference between a GPU VM being a one-off feat and being infrastructure.

Lessons

A workaround for a platform's API gap doesn't have to give up being declarative. Proxmox's token-auth restriction on args/hostpci/cpu looked at first like a forced retreat to manual configuration. Routing those specific fields through null_resource + SSH, still triggered by the same state Terraform tracks for everything else, kept the whole VM definition in one file instead of splitting it between version control and hypervisor-side tribal knowledge.

A passing health check only proves the check ran, not that the thing it's checking actually happened. fstrim.timer succeeding every week for a disk that never reclaimed space is the same shape of trap as a program returning 0 having done nothing — the exit code was never lying to me, it just wasn't checking what I assumed it was checking.

Cache topology is invisible until two things compete for it. Neither VM was individually slow or wrong; the interaction only showed up as unexplained variance under simultaneous load, and I only found the fix by going looking for a hardware-topology fact (lscpu -e) that no amount of staring at either VM's own config would have surfaced.