Teaching a Nix-Packaged Photo App About a GPU It Didn't Know It Had
Table of Contents
nixpkgs' Immich package has no CUDA wiring. Building it from source cost hours and an OOM kill. Fetching the vendor's own wheel cost minutes.
Immich is the self-hosted Google Photos alternative I run — upload photos, get automatic
face grouping and text/object search back, powered by a small machine-learning worker
running CLIP embeddings, face detection, and OCR locally. All of that runs meaningfully
faster on a GPU, and I've got one sitting idle on NitroAN515 most of the time (a GTX
1650 Ti). Getting Immich's ML worker to actually use it took me two different
approaches, and the first one wasted an afternoon.
What "no CUDA wiring by default" actually means
nixpkgs packages immich-machine-learning against a plain CPU build of onnxruntime —
the inference engine underneath the ML worker. There's no NixOS module option that
flips this to GPU; the CUDA-enabled build simply isn't what the package produces. My
first move, the obvious one in a Nix codebase, was an override:
onnxruntime = prev.onnxruntime.override { cudaSupport = true; };
This is the "correct" Nix way to ask for a different build configuration of a package,
and it works, in the sense that it eventually produces a working CUDA-enabled
onnxruntime. It also meant compiling onnxruntime from source, which for its CUDA
execution provider means compiling FlashAttention kernels for nine separate GPU
architectures, one at a time, on a machine with 15 GB of RAM. It took several hours. It
OOM-killed once, even at reduced build parallelism, and I had to restart it. And because
the override changes the derivation's hash, it rebuilds from scratch again on nearly
every nixpkgs bump — there's no cached binary anywhere for a locally-overridden
derivation this specific.
That's not a bug in the override; it's just what "build this yourself" costs for a package this size. The question I should have asked myself sooner wasn't "how do I make the build faster," it was "do I need to build this at all."
Using the binary NVIDIA and Microsoft already ship
onnxruntime-gpu is published as a prebuilt wheel on PyPI by Microsoft, targeting the
same CUDA/cuDNN ABI nixpkgs' own CUDA packages provide. Nothing about that wheel needs
compiling locally — it needs relinking, so its bundled shared libraries resolve
against nixpkgs' CUDA libs instead of whatever paths it was built expecting:
onnxruntime = pyprev.buildPythonPackage {
pname = "onnxruntime-gpu";
version = "1.24.4";
format = "wheel";
src = prev.fetchurl {
url = "https://files.pythonhosted.org/packages/.../onnxruntime_gpu-1.24.4-cp314-cp314-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl";
hash = "sha256-...";
};
nativeBuildInputs = [ prev.autoPatchelfHook ];
autoPatchelfIgnoreMissingDeps = [ "libnvinfer.so.10" "libnvonnxparser.so.10" ];
buildInputs = with prev.cudaPackages; [
prev.stdenv.cc.cc.lib cuda_cudart cuda_nvrtc libcublas libcurand libcufft libcusparse cudnn
];
};
autoPatchelfHook rewrites the wheel's binary dependencies to point at the Nix store
paths for those exact libraries instead of whatever the wheel's build environment had. I
ignored two deps on purpose — libnvinfer and libnvonnxparser are TensorRT
execution-provider libraries the wheel bundles but Immich never exercises, and pulling
in the entire TensorRT SDK to satisfy a code path I never hit would have defeated the
point of avoiding a heavy build in the first place. Minutes instead of hours, and stable
across nixpkgs bumps: I pinned this derivation to an explicit wheel URL and hash instead
of deriving it from whatever prev.onnxruntime happens to be on a given day, so it
doesn't rebuild just because nixpkgs moved.
The trade is real, not free: the wheel is ABI-locked to a specific Python minor version
(cp314 here), so bumping the pinned nixpkgs revision to one with a different default
Python means I have to go fetch the matching wheel by hand. A build-from-source override
tracks nixpkgs automatically; a pinned wheel needs me to notice and re-pin it myself.
For a dependency this expensive to rebuild, I'll take that trade.
The metadata mismatch that looked like a real bug
insightface (a face-detection library Immich's ML worker depends on) declares
Requires-Dist: onnxruntime in its own wheel's metadata. Nix's runtime dependency
checker cross-references that declared requirement against installed package names.
The onnxruntime-gpu wheel's own embedded metadata says onnxruntime_gpu, not
onnxruntime, because that's literally what its publisher named it. Renaming my Nix
derivation's pname doesn't change bytes already baked into a fetched .whl. The check
failed on a name it couldn't match, even though import onnxruntime — the thing that
actually matters — worked fine, which I verified directly with pythonImportsCheck. I
turned the check off for exactly this one package (dontCheckRuntimeDeps = true on the
insightface override), and wrote the reasoning for why it's safe right next to the
flag, since a bare "disable this check" reads as suspicious without it.
Lessons
"The Nix-idiomatic way" and "the fast way" aren't always the same package.
Overriding cudaSupport is the textbook move and it's not wrong — it's just expensive
for a dependency this large, and expensive in a way that recurs on every nixpkgs bump
rather than paying once. A prebuilt artifact from the same vendor, relinked rather than
recompiled, was available to me the whole time.
A dependency checker failing doesn't always mean the dependency is actually broken. The insightface/onnxruntime mismatch was a metadata-naming collision, not a real missing runtime dependency, and I only knew that because I tested the actual import instead of trusting the static check. Knowing which one I'm looking at before reaching for an override is the difference between a one-line justified skip and papering over a real bug.
Pin to an artifact, not to a moving target, when the point is to stop rebuilding. I
fetch the wheel by an explicit URL and hash, not derived from whatever prev.onnxruntime
happens to resolve to on a given day — that's what actually broke the "rebuilds on every
nixpkgs bump" cycle for me, not just avoiding the source build once.