The NVMe Drive That Only Failed While Charging
Table of Contents
Three PCIe kernel parameters fixed nothing. The actual bug was in the firmware blob, and the fix was pinning a five-day-old package.
New laptop, new NixOS install, and within the first week the NVMe drive started throwing I/O errors at me — but only sometimes, and it took me longer than it should have to see the pattern: it correlated with the charger being plugged in, not with load, not with temperature, not with anything in the actual storage stack.
The first fix, and the second, and the third
An NVMe drive going sideways in tandem with power-adapter state points straight at PCIe power management — ASPM renegotiating link states, a device dropping into a low-power mode it can't cleanly wake from, an AER (Advanced Error Reporting) event the kernel overreacts to. All well-trodden territory, all with well-known kernel parameters, so that's where I went first:
nvme_core.default_ps_max_latency_us=0
pci=noaer
pcie_aspm=off
The first disables NVMe's own runtime power-state transitions. The second stops the kernel from treating PCIe Advanced Error Reporting events as fatal. The third turns off ASPM link-power-management entirely, at the PCIe subsystem level, for every device on the bus. Between the three I thought I'd covered every plausible power-management culprit — the drive itself, the kernel's error handling around it, and the bus underneath it.
I shipped all three at once, rebooted, and the errors kept happening.
Reading the actual firmware, not just the driver
Three kernel parameters not fixing the problem told me the problem wasn't in
kernel-side power-management policy at all — those parameters are basically the entire
surface area of "the kernel decided to do something clever with power state," and
clever kernel policy wasn't it. What was left was the firmware blob itself: the binary
linux-firmware ships for this specific NVMe controller, running on the controller,
doing whatever it does before the kernel ever gets a say.
linux-firmware isn't versioned the way a kernel driver is — it's a loose bundle of
vendor binary blobs, one per device family, updated on its own schedule with no
per-blob changelog most of the time. There's no git bisect against a black box like
that in any normal sense. The only lever I had was which snapshot of the bundle gets
loaded, and NixOS makes that a one-line answer: pin the package to an older nixpkgs
revision and see if the drive stops erroring.
# hosts/ZenS13/default.nix
nixpkgs.overlays = [
(final: prev: {
inherit (inputs.nixpkgs-old.legacyPackages.${prev.system}) linux-firmware;
})
];
nixpkgs-old pins to a specific nixpkgs commit — dc9637876d0dcc8c9e5e22986b857632effeb727,
carrying linux-firmware 20250708 — fetched as its own flake input rather than the
system-wide one, so exactly one package on this one host is held back while everything
else tracks current nixpkgs normally.
It worked. I pulled the three kernel parameters back out in the same change: they'd done nothing, and shipping dead configuration alongside a real fix is how a future cleanup of mine would mistake load-bearing code for cruft, or vice versa.
Why a five-day gap mattered
The pinned revision is five days older than whatever nixpkgs-unstable carried at
install time. Five days isn't a long time for a regression window, and it's a strong
signal that this wasn't a long-standing, widely-hit bug — more likely a specific
firmware bundle update landed, broke power-state handling for this exact controller on
this exact charger/battery combination, and nobody upstream had connected the two yet.
That fits with my kernel-parameter attempts failing: a firmware regression in how the
drive itself handles a power-state transition isn't something I can route around from
the host side by asking the kernel to be less aggressive about power management. The
firmware is still the thing initiating the bad transition; the kernel parameters just
change how loudly it complains afterward.
What's actually pinned here, and for how long
This isn't a permanent fork. nixpkgs-old is a single flake input the rest of the
system doesn't touch, and I spelled out the exact commit, exact version, and the reason
both in flake.nix and next to the overlay itself — the two verification commands I
left there (nix eval .#nixosConfigurations.ZenS13.pkgs.linux-firmware.version --raw
against nix eval nixpkgs#linux-firmware.version --raw) turn "is this still necessary"
into a single command instead of a question I'd have to re-derive every time I bump
nixpkgs.
Lessons
A correlation with power state doesn't mean the fix is in power-management policy. It was the right category of hypothesis — ASPM, AER, and NVMe power states are exactly what I should suspect when something "fails when the charger state changes" — and every lever in that category coming up empty was itself useful information: it pointed me at the one thing sitting below all of those levers instead of at another knob in the same category.
A binary blob can still be bisected — one build at a time, by pinning the bundle
instead of the individual file. There was no meaningful diff for me to read inside
linux-firmware, but there's a meaningful diff between "nixpkgs commit before" and
"nixpkgs commit after," and Nix let me hold one package to an arbitrary point in that
history as a small, explicit, revertible change instead of a hand-patched system file
I'd forget I'd installed.
Remove the workaround that didn't work, in the same commit as the fix that did. Three kernel parameters that turned out to be irrelevant would have kept looking like part of the fix forever if I'd shipped them separately — cruft that reads as load-bearing is worse than no comment at all, because it actively teaches the wrong lesson to whoever reads it next, including me in six months.