The NVMe Drive That Only Failed While Charging

New laptop, new NixOS install, and within the first week the NVMe drive started throwing I/O errors at me — but only sometimes, and it took me longer than it should have to see the pattern: it correlated with the charger being plugged in, not with load, not with temperature, not with anything in the actual storage stack.

The first fix, and the second, and the third

An NVMe drive going sideways in tandem with power-adapter state points straight at PCIe power management — ASPM renegotiating link states, a device dropping into a low-power mode it can't cleanly wake from, an AER (Advanced Error Reporting) event the kernel overreacts to. All well-trodden territory, all with well-known kernel parameters, so that's where I went first:

nvme_core.default_ps_max_latency_us=0
pci=noaer
pcie_aspm=off

The first disables NVMe's own runtime power-state transitions. The second stops the kernel from treating PCIe Advanced Error Reporting events as fatal. The third turns off ASPM link-power-management entirely, at the PCIe subsystem level, for every device on the bus. Between the three I thought I'd covered every plausible power-management culprit — the drive itself, the kernel's error handling around it, and the bus underneath it.

I shipped all three at once, rebooted, and the errors kept happening.

Reading the actual firmware, not just the driver

Three kernel parameters not fixing the problem told me the problem wasn't in kernel-side power-management policy at all — those parameters are basically the entire surface area of "the kernel decided to do something clever with power state," and clever kernel policy wasn't it. What was left was the firmware blob itself: the binary linux-firmware ships for this specific NVMe controller, running on the controller, doing whatever it does before the kernel ever gets a say.

linux-firmware isn't versioned the way a kernel driver is — it's a loose bundle of vendor binary blobs, one per device family, updated on its own schedule with no per-blob changelog most of the time. There's no git bisect against a black box like that in any normal sense. The only lever I had was which snapshot of the bundle gets loaded, and NixOS makes that a one-line answer: pin the package to an older nixpkgs revision and see if the drive stops erroring.

# hosts/ZenS13/default.nix
nixpkgs.overlays = [
  (final: prev: {
    inherit (inputs.nixpkgs-old.legacyPackages.${prev.system}) linux-firmware;
  })
];

nixpkgs-old pins to a specific nixpkgs commit — dc9637876d0dcc8c9e5e22986b857632effeb727, carrying linux-firmware 20250708 — fetched as its own flake input rather than the system-wide one, so exactly one package on this one host is held back while everything else tracks current nixpkgs normally.

It worked. I pulled the three kernel parameters back out in the same change: they'd done nothing, and shipping dead configuration alongside a real fix is how a future cleanup of mine would mistake load-bearing code for cruft, or vice versa.

Why a five-day gap mattered

The pinned revision is five days older than whatever nixpkgs-unstable carried at install time. Five days isn't a long time for a regression window, and it's a strong signal that this wasn't a long-standing, widely-hit bug — more likely a specific firmware bundle update landed, broke power-state handling for this exact controller on this exact charger/battery combination, and nobody upstream had connected the two yet. That fits with my kernel-parameter attempts failing: a firmware regression in how the drive itself handles a power-state transition isn't something I can route around from the host side by asking the kernel to be less aggressive about power management. The firmware is still the thing initiating the bad transition; the kernel parameters just change how loudly it complains afterward.

What's actually pinned here, and for how long

This isn't a permanent fork. nixpkgs-old is a single flake input the rest of the system doesn't touch, and I spelled out the exact commit, exact version, and the reason both in flake.nix and next to the overlay itself — the two verification commands I left there (nix eval .#nixosConfigurations.ZenS13.pkgs.linux-firmware.version --raw against nix eval nixpkgs#linux-firmware.version --raw) turn "is this still necessary" into a single command instead of a question I'd have to re-derive every time I bump nixpkgs.

Lessons

A correlation with power state doesn't mean the fix is in power-management policy. It was the right category of hypothesis — ASPM, AER, and NVMe power states are exactly what I should suspect when something "fails when the charger state changes" — and every lever in that category coming up empty was itself useful information: it pointed me at the one thing sitting below all of those levers instead of at another knob in the same category.

A binary blob can still be bisected — one build at a time, by pinning the bundle instead of the individual file. There was no meaningful diff for me to read inside linux-firmware, but there's a meaningful diff between "nixpkgs commit before" and "nixpkgs commit after," and Nix let me hold one package to an arbitrary point in that history as a small, explicit, revertible change instead of a hand-patched system file I'd forget I'd installed.

Remove the workaround that didn't work, in the same commit as the fix that did. Three kernel parameters that turned out to be irrelevant would have kept looking like part of the fix forever if I'd shipped them separately — cruft that reads as load-bearing is worse than no comment at all, because it actively teaches the wrong lesson to whoever reads it next, including me in six months.