Six Days Inside a GPU That Would Not Wake Up
Table of Contents
- The thing you have to understand first
- The message that ate five patches
- A GPU that works, sitting right next to one that doesn't
- Ten thousand words of arguments, all correct
- The sentence that had been sitting in public for a year
- Test the version they named
- Asking a better question
- The ten-line patch
- What I can prove, and what I can't
- Where it lives now
- Six things I'd tell myself on day one
An RTX 5070 Laptop that no NVIDIA driver could initialise, and what it took to prove the bug was in firmware rather than in code anyone can read.
I just got a laptop with an RTX 5070.
nvidia-smi had exactly one thing to say, and it said it instantly:
No devices were found
Everything else looked healthy. All four kernel modules loaded. lspci cheerfully
reported Kernel driver in use: nvidia. No BAR errors, no IOMMU complaints, no
firmware load failures. The machine was quite certain nothing was wrong.
This is the story of the six days it took to find out why, which ends somewhere I did not expect: in a binary I am not allowed to read, reached by a route that should not have worked.
The thing you have to understand first
I lost most of the first day to a wrong mental model, so let me hand you the right one up front.
When you install "the NVIDIA driver" on a modern card, you are not installing the
driver. Since Turing, NVIDIA moved the bulk of the Resource Manager — the layer
that owns initialisation, memory, channels, clocks, display — off the CPU and
onto a RISC-V core inside the GPU called the GSP, the GPU System Processor.
nvidia.ko is largely a client that marshals RPCs to it.
On Blackwell there is no CPU-side fallback at all. NVreg_EnableGpuFirmware=0
is accepted and silently ignored, because the CPU-side Resource Manager those
chips would need was never written.
So "the driver fails to initialise the GPU" is, on this hardware, a misleading sentence. What actually happens is: firmware on the GPU declines to initialise itself, and tells the host why over an RPC.
That reframing is most of the battle. It also means that when things go wrong, a large part of the system you need to inspect is a signed binary running on a processor you cannot attach a debugger to.
The message that ate five patches
The kernel log had a line in it that looked like the answer:
NVRM: GPU0 _kgspBootGspRm: unexpected WPR2 already up, cannot proceed with booting GSP
NVRM: _kgspBootGspRm: (the GPU is likely in a bad state and may need to be reset)
WPR2 is Write-Protected Region 2, a locked area of video memory where GSP-RM's image and heap live. The message says: someone already set this up, I cannot set it up again, the card needs a reset.
That reads like a complete diagnosis. It is the line that appears in nearly every report of this failure I later found — on Arch, on Fedora, on Ubuntu, in NVIDIA's own tracker. It sent me down a long tunnel: FLR, secondary bus reset, remove-and-rescan, D3cold, ACPI power resources, vendor WMI power cuts. I wrote five kernel patches trying to tear WPR2 down before the driver touched it.
None of it worked, and eventually I did the thing I should have done on day one.
I stopped reading dmesg | tail and read the log from the top.
The first attempt gets all the way through. FSP accepts the chain of trust in
0.78 seconds. WPR2 is built. GSP-FMC runs. GSP-RM boots, runs, and then returns
NV_ERR_INVALID_ARGUMENT from its own initialisation.
Then the driver retries. The retry finds WPR2 up — because attempt one built it — and refuses. Four sixty-second timeouts later, that refusal is the last thing in the log, so it is the thing everyone quotes.
The message everybody chases is an artifact of a failure that already happened. Five patches at the wrong layer, because I read the end of a log instead of the beginning.
A GPU that works, sitting right next to one that doesn't
The thing that broke the investigation open was not a patch. It was a control.
nova-core is the in-tree Rust NVIDIA driver in upstream Linux — a completely
independent implementation of the same boot sequence, written by different people
against the same silicon. On kernel 7.2-rc2 I removed the NVIDIA stack entirely
and booted it.
nova-core 0000:01:00.0: NVIDIA (Chipset: GB206, Architecture: BlackwellGB20x, Revision: a.1)
nova-core 0000:01:00.0: GPU name: NVIDIA GeForce RTX 5070 Laptop GPU
[drm] Initialized nova-drm 0.0.0 for nova-core.nova-drm.0 on minor 0
That second line is not cosmetic. It is printed at gsp/boot.rs:159, after
wait_gsp_init_done() at line 152, and after the GetGspStaticInfo RPC at 155
whose reply carries that string. For nova to print it, GSP-RM must have booted,
reached init-done, and answered a structured RPC — on this exact card, in the
state UEFI leaves it, with no reset and no preparation.
In one measurement, every hardware and platform explanation died. The silicon is fine. The FSP path is fine. WPR2 is fine. The card is not broken.
nova gives no CUDA and never will — it is a display driver. But as a control it was worth more than any patch I wrote.
Ten thousand words of arguments, all correct
If the hardware works and another driver can drive it, then NVIDIA's driver must be telling GSP-RM something it doesn't like. So I built a harness: a patched driver with a regkey that selects which single difference to apply on each boot, and the GSP-RM error trace as the observable.
Then I went through everything the driver hands the firmware, one boot at a time.
Twenty-two experiments. The results were maddening in a specific way: nothing was
wrong. fbSize: no change. The nine layout offsets: not actually a difference,
ACR fills them in. bIsPrimary: already zero, the value I was about to "fix" it
to. Drop the libos regions that nova omits: the firmware panics, so they are
required, not spurious. Zero the message-queue geometry: the queue never comes
up.
Then the big one. I zeroed all thirty-nine fields NVIDIA sends that nova does not, reducing the struct to exactly what a working driver sends.
Byte-identical failure.
I remember sitting there genuinely stuck. Every input was correct. The firmware was rejecting an initialisation for which I had individually verified every single argument.
The sentence that had been sitting in public for a year
Out of ideas on the technical side, I went reading. Arch forums. Ubuntu Discourse. Launchpad. Fedora. Reddit. NVIDIA's own GitHub tracker.
And in issue #876 — a thread about somebody else's laptop, closed as resolved — an NVIDIA engineer had written this on 2025-06-17:
I'm sorry this wasn't called out explicitly in the changelog, but the GSP initialization failures on some blackwell notebooks should be fixed in both of today's releases (570.169, 575.64).
A fix. Named releases. Tracked internally as NVBug 5287221. Absent from the changelog, which is why nobody downstream had connected it to their own broken machine.
I had already tested four driver branches, and every one of them was newer than 570.169. So I reasoned: we already carry that fix, this must be a different bug. I wrote that conclusion into my report as its opening argument.
That was the worst mistake I made, and it cost weeks.
Test the version they named
What finally unstuck it was my own machine's owner asking a blunt question: maybe check with the exact firmware and driver?
I had tested versions after the fix. I had never tested the fix.
570.169 works. nvidia-smi lists the GPU. CUDA kernels compiled for sm_120
run and return correct results. 570.172.08 works too. 570.190 and everything
after it — ten versions, every branch head, including the newest driver NVIDIA
currently ships — fails.
The fix landed in 570.169 and regressed again before 570.190, and has stayed broken ever since.
A fix present in an ancestor release does not mean it survived into descendants. When a vendor names a version, test that version. I had direct evidence of this exact pattern — a Fedora-adjacent bug report bisecting an identical failure to 595.45.04 good, 595.58.03 broken — sitting in my notes, and I did not apply it to my own case.
Asking a better question
The bisect gave me a two-release window and a working version. Now I could stop guessing.
Everything up to this point had asked "does removing this fix it?" — one boot per hypothesis, and I was 0 for 3. So I asked a different question: what actually differs?
I patched both drivers — the one that works and the one that doesn't — to hexdump every structure they hand GSP-RM before initialisation, generated from the same edit applied to each source tree so the output compared line for line.
Six payloads. GspSystemInfo at 936 bytes. The packed registry table at 1383.
GSP_ARGUMENTS_CACHED, GspFwWprMeta, the FSP chain-of-trust payload, the libos
region table.
All identical, except for per-boot DMA addresses and one field that is literally the size of the firmware image.
And I had a free control for the addresses: the failing boot logs three initialisation attempts, so I could diff attempt one against attempt two of the same driver in the same boot. They differ at exactly the same offsets and nowhere else.
The host was not the difference. The host had never been the difference. Which left exactly one thing in the box.
The ten-line patch
The firmware image ships inside the driver package. So: put the old image under the new driver, and see what happens.
The driver refuses, of course:
_kgspFwContainerVerifyVersion: GSP firmware image version mismatch:
got version 570.172.08, expected version 570.190
I spent a while assuming this was a signature check, which would have ended the
investigation. It is not. It is a portStringCompare against NV_VERSION_STRING
on an ELF section called .fwversion. Plain text. No cryptography anywhere near
it.
Ten lines to downgrade that check to a warning. Rebuild. Reboot.
NVRM: loading NVIDIA UNIX Open Kernel Module for x86_64 570.190
NVRM: FWVER OVERRIDE: accepting image version '570.172.08' against driver 570.190
GPU 0: NVIDIA GeForce RTX 5070 Laptop GPU (UUID: GPU-b3c9b96b-...)
The broken driver, with the old firmware image, initialises the GPU.
Same driver binary. Same machine. Same cold boot. One file swapped, and the card wakes up.
What I can prove, and what I can't
This is the part I had to be talked down from overstating, twice.
The tempting claim is "the bug is in the 570.190 firmware image." What the evidence actually supports is narrower:
Holding the driver constant and changing only the image flips the outcome. That makes the image the variable. It does not distinguish between the image is defective and the image and driver are each fine and disagree about something on this specific hardware.
The test that would separate them is not available to me. It would mean running the 570.190 image under a different host — nova-core — and nova cannot load NVIDIA-packaged firmware at all, failing before GSP-RM even runs.
There is a second limit worth naming. I compared six payloads. I did not compare register writes, timing, or the ordering of the FSP handshake. A driver-side behavioural change outside those structures is not fully excluded.
I mention both because I got this wrong in the first two drafts of my report, and a firmware engineer who spots one overclaim will discard the entire document. Saying it yourself, first, is cheaper.
Where it lives now
The GPU works. It runs on 570.169, pinned, with a comment in the config that says in as many words: a driver update on this machine is a regression, not an upgrade. CUDA runs. GPU containers work.
There is a second working configuration if I want a newer driver — the head of r570 with the last-good firmware image and that ten-line patch — verified with CUDA. It holds within a branch and not across one: try the same trick with a 610-series driver and you get an RPC ABI mismatch, because the structures the two sides exchange changed shape between branches.
It's filed upstream now, with the log, the bisect, the byte dumps and the transplant. Whether it gets fixed is not up to me. The image is opaque and signed, and the remaining diff is between two firmware builds only NVIDIA holds.
Six things I'd tell myself on day one
Read the log from the top. The loudest error is often the last one, and the last one is often a consequence. Five patches at the wrong layer.
Get a control before you get a hypothesis. nova-core booting the same silicon killed every hardware theory in a single measurement. I should have reached for it in week one, not week three.
"Does removing this fix it?" is a weak question. It costs one experiment per guess and tells you nothing when the answer is no. "What actually differs?" cost two boots and answered everything.
When a vendor names a version, test that version. Not a later one. A fix present in an ancestor does not mean it survived.
Instrument the thing you're about to override. One of my experiments forced a field to zero that was already zero, and reported "no change" — which reads exactly like evidence. It only surfaced because the patch logged the value alongside changing it.
Say what you cannot prove. Twice I wrote a conclusion stronger than my measurement, and twice someone caught it. The audit table I keep now has a line at the top noting how many claims have been withdrawn. It is uncomfortable and it is the most useful line in the document.
The RPC boundary between nvidia.ko and GSP-RM is the same shape of problem I work on
professionally: two sides that have to agree on a wire contract — DDS/RTPS there, this
vendor's RPC schema here — where you can't read one side's source and have to reconstruct its
behaviour from the outside.
The full technical write-up, including every negative result, the byte dumps and the patches, is in a public gist. The upstream bug is open-gpu-kernel-modules#1332. If this kind of hunt is your thing, see also Debugging a Bluetooth Kernel Regression on NixOS, the same kind of chase on a different chip.