Six Days Inside a GPU That Would Not Wake Up

I just got a laptop with an RTX 5070.

nvidia-smi had exactly one thing to say, and it said it instantly:

No devices were found

Everything else looked healthy. All four kernel modules loaded. lspci cheerfully reported Kernel driver in use: nvidia. No BAR errors, no IOMMU complaints, no firmware load failures. The machine was quite certain nothing was wrong.

Everything looks fine except the one thing that matters

This is the story of the six days it took to find out why, which ends somewhere I did not expect: in a binary I am not allowed to read, reached by a route that should not have worked.

The RTX 5070 Laptop GPU that refused to wake up


The thing you have to understand first

I lost most of the first day to a wrong mental model, so let me hand you the right one up front.

When you install "the NVIDIA driver" on a modern card, you are not installing the driver. Since Turing, NVIDIA moved the bulk of the Resource Manager — the layer that owns initialisation, memory, channels, clocks, display — off the CPU and onto a RISC-V core inside the GPU called the GSP, the GPU System Processor. nvidia.ko is largely a client that marshals RPCs to it.

On Blackwell there is no CPU-side fallback at all. NVreg_EnableGpuFirmware=0 is accepted and silently ignored, because the CPU-side Resource Manager those chips would need was never written.

Where the Resource Manager actually runs

So "the driver fails to initialise the GPU" is, on this hardware, a misleading sentence. What actually happens is: firmware on the GPU declines to initialise itself, and tells the host why over an RPC.

That reframing is most of the battle. It also means that when things go wrong, a large part of the system you need to inspect is a signed binary running on a processor you cannot attach a debugger to.


The message that ate five patches

The kernel log had a line in it that looked like the answer:

NVRM: GPU0 _kgspBootGspRm: unexpected WPR2 already up, cannot proceed with booting GSP
NVRM: _kgspBootGspRm: (the GPU is likely in a bad state and may need to be reset)

WPR2 is Write-Protected Region 2, a locked area of video memory where GSP-RM's image and heap live. The message says: someone already set this up, I cannot set it up again, the card needs a reset.

That reads like a complete diagnosis. It is the line that appears in nearly every report of this failure I later found — on Arch, on Fedora, on Ubuntu, in NVIDIA's own tracker. It sent me down a long tunnel: FLR, secondary bus reset, remove-and-rescan, D3cold, ACPI power resources, vendor WMI power cuts. I wrote five kernel patches trying to tear WPR2 down before the driver touched it.

None of it worked, and eventually I did the thing I should have done on day one. I stopped reading dmesg | tail and read the log from the top.

The first attempt

The retry, tripping over what attempt one already built

The first attempt gets all the way through. FSP accepts the chain of trust in 0.78 seconds. WPR2 is built. GSP-FMC runs. GSP-RM boots, runs, and then returns NV_ERR_INVALID_ARGUMENT from its own initialisation.

Then the driver retries. The retry finds WPR2 up — because attempt one built it — and refuses. Four sixty-second timeouts later, that refusal is the last thing in the log, so it is the thing everyone quotes.

The message everybody chases is an artifact of a failure that already happened. Five patches at the wrong layer, because I read the end of a log instead of the beginning.


A GPU that works, sitting right next to one that doesn't

The thing that broke the investigation open was not a patch. It was a control.

nova-core is the in-tree Rust NVIDIA driver in upstream Linux — a completely independent implementation of the same boot sequence, written by different people against the same silicon. On kernel 7.2-rc2 I removed the NVIDIA stack entirely and booted it.

nova-core 0000:01:00.0: NVIDIA (Chipset: GB206, Architecture: BlackwellGB20x, Revision: a.1)
nova-core 0000:01:00.0: GPU name: NVIDIA GeForce RTX 5070 Laptop GPU
[drm] Initialized nova-drm 0.0.0 for nova-core.nova-drm.0 on minor 0

That second line is not cosmetic. It is printed at gsp/boot.rs:159, after wait_gsp_init_done() at line 152, and after the GetGspStaticInfo RPC at 155 whose reply carries that string. For nova to print it, GSP-RM must have booted, reached init-done, and answered a structured RPC — on this exact card, in the state UEFI leaves it, with no reset and no preparation.

In one measurement, every hardware and platform explanation died. The silicon is fine. The FSP path is fine. WPR2 is fine. The card is not broken.

nova gives no CUDA and never will — it is a display driver. But as a control it was worth more than any patch I wrote.

Same silicon, a different driver, and a control that killed every hardware theory


Ten thousand words of arguments, all correct

If the hardware works and another driver can drive it, then NVIDIA's driver must be telling GSP-RM something it doesn't like. So I built a harness: a patched driver with a regkey that selects which single difference to apply on each boot, and the GSP-RM error trace as the observable.

Then I went through everything the driver hands the firmware, one boot at a time.

Every argument, varied and measured

Twenty-two experiments. The results were maddening in a specific way: nothing was wrong. fbSize: no change. The nine layout offsets: not actually a difference, ACR fills them in. bIsPrimary: already zero, the value I was about to "fix" it to. Drop the libos regions that nova omits: the firmware panics, so they are required, not spurious. Zero the message-queue geometry: the queue never comes up.

Then the big one. I zeroed all thirty-nine fields NVIDIA sends that nova does not, reducing the struct to exactly what a working driver sends.

Byte-identical failure.

I remember sitting there genuinely stuck. Every input was correct. The firmware was rejecting an initialisation for which I had individually verified every single argument.


The sentence that had been sitting in public for a year

Out of ideas on the technical side, I went reading. Arch forums. Ubuntu Discourse. Launchpad. Fedora. Reddit. NVIDIA's own GitHub tracker.

And in issue #876 — a thread about somebody else's laptop, closed as resolved — an NVIDIA engineer had written this on 2025-06-17:

I'm sorry this wasn't called out explicitly in the changelog, but the GSP initialization failures on some blackwell notebooks should be fixed in both of today's releases (570.169, 575.64).

A fix. Named releases. Tracked internally as NVBug 5287221. Absent from the changelog, which is why nobody downstream had connected it to their own broken machine.

I had already tested four driver branches, and every one of them was newer than 570.169. So I reasoned: we already carry that fix, this must be a different bug. I wrote that conclusion into my report as its opening argument.

That was the worst mistake I made, and it cost weeks.


Test the version they named

What finally unstuck it was my own machine's owner asking a blunt question: maybe check with the exact firmware and driver?

I had tested versions after the fix. I had never tested the fix.

The bisect

570.169 works. nvidia-smi lists the GPU. CUDA kernels compiled for sm_120 run and return correct results. 570.172.08 works too. 570.190 and everything after it — ten versions, every branch head, including the newest driver NVIDIA currently ships — fails.

The fix landed in 570.169 and regressed again before 570.190, and has stayed broken ever since.

A fix present in an ancestor release does not mean it survived into descendants. When a vendor names a version, test that version. I had direct evidence of this exact pattern — a Fedora-adjacent bug report bisecting an identical failure to 595.45.04 good, 595.58.03 broken — sitting in my notes, and I did not apply it to my own case.


Asking a better question

The bisect gave me a two-release window and a working version. Now I could stop guessing.

Everything up to this point had asked "does removing this fix it?" — one boot per hypothesis, and I was 0 for 3. So I asked a different question: what actually differs?

I patched both drivers — the one that works and the one that doesn't — to hexdump every structure they hand GSP-RM before initialisation, generated from the same edit applied to each source tree so the output compared line for line.

Six payloads, byte for byte

Six payloads. GspSystemInfo at 936 bytes. The packed registry table at 1383. GSP_ARGUMENTS_CACHED, GspFwWprMeta, the FSP chain-of-trust payload, the libos region table.

All identical, except for per-boot DMA addresses and one field that is literally the size of the firmware image.

And I had a free control for the addresses: the failing boot logs three initialisation attempts, so I could diff attempt one against attempt two of the same driver in the same boot. They differ at exactly the same offsets and nowhere else.

The host was not the difference. The host had never been the difference. Which left exactly one thing in the box.


The ten-line patch

The firmware image ships inside the driver package. So: put the old image under the new driver, and see what happens.

The driver refuses, of course:

_kgspFwContainerVerifyVersion: GSP firmware image version mismatch:
  got version 570.172.08, expected version 570.190

I spent a while assuming this was a signature check, which would have ended the investigation. It is not. It is a portStringCompare against NV_VERSION_STRING on an ELF section called .fwversion. Plain text. No cryptography anywhere near it.

The transplant

Ten lines to downgrade that check to a warning. Rebuild. Reboot.

NVRM: loading NVIDIA UNIX Open Kernel Module for x86_64  570.190
NVRM: FWVER OVERRIDE: accepting image version '570.172.08' against driver 570.190
GPU 0: NVIDIA GeForce RTX 5070 Laptop GPU (UUID: GPU-b3c9b96b-...)

The broken driver, with the old firmware image, initialises the GPU.

Same driver binary. Same machine. Same cold boot. One file swapped, and the card wakes up.


What I can prove, and what I can't

This is the part I had to be talked down from overstating, twice.

The tempting claim is "the bug is in the 570.190 firmware image." What the evidence actually supports is narrower:

What the transplant shows

What it doesn't show

Holding the driver constant and changing only the image flips the outcome. That makes the image the variable. It does not distinguish between the image is defective and the image and driver are each fine and disagree about something on this specific hardware.

The test that would separate them is not available to me. It would mean running the 570.190 image under a different host — nova-core — and nova cannot load NVIDIA-packaged firmware at all, failing before GSP-RM even runs.

There is a second limit worth naming. I compared six payloads. I did not compare register writes, timing, or the ordering of the FSP handshake. A driver-side behavioural change outside those structures is not fully excluded.

I mention both because I got this wrong in the first two drafts of my report, and a firmware engineer who spots one overclaim will discard the entire document. Saying it yourself, first, is cheaper.


Where it lives now

The GPU works. It runs on 570.169, pinned, with a comment in the config that says in as many words: a driver update on this machine is a regression, not an upgrade. CUDA runs. GPU containers work.

There is a second working configuration if I want a newer driver — the head of r570 with the last-good firmware image and that ten-line patch — verified with CUDA. It holds within a branch and not across one: try the same trick with a 610-series driver and you get an RPC ABI mismatch, because the structures the two sides exchange changed shape between branches.

It's filed upstream now, with the log, the bisect, the byte dumps and the transplant. Whether it gets fixed is not up to me. The image is opaque and signed, and the remaining diff is between two firmware builds only NVIDIA holds.

Where most of the six days actually happened


Six things I'd tell myself on day one

Read the log from the top. The loudest error is often the last one, and the last one is often a consequence. Five patches at the wrong layer.

Get a control before you get a hypothesis. nova-core booting the same silicon killed every hardware theory in a single measurement. I should have reached for it in week one, not week three.

"Does removing this fix it?" is a weak question. It costs one experiment per guess and tells you nothing when the answer is no. "What actually differs?" cost two boots and answered everything.

When a vendor names a version, test that version. Not a later one. A fix present in an ancestor does not mean it survived.

Instrument the thing you're about to override. One of my experiments forced a field to zero that was already zero, and reported "no change" — which reads exactly like evidence. It only surfaced because the patch logged the value alongside changing it.

Say what you cannot prove. Twice I wrote a conclusion stronger than my measurement, and twice someone caught it. The audit table I keep now has a line at the top noting how many claims have been withdrawn. It is uncomfortable and it is the most useful line in the document.


The RPC boundary between nvidia.ko and GSP-RM is the same shape of problem I work on professionally: two sides that have to agree on a wire contract — DDS/RTPS there, this vendor's RPC schema here — where you can't read one side's source and have to reconstruct its behaviour from the outside.

The full technical write-up, including every negative result, the byte dumps and the patches, is in a public gist. The upstream bug is open-gpu-kernel-modules#1332. If this kind of hunt is your thing, see also Debugging a Bluetooth Kernel Regression on NixOS, the same kind of chase on a different chip.