The Watchdog That Watched a Disk Die Anyway

beta-cuda is the CI build server in my homelab — a GPU VM on Proxmox that runs pueue job queues for branch builds, each one dropping a cargo-target directory that can run to tens of gigabytes. To keep that from ever filling the disk, there's a watchdog: a Nushell script on a systemd timer, checking free space every ten minutes with a graduated response: log a notice, then delete old build artifacts, then delete all of them, then run the Nix garbage collector, each tier triggered at a lower free-space threshold.

On 2026-08-21 the disk hit zero anyway. The watchdog had seen it coming. Twice.


What the log actually said

Two runs, eleven minutes apart:

TimeFree spaceAction takenResult
01:0110 GBcleaned cargo targets older than 7 daysfreed nothing
01:1210 GBcleaned cargo targets older than 7 daysfreed nothing
01:210 GB—pueued couldn't write its state file, crash-looped out of systemd, killed a running build

Both runs reported the same number, ran the same action, changed nothing, and exited 0. A watchdog that is individually correct at every check and still misses the failure is a worse bug than one that's just wrong — it produces a clean log that looks like proof nothing was wrong.

The last thing standing between a warning and a crash-loop

The first bug: an off-by-one that was actually two bugs

The threshold ladder was written with strict <:

<  20 GB  — notice
<  15 GB  — clean targets older than 7 days
<  10 GB  — clean all targets
<   5 GB  — clean all targets + nix-collect-garbage -d

At a reported 10 GB free, 10 < 10 is false, so the guard fell into the 15 GB branch — the mildest action that matched, not the one that actually mattered. Switching every comparison to <= looked like the whole fix.

It wasn't the whole story. While tracking down why "10 GB" kept landing in the wrong branch, I found a second, independent bug sitting underneath it: df -BG rounds up, not down. 77.13 GiB free is reported as 78G. So a reported 10 doesn't mean somewhere in [10, 11) — it means the true figure is anywhere in (9, 10]. At the moment the guard read "10 GB" it may already have had barely over nine gigabytes free, one threshold tier further gone than the number on screen claimed. Both bugs pushed in the same direction: report a number gentler than reality, act one tier softer than the situation warranted.

I fixed both, shipped it, and moved on. Three hours later I went back and read the incident again, because something about the shape of it still didn't add up.

The fix that wasn't the fix

The off-by-one explained why the 15 GB action ran instead of the 10 GB action. It did not explain why the 15 GB action — delete cargo targets older than 7 days — freed nothing on either pass. That part turned out to be host-specific, and worse than a threshold bug: measured on beta-cuda that morning, there was 55 GB of cargo-target directories on disk, and the oldest of them was six days old. "Older than 7 days" was never going to match anything, on this host, at this moment, no matter which threshold tier picked it.

A mild action that matches zero files needs to fall through to the next, stronger action immediately. Instead the script was one if / else if chain — whichever branch matched first ran, and the rest were skipped for that entire pass, correct threshold or not. The guard had the harsher actions (delete all stale targets, run Nix GC) sitting right there in the same file the whole time. It just structurally couldn't reach for them without waiting for the next tick.

And the next tick was the other half of it. The timer polled every ten minutes. The disk went from a reported 10 GB free to 0 in nine — faster than the watchdog was ever asked to look again. No threshold, however well-tuned, survives a check interval longer than the time-to-failure. That's not a bug you fix by adjusting a number; it's a bug in the relationship between two numbers that were never chosen with each other in mind.

Reading the log from the top instead of the last line

The actual fix

Two changes, landed together three hours after the first one:

  1. Escalate within a single run. The if / else if chain became sequential if blocks, each re-measuring free space and falling through to the next tier if it's still low — so one run can go from "clean old targets" all the way to "run the garbage collector" without waiting for another invocation.
  2. Poll every two minutes, not ten. This was never a cost-based decision to begin with — nobody had measured what a run cost. I did: 868ms of CPU over 1.7 seconds of wall time, du walk over 55 GB of targets included. At a 2-minute cadence that's well under 1% of one core, on a VM whose whole job is to have spare cores sitting around.

The alert semantics changed too. The guard used to fire a warning any time a pass freed nothing — which is exactly the false alarm that fired twice before the incident, since a mild action legitimately freeing nothing is now expected to happen on the way to a stronger one. It now fires only after every tier has been tried in that run and the disk is still low: not "this step did nothing," but "I am out of levers and someone should look." Getting paged for the intermediate no-ops would have trained me to ignore the alert by the second week.

Lessons

A monitoring loop's interval is a variable, not a constant you get to assume away. Thresholds tuned in isolation from the polling rate can all be individually sensible and still leave a gap the failure walks straight through. The question isn't just "at what free-space level should I act," it's "is my sampling rate faster than the fastest way this can fail."

"Matched nothing" and "problem handled" are different outcomes, and conflating them is how a watchdog goes quiet at the worst moment. A guard that exits 0 after freeing zero bytes has told systemd everything is fine. If a mild remedy has nothing to remedy, that's information, not success — it should trigger the next remedy, not a clean exit.

The first fix that makes sense isn't always the fix. The off-by-one comparison was a real bug and a satisfying one — a one-character diff with an obvious cause. It shipped first because it was the bug I could see fastest, and it was genuinely necessary. It just wasn't sufficient, and I only found that out by going back and asking why the "working" threshold still freed nothing on a real host with real data on it.