The Watchdog That Watched a Disk Die Anyway
Table of Contents
A disk-space guard observed the exact failure condition it was built to catch, twice, and the disk hit zero nine minutes later. Two bugs, two fixes, three hours apart.
beta-cuda is the CI build server in my homelab — a GPU VM on Proxmox that runs pueue
job queues for branch builds, each one dropping a cargo-target directory that can run to
tens of gigabytes. To keep that from ever filling the disk, there's a watchdog: a Nushell
script on a systemd timer, checking free space every ten minutes with a graduated
response: log a notice, then delete old build artifacts, then delete all of them, then
run the Nix garbage collector, each tier triggered at a lower free-space threshold.
On 2026-08-21 the disk hit zero anyway. The watchdog had seen it coming. Twice.
What the log actually said
Two runs, eleven minutes apart:
| Time | Free space | Action taken | Result |
|---|---|---|---|
| 01:01 | 10 GB | cleaned cargo targets older than 7 days | freed nothing |
| 01:12 | 10 GB | cleaned cargo targets older than 7 days | freed nothing |
| 01:21 | 0 GB | — | pueued couldn't write its state file, crash-looped out of systemd, killed a running build |
Both runs reported the same number, ran the same action, changed nothing, and exited 0. A watchdog that is individually correct at every check and still misses the failure is a worse bug than one that's just wrong — it produces a clean log that looks like proof nothing was wrong.
The first bug: an off-by-one that was actually two bugs
The threshold ladder was written with strict <:
< 20 GB — notice
< 15 GB — clean targets older than 7 days
< 10 GB — clean all targets
< 5 GB — clean all targets + nix-collect-garbage -d
At a reported 10 GB free, 10 < 10 is false, so the guard fell into the 15 GB
branch — the mildest action that matched, not the one that actually mattered. Switching
every comparison to <= looked like the whole fix.
It wasn't the whole story. While tracking down why "10 GB" kept landing in the wrong
branch, I found a second, independent bug sitting underneath it: df -BG rounds up,
not down. 77.13 GiB free is reported as 78G. So a reported 10 doesn't mean
somewhere in [10, 11) — it means the true figure is anywhere in (9, 10]. At the
moment the guard read "10 GB" it may already have had barely over nine gigabytes free,
one threshold tier further gone than the number on screen claimed. Both bugs pushed
in the same direction: report a number gentler than reality, act one tier softer than
the situation warranted.
I fixed both, shipped it, and moved on. Three hours later I went back and read the incident again, because something about the shape of it still didn't add up.
The fix that wasn't the fix
The off-by-one explained why the 15 GB action ran instead of the 10 GB action. It
did not explain why the 15 GB action — delete cargo targets older than 7 days — freed
nothing on either pass. That part turned out to be host-specific, and worse than a
threshold bug: measured on beta-cuda that morning, there was 55 GB of cargo-target
directories on disk, and the oldest of them was six days old. "Older than 7 days" was
never going to match anything, on this host, at this moment, no matter which threshold
tier picked it.
A mild action that matches zero files needs to fall through to the next, stronger action
immediately. Instead the script was one if / else if chain — whichever branch matched
first ran, and the rest were skipped for that entire pass, correct threshold or not. The
guard had the harsher actions (delete all stale targets, run Nix GC) sitting right
there in the same file the whole time. It just structurally couldn't reach for them
without waiting for the next tick.
And the next tick was the other half of it. The timer polled every ten minutes. The disk went from a reported 10 GB free to 0 in nine — faster than the watchdog was ever asked to look again. No threshold, however well-tuned, survives a check interval longer than the time-to-failure. That's not a bug you fix by adjusting a number; it's a bug in the relationship between two numbers that were never chosen with each other in mind.
The actual fix
Two changes, landed together three hours after the first one:
- Escalate within a single run. The
if/else ifchain became sequentialifblocks, each re-measuring free space and falling through to the next tier if it's still low — so one run can go from "clean old targets" all the way to "run the garbage collector" without waiting for another invocation. - Poll every two minutes, not ten. This was never a cost-based decision to begin
with — nobody had measured what a run cost. I did: 868ms of CPU over 1.7 seconds of
wall time,
duwalk over 55 GB of targets included. At a 2-minute cadence that's well under 1% of one core, on a VM whose whole job is to have spare cores sitting around.
The alert semantics changed too. The guard used to fire a warning any time a pass freed nothing — which is exactly the false alarm that fired twice before the incident, since a mild action legitimately freeing nothing is now expected to happen on the way to a stronger one. It now fires only after every tier has been tried in that run and the disk is still low: not "this step did nothing," but "I am out of levers and someone should look." Getting paged for the intermediate no-ops would have trained me to ignore the alert by the second week.
Lessons
A monitoring loop's interval is a variable, not a constant you get to assume away. Thresholds tuned in isolation from the polling rate can all be individually sensible and still leave a gap the failure walks straight through. The question isn't just "at what free-space level should I act," it's "is my sampling rate faster than the fastest way this can fail."
"Matched nothing" and "problem handled" are different outcomes, and conflating them is how a watchdog goes quiet at the worst moment. A guard that exits 0 after freeing zero bytes has told systemd everything is fine. If a mild remedy has nothing to remedy, that's information, not success — it should trigger the next remedy, not a clean exit.
The first fix that makes sense isn't always the fix. The off-by-one comparison was a real bug and a satisfying one — a one-character diff with an obvious cause. It shipped first because it was the bug I could see fastest, and it was genuinely necessary. It just wasn't sufficient, and I only found that out by going back and asking why the "working" threshold still freed nothing on a real host with real data on it.