One Flake, Seven Hosts, and the Backup That Died Without Telling Anyone
Table of Contents
Two hosts each carried their own copy of the same backup config, and the same structural mistake in both of them let one backup silently stop running for two days.
My homelab flake covers seven hosts today — laptops, a desktop, and a Proxmox-hosted
VM fleet — from one flake.nix and one hosts/ directory. Each host imports a shared
set of modules for the things that are genuinely uniform (Tailscale join, off-box
backup, SSH) and keeps only what's actually host-specific in its own default.nix. I
like to think that's the obviously correct shape for managing more than one machine,
and it is — but I didn't always have it factored this way, and the version before I
extracted the shared module is where the interesting bug was.
What "single source of truth" actually looks like here
# flake.nix — a helper per host class, not one per machine
mkNixosConfig = host: profile: lib.nixosSystem {
modules = [ { nixpkgs.hostPlatform = system; } ./profiles/${profile} ];
};
mkServerConfig = host: lib.nixosSystem {
modules = [ inputs.disko.nixosModules.disko ./profiles/server ./hosts/${host} ];
};
nixosConfigurations = {
ZenS13 = mkNixosConfig "ZenS13" "amd";
NitroAN515 = mkNixosConfig "NitroAN515" "nvidia";
beta-cuda = mkServerConfig "beta-cuda";
beta-meerkat = mkServerConfig "beta-meerkat";
alpha-otter = mkServerConfig "alpha-otter";
# ...
};
Adding a host is a new entry in this list plus a hosts/<name>/default.nix importing
whichever shared modules apply — modules/nixos/tailscale-client.nix,
modules/nixos/restic-client.nix, and so on. The shared modules carry the actual
behavior; the host file carries only what makes that machine different: its hostname,
its hardware quirks, which shared modules it opts into and with what options.
That's the shape I aim for. Before I wrote modules/nixos/restic-client.nix, though, I
had two hosts — NitroAN515 and ZenS13 — each carrying its own hand-written,
nearly-identical restic backup block. "Nearly identical" is the load-bearing phrase.
Where the copies actually agreed, which was the problem
The two inline blocks weren't a bug on their own — each host's restic job ran fine in
isolation, for a long time. The bug was structural, and I'd given both hosts the exact
same version of it: on both NitroAN515 and ZenS13, I'd defined the remotebackup
job (the actual home-directory backup) inside home-manager.users.circle. A user-level
systemd unit, scoped to whichever home-manager generation is currently active, not to
the NixOS generation. (NitroAN515 also had a separate, genuinely root-level job —
pgbackup, for Postgres dumps — which is real config asymmetry between the two hosts,
just not the one that bit me here.)
That home-manager scoping is invisible in day-to-day use. The unit shows up under
systemctl --user, runs on its timer, backs up to the restic REST server, same as any
other service — until something rewrites just the home-manager generation without
touching the NixOS one, which is exactly what a standalone nh home switch does. On
2026-07-24 that's what happened on NitroAN515: a home-manager-only switch regenerated
~/.config/systemd/user from the current home-manager config, and the backup timer,
existing only inside the previous home-manager generation, wasn't in the new one. It
got deleted. Silently — no error, no failed unit, nothing to notice unless I went
looking for a timer that was no longer there. I didn't notice for two days.
ZenS13 had the identical exposure the whole time; it just hadn't been hit yet.
The fix was extraction, not a bigger workaround
My first instinct after finding a bug like that is to patch the specific failure — pin
the timer so nh home switch can't touch it, or add a check that alerts if the unit
goes missing. Both are real options and both would have worked. What I actually did was
upstream of either: stop having two copies of this config in the first place, and
while I was at it, put the surviving copy at the level that isn't vulnerable to this
failure mode at all.
# modules/nixos/restic-client.nix
{ config, lib, ... }:
{
options.resticClient = {
enable = lib.mkEnableOption "off-box restic backup";
name = lib.mkOption { type = lib.types.str; };
};
config = lib.mkIf config.resticClient.enable {
systemd.services.remotebackup = { /* ... */ };
systemd.timers.remotebackup = { /* ... */ };
# paths, schedule, sops secrets — parameterized by resticClient.name
};
}
I folded both hosts' blocks into one module, with the only real per-host variation (the
client name used for the htpasswd user, the repo namespace, the sops secret prefix)
exposed as resticClient.{enable,name}. I moved both hosts' jobs from home-manager
scope to NixOS scope in the same change, because a root-level systemd unit only
changes when the NixOS generation changes, and a nixos-rebuild switch is the
operation that's actually meant to change what services exist on a machine. The whole
nh home switch failure mode was only possible because I'd put the job somewhere a
home-manager-scoped switch had the authority to touch.
I left one thing inline on NitroAN515 rather than folding it fully into the module:
its Postgres dump job, which needs access to the database's own dump directory in a way
the generic module has no reason to know about. I moved it to 03:15 instead of the
shared module's 03:00, so the two jobs backing up to the same restic repository don't
try to lock it at the same moment.
What one flake actually buys, concretely
The abstract case for "one source of truth" is easy to state and easy to nod along
with. The concrete version, after this, is what actually convinced me: a bug that
existed because I'd let two hosts each carry their own copy of a piece of config —
identical enough that I had no reason to suspect either one — became structurally
impossible to reintroduce once there was only one copy. Every host that turns
resticClient.enable on now gets the version that lives at NixOS scope, not whichever
version I happened to hand-write for that machine on whatever day I set it up. The next
host I add inherits the current, already-debugged shape by default — it doesn't get to
make the mistake either of the first two made, because the mistake isn't a copy-paste
away anymore.
Lessons
"Nearly identical" config across hosts is not a safety net — a shared mistake, copied
twice, is still a mistake. Two hand-written blocks that agree with each other invited
me to assume agreement meant correctness. Here it meant the opposite: I'd given both
hosts the same home-manager-scoping problem, and having two independent copies of it
didn't help me catch the bug any faster than one copy would have — it just meant the
failure was waiting on whichever host got a nh home switch first.
A silent failure is worse than a loud one, and "no error" is not the same as "working." The backup job that died left nothing behind for me to notice — no failed unit, no red text, just a timer that quietly stopped existing. I didn't fix this by adding monitoring for that specific failure; I fixed it by moving the config to a layer where that class of failure has no mechanism to occur.
Extracting a shared module is itself a bug fix, not just tidying. The value wasn't "less duplicated code" as a style preference — deduplicating forced me to ask why the two copies differed, and that question is what surfaced the actual bug.