Why a Wayland Kiosk Can't Live in a Container
Table of Contents
Three fixes, three dead ends, and a root cause underneath all of them: no container runtime can give a process its own virtual terminal.
I wanted a TV in the living room to just turn on and show Kodi, no keyboard, no monitor, nothing plugged in but the HDMI cable and a phone remote. The plan was cage (a minimal Wayland kiosk compositor) plus Kodi, tucked away in an unprivileged Linux system container on the home hypervisor so it lived alongside everything else instead of needing its own box. It never worked, and by the time I found out why, I'd learned more about virtual terminals than I ever expected to need.
The short version: virtual terminals (VTs) are host-global kernel state, one register
the whole machine shares, and no container runtime hands a container its own slice of
it. Not Incus, not Docker, not Podman, not plain LXC. Any Wayland compositor using
seatd/libseat for display access is VT-bound underneath, so every container
technology hits the exact same wall.
The symptom that lied first
systemctl status said active (running). The unit was actually crash-looping every
fifteen seconds, for hours, and nothing was reaching the journal. Restart=always on a
short RestartSec was outrunning anything that might have logged the real error before
the next restart clobbered it, so the status line was actively lying to me. I only got
the real stderr by running the command by hand, outside the unit:
systemd-run --wait --pipe <the command>
Fix one: an environment variable that "worked" for the wrong reason
First guess: a runtime-directory environment variable pointing at the wrong path. I
hardcoded the right value, checked the process environment, saw exactly what I expected
to see there, and the failure didn't budge. Turns out the check itself was worthless:
isolating that one variable in a minimal systemd-run test, outside the real unit,
showed that PAM's own session setup was independently supplying the same correct value
in both the broken and the "fixed" version. Reading the value back couldn't tell the two
apart, because something else in the system was already handing out the value my fix
was taking credit for.
That's the trap worth naming: a fix can pass its own verification and still be inert, if something else already produces the same observable state regardless of whether the fix ran. The way out was isolating one variable at a time in a minimal reproduction, not trusting an observed value inside the full, complicated unit.
Fix two: pass a spare VT into the container
With the environment variable ruled out, the real failure came into focus. cage needs a
"seat" for DRM access, and none of the paths libseat tries succeed inside a container:
no seat daemon, no resolvable login session, no /dev/tty0.
So, next attempt: pass an unused virtual terminal into the container so a seat daemon inside it could claim that VT on its own. Didn't work. The seat daemon asks one question, "am I the foreground console right now," and the kernel answers from a single global register no matter who's asking or from where:
There's no way to give a container an isolated VT for this, confirmed live rather than inferred from a manual page. wlroots#3200 and systemd#31227 are the upstream threads covering the same gap, one from the compositor side and one from session management.
Fix three: cage's documented setuid escape hatch
Cage's own wiki documents a setuid-root path meant to
skip seat management entirely.
I built it properly: a real suid wrapper binary, since the package store is mounted
nosuid and a bare chmod on the built executable does nothing there. The failure was
byte-for-byte identical with the wrapper in place. Best guess: the documented behavior
is vestigial, a leftover from a pre-libseat version of the compositor, no longer wired
into anything the current code path touches.
The fix: run it on the host, on a real VT
Three levers, three dead ends, and what was left underneath all three was the one thing
a container structurally cannot manufacture: a real virtual terminal. So the kiosk runs
directly on the hypervisor now, on a dedicated, otherwise-unused VT (checked first that
nothing else on the host was using its display). It runs under the systemd unit template
from cage's
own boot-time example,
ConditionPathExists=/dev/tty0 included.
Two smaller traps on the way to something usable
Getting something to render was only half the job. First it rendered at the wrong resolution: the compositor's default output mode spans every connected display into one combined canvas, and the hypervisor's own built-in screen, otherwise sitting there doing nothing, still counted as connected. Pinning the compositor to a single output fixed it.
Then came the part where I wanted the whole thing controllable from a phone, no keyboard ever touching it. Three separate Kodi settings each "worked" in the sense that they parsed without a single error and changed nothing:
- One settings file is an override layer that never triggers the callback that actually starts a service. It just describes preferred state to nobody in particular.
- The settings file that does work can lose a race against the app's own default file if it's seeded too late on first boot.
- A webserver with authentication on and a blank password just refuses to bind the port. No log line, no error, nothing.
Turns out the advancedsettings.xml symptom specifically is a known, still-unresolved community issue, not something particular to this setup.
Lessons
A fix that passes its own verification can still be doing nothing. If something else in the system can independently produce the same observable state, reading that state back doesn't tell you whether your change caused it. Isolate the one variable in a minimal reproduction instead of re-checking the same observable inside the full, busy system.
When three fixes in the same category all fail the same way, that's usually the signal to stop looking in that category. An environment variable, a passed-in VT, a documented escape hatch: all three assumed the constraint was something a container could be configured around. Three confirmed dead ends pointed at the constraint sitting one layer lower, in what a container fundamentally is, not in how this particular one was set up.
And a setting that parses without error hasn't been verified to do anything at all. Three separate Kodi settings cleared that bar and changed nothing. The only check that means anything is behavioral: does the thing the setting is supposed to cause actually happen.