TL;DR
- I joined eight surplus all-in-ones to my office k3s cluster as workers (the build is in the previous post). Six things went wrong, and all six failed quietly.
- A scheduling taint left 96 cores running zero workload pods while three small VMs ran 66.
- Prose inside a YAML folded scalar became k3s command-line arguments, and k3s silently dropped every flag after it.
- A DHCP lease and a pinned node IP disagreed. The login greeter suspended WiFi-only machines after 20 minutes. One
mokutilcommand broke TPM auto-unlock. A weekly patch play skipped its own job for weeks. - The common fix is boring: read back what the system actually did and assert on that, not on the exit code.
1. The taint that idled the fleet
My original plan, from a proposal doc, was to taint the panels so nothing but the kiosk would run on them. What shipped was a softer version, PreferNoSchedule, because I also wanted Longhorn on the panels, and Longhorn’s DaemonSets would be locked out by a hard taint. Changing Longhorn’s toleration setting needs every volume detached, which means downtime just to satisfy a taint.
PreferNoSchedule is a preference. The scheduler won’t put anything on a tainted node until the other nodes are genuinely full, and my other nodes are small VMs. “Full” arrived long before the hardware ran out.
When I finally counted, it looked like this:
| Nodes | Allocatable | Workload pods |
|---|---|---|
| 3 small agent VMs | 12 cores, 23 GiB | 66 |
| 8 all-in-one panels | 96 cores, 170 GiB | 0 (DaemonSets only) |
One VM was at 72 percent of its CPU requests on four cores with 96 idle cores next to it. The cluster believed it was nearly out of room.
The taint wasn’t what was protecting the kiosk anyway. system-reserved is: a hard kubelet guarantee that holds CPU and memory back from pods. With the wall live, a panel uses about 3.5 GiB and a tiny slice of CPU. I removed the taint, kept a form-factor label so any workload that really must avoid a desktop can say so with node affinity, and bumped the memory reservation from 3 GiB to 4 GiB. My first number was below the machine’s real usage, which isn’t a reservation at all.
The lesson: a protective control you can’t name the benefit of is costing you something. Write down what the taint is for before you add one, and count what it’s costing.
2. YAML turned prose into argv
This one stung. The install arguments live in a folded scalar:
k3s_extra_agent_args: >-
--node-ip={{ _aio_node_ip }}
--node-label=example.io/form-factor=aio-panel
I’d added a helpful note inside that block explaining a flag’s spelling. In a >- scalar every line is literal text. A # isn’t a comment, it’s text. The note became words on a command line, roughly:
... --node-ip=10.x.x.x defined: -register-with-taints # Verified ... --node-taint=... --kubelet-arg=...
k3s parses flags with a library that stops at the first positional argument. Everything after the prose, including the taint and the kubelet reservation, was silently dropped. The agent started happily. Seven panels registered with their full 16 CPUs and 24 GiB allocatable and no reservation, which is the opposite of what the playbook exists to do.
It got worse: the note contained backticks, and the install variable is expanded by a shell, so they ran as command substitution.
What I did about it:
- Prose goes in real YAML comments above the block, with a warning there that says “nothing but flags below”.
- A post-join task reads the installed
k3s-agent.servicefile and asserts that the label is in it and that no#appears in the command line (which would mean prose leaked in again). - It also asserts the reservation is in
config.yaml. - A CI regression test pins the shape.
The fail message tells you to uninstall and rejoin, because a panel without its reservation looks healthy and is invisible to the GPU plugin and the storage guard. Which brings me to the general rule: “did it join?” is not a sufficient check. Seven nodes were Ready, labelled in some places, and completely unguarded.
One more footnote from the same episode: the k3s flag is --node-taint. --register-with-taints is the kubelet’s flag, and k3s refuses it at startup.
3. DHCP vs a pinned node IP
I pin --node-ip at install time to the wired interface’s address. If the DHCP lease moves, the node’s registered address is wrong and it goes NotReady until reinstalled. These machines are dual-homed (wired for the cluster, WiFi for management), and two things happened before I had reservations:
- One unit’s WiFi lease disappeared and I managed it over the cluster NIC for a while.
- Another’s wired interface sat up at gigabit with no IPv4 address for a day. Its lease had earlier been held by its own WiFi adapter, and NetworkManager’s conflict detection saw the other holder’s MAC address, which was the machine’s own, and refused the address.
Fix: fixed-IP reservations on the router for both NICs on all eight panels, and a cable-pull test on one panel to prove the node survives. The wired-versus-WiFi default route is another quiet trap, since the metrics only decide which interface wins today. Pinning --node-ip removes the dependency.
4. The greeter that suspended the fleet
After a patch reboot on 2026-08-31, the machines went dark roughly 20 minutes later. They’d come up healthy. I verified them. Then they vanished.
After a reboot these units sit at the GDM login screen, because I deliberately don’t use autologin. And the GNOME greeter’s shipped default suspends the whole machine on AC after 20 minutes of inactivity. WiFi-only management means there’s no Wake-on-LAN. A suspended unit stays dark until someone presses a key.
I’d also been chasing “random unreachable” units that I’d blamed on WiFi power saving, with pings that took hundreds of milliseconds. Later evidence says some of those slow readings were the machines suspending.
The fix is a playbook that sets the greeter’s sleep behaviour to never suspend, written to the system dconf so a user-level idle delay can’t win, and shortens display blanking to 10 minutes instead. The box stays reachable and the screen goes off sooner. An all-in-one’s panel dominates its idle draw, so that’s where the saving should come from (reasoning, not a measurement).
I didn’t switch on autologin to hide this. It would have fixed the symptom, but it would also have traded a real login for passwordless sudo on an unattended machine, to fix something SSH never depended on.
5. One mokutil command broke TPM unlock
Root is LUKS2, unlocked by the TPM through Clevis, sealed to PCR 7. It’s easy to lump every Secure Boot change together. The actual rules, one of them learned the hard way on one unit:
| Action | Measured into | Breaks auto-unlock? |
|---|---|---|
| Enrolling a MOK for the signed NVIDIA module | PCR 14 (on Debian 13’s shim 15.4 or newer) | No |
mokutil --disable-validation, BIOS Secure Boot toggles | PCR 7 | Yes |
| Firmware KEK/db/dbx updates | PCR 7 | Yes |
mokutil --disable-validation on one machine changed PCR 7, auto-unlock stopped working, and that unit had to be reimaged. A kernel upgrade doesn’t move PCR 7, which I confirmed rather than assumed.
The consequence: firmware updates are report-only in my playbooks and never applied automatically, because one would break unattended unlock on every machine at once.
6. The weekly patch play that skipped its own job
Weekly patching drains a panel, reboots it, waits for Ready, then uncordons. It ran green for a while. It was green because it wasn’t doing the work. Four separate causes:
- The reboot gate checked
/var/run/reboot-required. That file comes fromupdate-notifier-common, which these machines don’t have. The flag was always absent. A kernel could install every week and never be booted. It now compares the running kernel to the newest installed one, and also the loaded NVIDIA module to the one on disk, since a driver update with the old module still loaded gives you “Driver/library version mismatch” until you reboot. - The posture check (GPU alive, lockdown enforcing, LUKS unlock intact) was gated on the reboot. A quiet week verified nothing and reported success. On one run the log said
skipped=7on all eight units. It’s unconditional now. serial: 2plus an unreachable host aborts the batch. On 2026-08-30, two units were off, the play died on the first batch, and nothing was patched, including six healthy units that then sat 55 packages behind for weeks. Now there’s a reachability probe first, patching only runs on units that answered, and the run fails at the end naming the missing ones.- A drain that can’t complete shouldn’t stop the fleet. A blocked drain is recorded, the unit is uncordoned and left alone, the rest proceed, and the final step fails naming every unit it couldn’t reboot.
There’s a fifth, from the NVIDIA module. Replacing Debian’s packages leaves the nouveau blacklist parked as nvidia.conf.dpkg-new, and modprobe only reads files ending in .conf, so nouveau grabs the GPU. The patch play also refuses to reboot into a kernel whose NVIDIA module isn’t signed, because on a Secure Boot machine with lockdown that kills the GPU on every unit at once, silently.
The pattern
Every one of these passed its own check:
- the taint “worked” (nothing broke)
- the join “worked” (the node was Ready)
- the reboot “worked” (the job was green)
- the greeter machine was “healthy” (when I looked)
None of those checks measured the thing I wanted. Read the unit file the installer wrote. Count where the pods run. Compare the running kernel to the installed one. Ask what the check can see in the failure state. This is the same lesson as my failures-that-look-like-success post, and I keep needing it.