Six ways my surplus all-in-one k3s nodes failed quietly
TL;DR I joined eight surplus all-in-ones to my office k3s cluster as workers (the build is in the previous post). Six things went wrong, and all six failed quietly. A scheduling taint left 96 cores running zero workload pods while three small VMs ran 66. Prose inside a YAML folded scalar became k3s command-line arguments, and k3s silently dropped every flag after it. A DHCP lease and a pinned node IP disagreed. The login greeter suspended WiFi-only machines after 20 minutes. One mokutil command broke TPM auto-unlock. A weekly patch play skipped its own job for weeks. The common fix is boring: read back what the system actually did and assert on that, not on the exit code. 1. The taint that idled the fleet My original plan, from a proposal doc, was to taint the panels so nothing but the kiosk would run on them. What shipped was a softer version, PreferNoSchedule, because I also wanted Longhorn on the panels, and Longhorn’s DaemonSets would be locked out by a hard taint. Changing Longhorn’s toleration setting needs every volume detached, which means downtime just to satisfy a taint. ...