Debugging quiet failures on a fleet of desktop machines used as cluster nodes Debugging quiet failures on a fleet of desktop machines used as cluster nodes

Six ways my surplus all-in-one k3s nodes failed quietly

TL;DR I joined eight surplus all-in-ones to my office k3s cluster as workers (the build is in the previous post). Six things went wrong, and all six failed quietly. A scheduling taint left 96 cores running zero workload pods while three small VMs ran 66. Prose inside a YAML folded scalar became k3s command-line arguments, and k3s silently dropped every flag after it. A DHCP lease and a pinned node IP disagreed. The login greeter suspended WiFi-only machines after 20 minutes. One mokutil command broke TPM auto-unlock. A weekly patch play skipped its own job for weeks. The common fix is boring: read back what the system actually did and assert on that, not on the exit code. 1. The taint that idled the fleet My original plan, from a proposal doc, was to taint the panels so nothing but the kiosk would run on them. What shipped was a softer version, PreferNoSchedule, because I also wanted Longhorn on the panels, and Longhorn’s DaemonSets would be locked out by a hard taint. Changing Longhorn’s toleration setting needs every volume detached, which means downtime just to satisfy a taint. ...

September 18, 2026 · 8 min · zolty
Surplus all-in-one computers joined to a small k3s cluster Surplus all-in-one computers joined to a small k3s cluster

Nine surplus all-in-ones became k3s nodes for $134 each

TL;DR I bought nine Dell OptiPlex all-in-ones (a mix of 7760 and 7780) in one government surplus lot for $1,209.37 all-in, about $134 a unit. Tax and shipping were zero; I picked them up myself. They run Debian 13 with a LUKS-encrypted root that unlocks through the TPM, Secure Boot on, and a signed NVIDIA module. Getting all three to coexist was most of the work. Eight of them joined my second, small k3s cluster (the “office cluster” at the shop where I do resale and bench work) as workers. Together they add 96 cores and 170 GiB of allocatable memory to a cluster whose three agent VMs had 12 cores and 23 GiB between them. The join is a dedicated Ansible playbook that reuses the existing k3s_agent role. One canary first, then the other seven in two waves. I didn’t measure power draw on these, so there’s no watts figure in this post. The older OptiPlex desktop cluster number doesn’t transfer. Why all-in-ones Most of what I buy for the hobby that funds the hobby is resale stock. This lot was different: nine machines, all wiped, “last powered on six months ago”, sold as-is, pickup by appointment at a school. The seller doesn’t box or ship. The hammer price was $1,075.00, plus a $134.37 buyer premium. I allocated the cost per unit in the inventory database and marked all nine as kept rather than for sale. ...

September 14, 2026 · 7 min · zolty

Affiliate Disclosure: Some links on this site are affiliate links (Amazon Associates, DigitalOcean referral). As an Amazon Associate, I earn from qualifying purchases. This does not affect the price you pay or my editorial independence — I only recommend products and services I personally use and trust.