TL;DR

  • The screens on eight surplus all-in-ones show an ops board: bids, ship queue, pickups, repair backlog, cluster health, weather and traffic. They sit on desks around the office and do double duty: a workstation when someone sits down, the board the rest of the time. The same machines are on their way to becoming k3s worker nodes (one has joined so far), and that build gets its own post.
  • It’s one static HTML page in Chrome kiosk mode. No framework server, no Kubernetes workload. Ansible deploys it.
  • 15 scenes sit in a weighted registry and rotate off the wall clock, about 34 seconds each, so eight screens agree without talking to each other.
  • When a fetch fails the page keeps the last good data and shows its age. A board with stale data and a warning beats a blank screen.
  • It’s a different animal from my older home signage system, which is a k8s app with Raspberry Pi kiosks.
  • One real thing it did: on 2026-09-06 the home cluster’s daily Longhorn backup had stopped completing, and the board’s home-cluster card showed the critical alert.

Not the same project as the home signage

I’ve written about digital signage on k3s before: an Angular app, a handful of Flask services, MQTT, Raspberry Pis driving displays. That’s a proper application and it lives in the home cluster.

This is the opposite design. The repo still calls it the “video wall”, but there’s no wall: the screens are the all-in-ones on the office desks, and each one is somebody’s workstation as well as a panel of the board. The office cluster, a small second k3s cluster at the shop where I do resale and bench work, has no signage app in its manifests at all. The board is a single page, index.html, about 1,400 lines, and a handful of helper scripts. Its “deployment” is an Ansible playbook that copies files onto each machine and starts a browser.

I think this is the right call for a board like this. An unattended display’s job is to still work when everything else is broken, and the less it depends on, the more often that’s true.

The data path

The data comes from my flipping tracker’s /api/signage/board endpoint. That endpoint sits behind SSO forward-auth, and a kiosk browser can’t attach an Authorization header. So each panel does the fetching itself:

  1. A systemd --user timer runs fetch-board.sh every minute.
  2. It calls the endpoint with curl, reading the bearer token from a 0600 file and handing it to curl through --config so the token never appears on a command line (argv is world-readable through /proc).
  3. It writes board.json next to the page. Weather comes from Open-Meteo, which needs no API key, and is fetched by the script too. A weather outage can’t blank the board.
  4. A loopback-only static server serves the directory. Chrome runs --kiosk --incognito against that local URL.

Same-origin through loopback means no CORS to relax. That’s the whole reason for the indirection.

Before this page existed, the panels played a live stream through mpv. I’d tried Chrome kiosk on a video embed first and got a refusal, because embeds won’t load as the top-level page. The two modes are mutually exclusive in the playbook, selected by a variable.

One rule: keep the last good board

The fetch script’s comment says it best, so I’ll quote it: on failure the previous board.json stays in place, because the page renders the last known state and flags its age, “which is far more useful on a wall than a blank screen every time the network hiccups.”

On the page there’s a freshness chip. It turns amber past a stale threshold and red past a dead one. Anyone glancing at the wall can tell “the data is old” from “the world is fine”. That distinction is what you actually want from an unattended display.

The other rules in the same spirit:

  • A scene with no data is dropped, not shown with a caption saying there’s nothing to show.
  • For traffic, the delta is the scene: the interesting fact is how much worse than normal, not the minutes.
  • Only critical cluster alerts escalate. My home cluster routinely carries a pile of high-severity Longhorn alerts that are just weather. Escalating them would train everyone to ignore the wall.

15 scenes, weighted

The page has a registry of scenes, each with a render function and a weight:

GroupScenes
Resale flowbids, ship queue, pickups queued, inventory, KPIs
Benchrepair backlog, repair ROI
Infrastructurealerts, home cluster health, office cluster health, changelog of merged work
The outside worldtraffic deltas, weather radar, local news, a nearby-incidents feed

Most scenes weigh 1. The incidents, news and changelog scenes weigh half, because I don’t need a long look at those. Alerts weigh 2, and 4 when something is actually failing, so a problem gets more airtime exactly when there’s a problem.

The persistent chrome around the scene stays on screen the whole time: clock, date, a weather card, an item count, the freshness chip, and a backdrop made from photos of what’s currently for sale.

Lockstep from the wall clock

Eight screens that rotate on local timers drift apart. They’d be on different scenes within an hour. I wanted them in agreement without a server telling them what to do.

So the scheduler is a pure function of Date.now() and the registry’s weights. Nominal dwell is 34 seconds. After 25 nominal slots, every panel rests to black for 20 seconds, and the whole cycle repeats. Given the same weights, every panel computes the same answer at the same instant.

const SCENE_MS   = 34000;
const REST_EVERY = 25;
const REST_MS    = 20000;
const PERIOD_MS  = SCENE_MS * REST_EVERY;
const WINDOW_MS  = PERIOD_MS - REST_MS;
// schedule(weights, now) maps the wall clock onto a weighted timeline

Two details I got wrong the first time, both recorded as comments in the source:

  • The first version indexed by floor(now / SCENE_MS) % scenes. That makes airtime a function of a scene’s position in the list, not its importance. The alerts scene sat at index 0 behind an always-populated news scene and got one slot no matter what.
  • The rest started 14 seconds into a slot (830,000 divided by 34,000 isn’t a whole number), so whichever scene landed there was cut short every cycle.

The fix lays the timeline over the whole registry, not just the scenes a given panel has data for, and fits it to a whole number of passes so the rest always starts on a boundary. A panel with no radar frame yet walks forward to its next scene for that slot only. Everyone else still agrees.

The weight function is a function, so alerts can check whether anything is failing. The scheduler has its own CI test, and a schema check fails the pipeline if a scene’s builder exists without being registered, or the other way round.

Deploying to eight desks is its own testing problem

Two lessons that I’d have missed if I’d only watched green jobs:

  • A deploy can succeed and change nothing on screen. The kiosk never reloaded a changed page. The playbook finished, the file was new, and the panel kept showing the old one. There’s now a CI guard that checks the kiosk restart path.
  • Never sample at 34 seconds. A verification script that grabs a screenshot every 34 seconds aliases against the dwell and keeps seeing the same scene. Pick a sampling interval that doesn’t divide the dwell.

The one real thing it did

I’m not going to invent a pile of anecdotes. I have one that holds up.

On 2026-09-06 the home cluster’s daily Longhorn backup to S3 stopped completing. The last good run was about 44 hours old. DailyBackupStale fired critical, and because the board’s home-cluster card pulls from Alertmanager, the critical alert escalated onto it, no dashboard login required.

The cause turned out to be interesting. Longhorn takes a per-volume lock in the backup store and normally removes it when it’s done. A process killed mid-backup leaves its lock behind and nothing reclaims it. The backup run treats a lock timeout as fatal, so a single orphan aborts the whole 80-volume run, which restarts from zero and fails the same way. Four orphaned locks, most likely left by pods being killed during a control-plane consolidation the day before, had stopped the daily backup entirely.

Deleting the four locks let the next run finish in 23 minutes. I added a small CronJob that prunes stale locks fifteen minutes before the scheduled backup.

What the wall did is narrower. It put a critical alert in front of everyone in the room without anyone opening a tab. My handoff notes from that day say the board “reported it correctly”. I wouldn’t claim more than that.