TL;DR
- I built a private trip page for a week at a backcountry cabin in a large Ontario park, and nobody ever edited its HTML. Four JSON page specs go through a generator; every change is a typed operation that becomes a self-merging merge request.
- A GitLab CI schedule runs every 3 hours, pulls public Canadian weather, park and fire data, and commits a “conditions: refresh”. It made about 120 of those commits.
- The conditions script is stdlib-only, deterministic and offline-safe. If a source is down it keeps the last good copy, marks it stale, and never invents a number.
- The cabin has no cell service. Power and internet came from a Starlink Mini, a portable power station and a 120 W panel, and the whole workflow assumed a slow, metered link.
- Two real bugs: a bot that pushed as the wrong identity and got a 403, and a “fire” layer that turned out to be a static 2011 zone map.
Why a page at all
A week in the backcountry with a friend means a pile of shared decisions: gear, a meal plan with calorie math, a route with a water taxi and a portage, a charging plan for the drive, and some sense of whether we would be crossing a big lake in a 35 km/h wind. The usual answer is a group chat and a spreadsheet.
This time I wanted one page we could both read on a phone, tick things off on, and comment on. I also wanted an agent to do the boring parts: keep the conditions current, apply edits we asked for, and write the journal while I was out there with a bad connection.
The constraint I set early: nobody edits HTML. Not me, not the agent, not the page. Everything is data in Git.
Four specs, one generator
The page is rendered by a generator on a self-hosted share service. It takes a JSON spec and produces a themed, mobile-friendly page. There are four specs:
| Spec | What it holds |
|---|---|
trip.json | The mid-trip page: conditions, route, power, logistics |
planning.json | Gear checklist, shopping list, calorie math |
journal.json | One entry per day |
gallery.json | Every photo and video |
It started as one page. It got too long to scroll on a phone, so I measured it (screens of scroll at phone and desktop widths), wrote up a short UX review, and split it. Splitting was easy because the generator only cares about the spec, not the HTML.
Typed ops instead of edits
Readers do not change the page directly, and neither does the agent. A change is a typed operation against the spec. The vocabulary is small: set_status, add_row, remove_row, edit_cell, edit_field, add_section, edit_section, move and add_log.
There are two ways to produce one:
- Tick a box or edit a row on the page. The page turns that into an op.
- Ask the in-page assistant. It proposes ops, the reader confirms, and only then do they go anywhere.
The service then applies the ops to the JSON at the current main HEAD, rebuilds, runs the generator’s own checks, opens a merge request on a share/<space>/... branch with a Requested-by trailer, and merges it. Git history ends up being a log of who asked for what. The first day of packing produced dozens of those tiny gear-tick merges, which is exactly what you want from a checklist: boring and attributable.
Comment threads on the page worked the same way. A comment like “the rope is mine, and the tarp is optional since we have the cabin” turned into edits to the gear list. The comment stayed visible, the change landed in Git, and I did not have to remember either.
This is the same shape as the trust ladder I wrote about in the autonomy post: a narrow, typed action surface that is cheap to apply and easy to revert, so self-merge is a reasonable default.
Guardrails that made self-merge safe
- A QA script checks things that are easy to get wrong in a living document: stale gear bars, status words, log ordering, contradictions between sections.
- Agents never publish and never touch comments. One coordinator session does both.
- Only one writer touches the media index, so nothing races.
- A shrink guard refuses to publish a page with fewer sections, tables or log rows than the live one, unless the removal is forced on purpose. An unexpected shrink usually means someone else published first.
- A plain
publishwould strip the reader controls, so the publishing tool refuses it for a live page and only the live variant is used.
The conditions bot
The most useful feature turned out to be the least clever. A script called update_conditions.py rewrites the conditions blocks of the page from public sources:
- Environment Canada (MSC GeoMet city page forecasts for the nearest park station and the nearest town, plus a weather-alerts bounding-box query)
- Ontario Parks (park alerts and the fall colour report)
- ontario.ca, the “restricted fire zones” paragraph
- Ontario’s forest fire danger rating, a point query at the cabin
- Solar equations for civil dawn, sunrise, sunset and dusk, computed locally
It uses the standard library only, so it runs on any runner with no virtualenv drama. The rules matter more than the sources:
- Deterministic. Same inputs, same output. No LLM writes the numbers.
- Never guesses a number. If a field is missing, it says so.
- Offline-safe. Every successful fetch is written to a cache. If a source fails, the page shows the last good copy marked stale rather than blank or fabricated data.
The flags
The flags are plain rules, because on a big lake plain rules are what you want at 6 AM:
# Rule-of-thumb paddling thresholds. 20 km/h is the bottom of Beaufort
# force 4 (whitecaps start); gusts of 35 km/h are well into force 5.
WIND_SUSTAINED_KMH = 20
WIND_GUST_KMH = 35
A day gets a wind flag at 20 km/h sustained or 35 km/h gusts, and the message tells you to cross early and hug the shore. Other flags cover rain, thunder, cold nights (frost below 2 C) and fire. Because the logic is just thresholds, you can read the page and know exactly why it flagged a day.
The schedule
A GitLab pipeline schedule runs the job every 3 hours in the local timezone. The job commits the regenerated section to main as a bot, then republishes. Two details keep it from going wrong:
- The refresh job only runs for
scheduleandwebpipelines. The bot’s own push triggers a normalpushpipeline that runs validation only, so a refresh can never trigger another refresh. - The script exits 0 once the trip is over. GitLab schedules have no end date, so the script supplies one.
The result was roughly 120 conditions: refresh commits on main. That is a noisy history, and I am fine with it. It is a time series of what the forecast said when, in a repo I can diff.
Bug 1: the bot pushed as the wrong person
The first real scheduled run failed with a 403. The token was fine. The identity was wrong.
The CI runner is a shell runner on a Mac. It injects its own credentials into the build checkout (a job token that cannot push to a protected branch), and the system git config on that machine sets a macOS keychain credential helper. Both of those win over GIT_ASKPASS, so git quietly authenticated as something else.
The fix was to stop using the runner’s checkout. The job clones into a fresh temp directory and turns off every other source of config:
export GIT_ASKPASS="$askpass" \
GIT_TERMINAL_PROMPT=0 \
GIT_CONFIG_NOSYSTEM=1 \
GIT_CONFIG_GLOBAL=/dev/null
With system and global config off, the only credential git can find is the bot token supplied by the askpass helper. As a side effect that keeps the token out of URLs, argv and logs. It is a good pattern for any bot that has to push from a machine with its own opinions about credentials, and I covered a related problem in secrets into CI without leaking.
Bug 2: the fire layer that was a 2011 map
I wanted a real fire signal at the cabin. Ontario publishes a GIS layer called Restricted Fire Zone, so I wired it up as “is there a fire ban?”
It is not that. It is a static zone map from 2011. It tells you which zones can be restricted, not whether a restriction is in force today. The page would have shown a confident, wrong “no ban”.
What actually answers the question is a sentence on the ontario.ca forest fires page, under the “Restricted fire zones” heading, plus any park alert that mentions a fire ban. The script now parses that paragraph and fails loudly if the heading or paragraph is missing. The danger rating comes from a separate live layer, queried at the cabin as a point. The lesson is an old one: a layer with the right name is not a layer with the right data. Read the metadata before you build a flag on it.
Power and internet at a cabin with no cell service
The cabin has no cell coverage, so the page’s “live” part ran over a Starlink Mini on a portable power station, with a 120 W panel to refill it. If you want to copy this, the whole kit was a Starlink Mini, a roughly 300 W-class portable power station and a 120 W solar panel. Carried weight: about 14 lb for the battery bag and 8 lb for the panel.
The budget was simple:
- Starlink Mini draws roughly 20 W.
- Plan: off around 10 PM, on at 7 AM, so it is not burning battery while everyone sleeps.
- Fallback: if the battery drops below 30%, go to check-in windows of about 30 minutes morning and evening until a sunny stretch refills it.
Measured numbers beat estimates. In direct sun the battery went from 10% to 100% in about 3 hours. One afternoon read 73 W in against 13 W out with the dish on. The surprise was the evening. The cabin sits below a ridge, and the evening sun never reached the panel, so one evening reading was 85% battery with about 10 hours estimated. Panel placement mattered far more than panel size. The plan had to account for where the sun actually falls, not what the spec sheet promises.
The slow-link workflow
Once I was out there, the rules for the agent changed. The coordinator’s instructions said to assume a slow, metered link: small commits, no screenshots, no browser automation.
The best example is the loop for status messages. I would send something like:
battery 62%, caught 2 bass, windy
and that became one log entry (plus the matching power figure update). One message in, one typed op, one self-merging change, a few kilobytes over the air.
Voice notes I recorded in the field were shaped into journal entries in first person, one section per day, with the day’s photos attached by a separate media job.
The restart handoff
A long-running coordinator session is great until it is not. Before I stopped it, it wrote a restart brief: how to run each script, what had been verified that week, what was in flight, which decisions were still mine, the voice rules for the journal, and which blocks of the page are bot-owned and never hand-edited. When a fresh session read it, there was no “wait, what were we doing”. It is the same idea as the knowledge base that agents write, but scoped to a single operation and dated.
By the numbers at handoff: 8 journal entries, a trip log of 70 entries, and about 660 tests across the QA and script suites.
What I would do again, and what I would not
Again:
- Typed ops and self-merge for anything a human would otherwise edit by hand.
- A deterministic data bot with a stale cache, rather than an agent that “checks the weather”.
- Hard thresholds, written down, that anyone can read.
- Measuring the solar setup instead of trusting the spec sheet.
Not:
- Trusting a layer by its name.
- Letting a runner’s ambient credentials decide who the bot is.
- Leaving the schedule running without an end date baked into the script.
The page did what I wanted: it was boring when everything worked and obviously stale when something did not. If the next project gets a similar treatment, it will be a pantry list or a project tracker, anything where the content is structured and the edits are small.