Files
felhom-controller/REPORT.md
T
admin 93cee16843
gates / gates (push) Successful in 25s
CONTEXT + REPORT for v0.261.0 (R-608, R-609)
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
2026-09-21 15:00:29 +02:00

5.6 KiB
Raw Blame History

REPORT — controller v0.261.0: the box stops updating itself out from under an app update

R-608 (P2) + R-609 (P3). Base d0d431b42b51 (v0.260.0) → v0.261.0 (811f75736ec8). MinAgent 0.131.0 unchanged. Architecture read first and named: felhom.eu/documentation/architecture/09-update-architecture.md §3 (the ten decisions), §3b (the seven open questions), §6.1 (the guarded update's phases), §6.2.

Not done / changed from the brief — first, because it is the point

item state
Part 2 (the lock, the reason on the wire, the red-proofs) done in full
Part 0 (floor to 0.260.0) done — see felhom.eu/REPORT.md
Parts 1, 3, 4 not in this repo — they are measurements and documents; see felhom.eu
the caller script's ≤150-line budget 167 lines. Over by 17. Not trimmed: the excess is the same_major rule and its comment, and shortening it would have meant a terser rule, which is the one thing in that file that must be readable. Named rather than hidden.

What was wrong

The controller updates itself — daily at self_update.auto_update_time, default 04:30 (config/config.go L422, scheduled cmd/controller/main.go ~L1365), and again from MaybeAutoUpdate after any hub report once a floor sits above the box, so at any hour. That swap restarts the controller container, which is the supervisor of a running app update. 09 §3b Q1 proposes 02:30–05:00 for automatic app updates. It contains 04:30.

What the brief got wrong, measured rather than assumed

The gap was narrower than "the whole update", and that matters for where the fix goes. The brief said the self-updater's only busy gate is backupRunning and left open whether that covers the update's backing-up phase. It does: RunAppBackupNow calls acquireRunning (internal/backup/update_guard.go:333), so backupMgr.IsRunning() was already true for that one phase. It was false for checking, safety-dump, pinning, pulling, starting and verifying — and the last two are exactly where the new version may already have touched the customer's data. Everything else the brief asserted about the three call sites and the wiring held at source.

What shipped

  • stacks.Manager.AnyUpdating() — is a guarded update in flight for ANY app.
  • Updater.SetAppUpdatingCheck — a deliberate sibling of SetBackupRunningCheck, consulted in the same three places (the dry run, TriggerUpdate, maybeAutoUpdate). One busy-gate pattern in that file, not two.
  • Manager.SetSelfUpdatingCheck — the reverse. UpdatePreflight refuses self_updating.
  • Both halves wired in main.go, the only place holding both objects. stacks never imports selfupdate — the dependency is inverted with a plain callback rather than by widening UpdateGuards, which is the backup side's interface and has nothing to do with this.
  • Two sentences, born as bundle keys, in both bundles, registered in i18n_go_keys.json.
  • data.reason on every update refusal (R-609) — additive; the sentence is unchanged, so no page moves. busy/updating/deploying/migrating/self_updating are transient; held/downgrade terminal; memory/disk/no_backup need a person.

The two things the tests found that reading did not

  1. The router refuses a HELD app on its own line, BEFORE UpdatePreflight (api/router.go ~L601). Without a second edit, held — the single reason an unattended caller most needs — would have been the one missing from the wire, and such a caller would press a terminally-refused button on every pass for ever. Found while writing the table, not while reading the code.
  2. The lock must not latch. Stack.Updating is cleared on done, failed and held, so a held app does not block the controller's own updates — including the release that might fix whatever held it. A latching gate would be a worse failure than the one prevented, and silent for weeks. TestR608_LockReleasesAfterHold is a consequence test and exists for that alone.

Red-proofs — five, each SEEN to fail

# the mutation what failed
1 AnyUpdating always false TestR608_AnyUpdatingSeesAnUpdateInFlight — the gate reads false with an update in flight
2 AnyUpdating true for a held app too TestR608_LockReleasesAfterHold — "a HELD app must not hold the self-update lock for ever"
3 delete the preflight's self_updating block TestR608_PreflightRefuses… — "while the controller swaps itself the app update must be REFUSED"
4 delete TriggerUpdate's app-update block TestR608_TriggerUpdateRefused… — the swap proceeds over a live app update
5 drop Data from the refusal TestR609_EveryRefusalCarriesItsReason — reason = "" on no_backup and self_updating

Green gate

go build ./... && go vet ./... && go test ./... — all green, full suite. controller_gates.py --fast — 17 gates OK; go-parity convicted the two new keys first and was satisfied properly by registering both as BORN-AS-KEYS with the test that pins each. No --no-verify; the pre-push hook ran and passed.

Image gitea.dooplex.hu/admin/felhom-controller:0.261.0 built and pushed. Deployed to guest 9202 (the scratch guest) ONLY — the demo boxes and the fleet stay on 0.260.0. The floor is NOT raised to 0.261.0; that is a separate ask for the operator, recorded in felhom.eu/STATUS.md.

The measurements this release was built for — the three power cuts and the unattended night — are in felhom.eu/documentation/audits/update-arc-gaps-2026-09-21/.