Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
+17
-1
@@ -7,7 +7,23 @@
|
||||
>
|
||||
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
|
||||
|
||||
Last updated: 2026-09-21 (v0.261.0 — the controller stops swapping itself out from under an app update)
|
||||
Last updated: 2026-09-23 (v0.263.2 — a failed update puts the app back by itself)
|
||||
|
||||
> **2026-09-23 — v0.263.0 → v0.263.2 (R-637, `09` §3 decisions 15 + 19).** A failed health check after
|
||||
> an update is now UNDONE: phase `copying` (after the pull, the app stopped anyway) copies every NAMED
|
||||
> volume `cp -a` into `<vol>.pre-update-<stamp>` (label `felhom.undo-copy-of`, finished-marker last);
|
||||
> on failure `undoing` validates every copy first, refills the volumes, puts back definition + pin +
|
||||
> the pinned version's `.felhom.yml` from the job's OWN copies, and checks the old version with ITS
|
||||
> probe → `undone` (`app.yaml` `last_update_undone`, one page line) or a HOLD prefixed with the undo's
|
||||
> failure and the data state. Bind folders are never touched. Journal recovery: `copying` → old version
|
||||
> back; `undoing` → resumed. **Two live-only lessons, both now in code:** (1) the periodic probe with
|
||||
> the CURRENT `.felhom.yml` flips the app `unhealthy`, so the undo's health wait must probe an
|
||||
> `unhealthy` app with its own override (`waitUpdateHealthyMeta`, v0.263.1); (2) `.felhom.yml` flows
|
||||
> into the stack dir on every catalog SYNC, so the pinned version's file must be recorded AT PIN TIME
|
||||
> — `<stack>/applied-meta/.felhom.yml`, written at deploy, adoption and pin advance (v0.263.2). Apps
|
||||
> pinned earlier have no record until their next pin (R-646). R-642: start/restart never say
|
||||
> "completed". **The fleet floor is still 0.262.1** — raising it is the operator's decision. Evidence:
|
||||
> `felhom.eu/documentation/audits/undo-bakeoff-2026-09-23/`, `…/undo-live-2026-09-23/`.
|
||||
|
||||
> **2026-09-21 (evening) — v0.261.0 (R-608 + R-609).** The controller self-updates daily at **04:30** by default and after ANY hub report once a floor sits above the box; that swap restarts the controller, which is the supervisor of a running app update. The window `09` §3b Q1 proposes for automatic app updates is 02:30-05:00. **It contains 04:30.** **A two-way lock, wired in `main.go` — `stacks` never imports `selfupdate`:** `Manager.AnyUpdating()` -> `Updater.SetAppUpdatingCheck` (a sibling of the existing `SetBackupRunningCheck`, consulted in the SAME three places), and `Updater.IsUpdateRunning` -> `Manager.SetSelfUpdatingCheck` with `UpdatePreflight` refusing `self_updating`. **MEASURED: the gap was NARROWER than assumed** — the update's `backing-up` phase already took the backup single-flight, so only the other six phases were exposed. The live probe landed in `safety-dump`, i.e. in the real gap. **THE PROPERTY THAT MATTERS: the lock must NOT latch** — `Stack.Updating` clears on done, failed AND held, so a held app does not block the controller's own updates for ever. **R-609:** the 409 now carries `data.reason` — transient (`busy`,`updating`,`deploying`,`migrating`,`self_updating`) vs terminal (`held`,`downgrade`). **Found while writing the test: the router refuses a HELD app on its OWN line BEFORE `UpdatePreflight`**, so `held` — the reason an unattended caller needs most — would have been the one missing. Five red-proofs, each seen to fail. **Proven live on guest 9202 with a negative control:** with no app update the manual swap gives the AGENT refusal; with one in flight it gives OURS. The sentence changing IS the proof. Deployed to **9202 only**; the fleet floor is 0.260.0 and 0.261.0 is a separate operator ask. Measurements: `felhom.eu/documentation/audits/update-arc-gaps-2026-09-21/`.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user