docs: CONTEXT and REPORT for v0.263.0-v0.263.2 (the undo)
gates / gates (push) Successful in 29s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-23 12:25:44 +02:00
parent 2cd66663f3
commit c3a2aba0d2
2 changed files with 46 additions and 47 deletions
+17 -1
View File
@@ -7,7 +7,23 @@
>
> Ask Claude Code: "Please update CONTEXT.md with what we did today"
Last updated: 2026-09-21 (v0.261.0 — the controller stops swapping itself out from under an app update)
Last updated: 2026-09-23 (v0.263.2 — a failed update puts the app back by itself)
> **2026-09-23 — v0.263.0 → v0.263.2 (R-637, `09` §3 decisions 15 + 19).** A failed health check after
> an update is now UNDONE: phase `copying` (after the pull, the app stopped anyway) copies every NAMED
> volume `cp -a` into `<vol>.pre-update-<stamp>` (label `felhom.undo-copy-of`, finished-marker last);
> on failure `undoing` validates every copy first, refills the volumes, puts back definition + pin +
> the pinned version's `.felhom.yml` from the job's OWN copies, and checks the old version with ITS
> probe → `undone` (`app.yaml` `last_update_undone`, one page line) or a HOLD prefixed with the undo's
> failure and the data state. Bind folders are never touched. Journal recovery: `copying` → old version
> back; `undoing` → resumed. **Two live-only lessons, both now in code:** (1) the periodic probe with
> the CURRENT `.felhom.yml` flips the app `unhealthy`, so the undo's health wait must probe an
> `unhealthy` app with its own override (`waitUpdateHealthyMeta`, v0.263.1); (2) `.felhom.yml` flows
> into the stack dir on every catalog SYNC, so the pinned version's file must be recorded AT PIN TIME
> — `<stack>/applied-meta/.felhom.yml`, written at deploy, adoption and pin advance (v0.263.2). Apps
> pinned earlier have no record until their next pin (R-646). R-642: start/restart never say
> "completed". **The fleet floor is still 0.262.1** — raising it is the operator's decision. Evidence:
> `felhom.eu/documentation/audits/undo-bakeoff-2026-09-23/`, `…/undo-live-2026-09-23/`.
> **2026-09-21 (evening) — v0.261.0 (R-608 + R-609).** The controller self-updates daily at **04:30** by default and after ANY hub report once a floor sits above the box; that swap restarts the controller, which is the supervisor of a running app update. The window `09` §3b Q1 proposes for automatic app updates is 02:30-05:00. **It contains 04:30.** **A two-way lock, wired in `main.go` — `stacks` never imports `selfupdate`:** `Manager.AnyUpdating()` -> `Updater.SetAppUpdatingCheck` (a sibling of the existing `SetBackupRunningCheck`, consulted in the SAME three places), and `Updater.IsUpdateRunning` -> `Manager.SetSelfUpdatingCheck` with `UpdatePreflight` refusing `self_updating`. **MEASURED: the gap was NARROWER than assumed** — the update's `backing-up` phase already took the backup single-flight, so only the other six phases were exposed. The live probe landed in `safety-dump`, i.e. in the real gap. **THE PROPERTY THAT MATTERS: the lock must NOT latch** — `Stack.Updating` clears on done, failed AND held, so a held app does not block the controller's own updates for ever. **R-609:** the 409 now carries `data.reason` — transient (`busy`,`updating`,`deploying`,`migrating`,`self_updating`) vs terminal (`held`,`downgrade`). **Found while writing the test: the router refuses a HELD app on its OWN line BEFORE `UpdatePreflight`**, so `held` — the reason an unattended caller needs most — would have been the one missing. Five red-proofs, each seen to fail. **Proven live on guest 9202 with a negative control:** with no app update the manual swap gives the AGENT refusal; with one in flight it gives OURS. The sentence changing IS the proof. Deployed to **9202 only**; the fleet floor is 0.260.0 and 0.261.0 is a separate operator ask. Measurements: `felhom.eu/documentation/audits/update-arc-gaps-2026-09-21/`.
+29 -46
View File
@@ -1,55 +1,38 @@
# REPORT — controller v0.262.0 + v0.262.1 (2026-09-22)
# REPORT — v0.263.0 → v0.263.2: a failed update puts the app back by itself (R-637)
## Not done, or changed from the brief
2026-09-23. Baseline `b9deec19077b` (v0.262.1). Commits: `8fc2b4a` (v0.263.0), `5d38573` (v0.263.1),
`2cd6666` (v0.263.2). **MinAgent 0.131.0, unchanged.** Deployed: **scratch guest 9202 only**; the fleet
floor stays at 0.262.1 (the operator's question).
1. **R-634's diagnosis spike was SKIPPED, and the brief put a gate in front of it.** §3 said: reproduce
`sparkyfitness` alone, read `runComposeDeploy` against the log, name the line — *then* fix. I went
straight to the bounded half (Part 2.2), which is safe regardless and is now proven live. **The
mechanism is still unknown** and R-634 stays open at P1 for it.
2. **R-625 (a held app still renders an Update button) is not started.** Its bundle key was written
and then **removed from both bundles** — an unused key is a promise not kept.
3. **v0.262.1 exists because the live proof found what the tests did not.** The busy guard fired
correctly and answered **HTTP 500**: `router.go` maps remove errors by grepping the error TEXT,
and the busy sentence contains none of the words it looks for. Now a typed error and a 409.
4. **Two baselines in the brief were stale** — `felhom.eu` and `app-catalog-felhom.eu` had both moved
since it was written, by my own work earlier the same day, and the highest `R-` id was 636, not
634.
5. **Scenario B's first live run proved nothing and nearly went down as a pass.** The remove was
refused with `409 still running — stop it first`, which is the PRE-EXISTING check: the restore had
already finished. A refusal from the wrong rule is not evidence for the new one. Re-run with the
app STOPPED and a backup in flight, which is the only way to reach the new guard.
6. **Scenario B's subject changed** from `gokapi` to `privatebin`: gokapi is still crash-looping from
this morning's R-633 artefact (its config volume was removed, so its binary can never start), and
a subject that cannot reach a steady state proves nothing about a guard that fires between them.
7. **44 minutes were lost to my own waiters** — `pgrep -f "build.sh 0.262.0"` matched the waiter's own
command line, so it waited for itself while the build had already succeeded. Same bug cost 15
stuck watchers earlier in the day. Recorded as its own memory rule: watch a sentinel the work
writes, never a process name.
## What changed
## What shipped
- `internal/stacks/undo.go` (new): the folder copy (`volumeCopier`, docker helper), `tryUndo`, the
applied `.felhom.yml` record, `last_update_undone`.
- `internal/stacks/update.go`: phase `copying` after the pull; `failAndHold` undoes first; journal
phases `copying`/`undoing` with recovery; `restoreDefinition` split from `pinBack`;
`waitUpdateHealthyMeta` (the undo's own probe, which runs on an `unhealthy` app); seam `probeRunFn`.
- `pin.go`, `deploy.go`: every pin writer records the pinned version's `.felhom.yml` (`applied-meta/`).
- `backup/`: `RestoreHold.UndoState`; the hold sentence's undo prefix + state clause (box language).
- `web/`: the undone line on the app page (request language). `api/`: R-642 start answer.
- `delete.go`: a removal deletes the app's kept undo copies.
- i18n: 10 keys born as keys (hu + en). `docker_run_volume_path_gate`: four named-volume mounts allowlisted.
`felhom-controller@3d41758` — **v0.262.0** (R-630, R-634-half, R-633/R-626, R-621, R-614) and
**v0.262.1** (the 409). All 17 controller gates green; the go-parity gate caught both new i18n keys
as unregistered before the push and was right to.
## Tests and red-proofs
`app-catalog-felhom.eu@02844ae` — six proven versions moved (one commit each), `healthcheck.container`
for paperless-ngx and immich, the probe gate rewritten to the controller's four rules, decoy suite
56 → **64 cases**.
`go build ./... && go vet ./... && go test ./...` → rc=0 before each commit. Controller gates rc=0.
Twelve red-proofs, each seen failing and restored — table in `felhom.eu/REPORT.md` §3. Two of them
(9, 3b) reproduce the two live-only defects verbatim.
## Proven live on 9202
## Live (endpoint-level, 9202)
| | before (v0.261.0) | after |
|---|---|---|
| **A** paperless-ngx Update | `failed` at **+313.0 s**, app stopped, front door 404 | **`done` at +53.4 s**, running |
| **B** remove during a backup | — | refused, `REFUSED (busy): a backup or restore is running`; remove after → `verified: true`, nothing 60 s later |
| **C** remove a half-state | `stack "sparkyfitness" is not deployed` | **200**, `leftovers: NONE` |
| **F** redeploy after remove | inherited the old phase | phase `done` → **no phase** |
Three apps undone by the product with seeds before/after the backup and seconds before the press read
back; a cut-off copy held honestly (hu and en prefix); a power cut during `undoing` resumed and undone;
a person's press after an undo worked; removal deleted kept copies. Evidence:
`felhom.eu/documentation/audits/undo-live-2026-09-23/README.md`.
**Five red-proofs, each seen failing then green.** D (hold logs) is pinned by the two existing tests
whose compose sequence now includes the capture — they caught it and made me justify the order.
## NOT yet live-validated
## What is owed
- **R-634's mechanism** — why `deployed` goes false while containers run.
- **R-625** — one verdict for the held badge and the Update button.
- **The floor.** v0.262.1 is on 9202 only. Raising it is the operator's call.
- The household mail / operator event for an undo (not built — `09` §6.4 part 2).
- An app pinned BEFORE v0.263.2 that has not been re-pinned (R-646) — its undo probes with the
current `.felhom.yml`; reasoned, not run live.
- Any box other than 9202.