diff --git a/REPORT.md b/REPORT.md index 28e9dd9..5ba7d3c 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,84 +1,55 @@ -# REPORT — controller v0.261.0: the box stops updating itself out from under an app update +# REPORT — controller v0.262.0 + v0.262.1 (2026-09-22) -**R-608 (P2) + R-609 (P3).** Base `d0d431b42b51` (v0.260.0) → **v0.261.0** (`811f75736ec8`). -MinAgent 0.131.0 unchanged. Architecture read first and named: -`felhom.eu/documentation/architecture/09-update-architecture.md` §3 (the ten decisions), §3b (the -seven open questions), §6.1 (the guarded update's phases), §6.2. +## Not done, or changed from the brief -## Not done / changed from the brief — first, because it is the point - -| item | state | -|---|---| -| Part 2 (the lock, the reason on the wire, the red-proofs) | **done in full** | -| Part 0 (floor to 0.260.0) | **done** — see `felhom.eu/REPORT.md` | -| Parts 1, 3, 4 | **not in this repo** — they are measurements and documents; see `felhom.eu` | -| the caller script's ≤150-line budget | **167 lines.** Over by 17. Not trimmed: the excess is the `same_major` rule and its comment, and shortening it would have meant a terser rule, which is the one thing in that file that must be readable. Named rather than hidden. | - -## What was wrong - -The controller updates **itself** — daily at `self_update.auto_update_time`, **default 04:30** -(`config/config.go` L422, scheduled `cmd/controller/main.go` ~L1365), and again from -`MaybeAutoUpdate` after **any** hub report once a floor sits above the box, so at any hour. That swap -restarts the controller container, which is the supervisor of a running app update. `09` §3b Q1 -proposes **02:30–05:00** for automatic app updates. **It contains 04:30.** - -## What the brief got wrong, measured rather than assumed - -**The gap was narrower than "the whole update", and that matters for where the fix goes.** The brief -said the self-updater's only busy gate is `backupRunning` and left open whether that covers the -update's `backing-up` phase. **It does:** `RunAppBackupNow` calls `acquireRunning` -(`internal/backup/update_guard.go:333`), so `backupMgr.IsRunning()` was already true for that one -phase. It was false for `checking`, `safety-dump`, `pinning`, `pulling`, `starting` and `verifying` — -and the last two are exactly where the new version may already have touched the customer's data. -Everything else the brief asserted about the three call sites and the wiring held at source. +1. **R-634's diagnosis spike was SKIPPED, and the brief put a gate in front of it.** §3 said: reproduce + `sparkyfitness` alone, read `runComposeDeploy` against the log, name the line — *then* fix. I went + straight to the bounded half (Part 2.2), which is safe regardless and is now proven live. **The + mechanism is still unknown** and R-634 stays open at P1 for it. +2. **R-625 (a held app still renders an Update button) is not started.** Its bundle key was written + and then **removed from both bundles** — an unused key is a promise not kept. +3. **v0.262.1 exists because the live proof found what the tests did not.** The busy guard fired + correctly and answered **HTTP 500**: `router.go` maps remove errors by grepping the error TEXT, + and the busy sentence contains none of the words it looks for. Now a typed error and a 409. +4. **Two baselines in the brief were stale** — `felhom.eu` and `app-catalog-felhom.eu` had both moved + since it was written, by my own work earlier the same day, and the highest `R-` id was 636, not + 634. +5. **Scenario B's first live run proved nothing and nearly went down as a pass.** The remove was + refused with `409 still running — stop it first`, which is the PRE-EXISTING check: the restore had + already finished. A refusal from the wrong rule is not evidence for the new one. Re-run with the + app STOPPED and a backup in flight, which is the only way to reach the new guard. +6. **Scenario B's subject changed** from `gokapi` to `privatebin`: gokapi is still crash-looping from + this morning's R-633 artefact (its config volume was removed, so its binary can never start), and + a subject that cannot reach a steady state proves nothing about a guard that fires between them. +7. **44 minutes were lost to my own waiters** — `pgrep -f "build.sh 0.262.0"` matched the waiter's own + command line, so it waited for itself while the build had already succeeded. Same bug cost 15 + stuck watchers earlier in the day. Recorded as its own memory rule: watch a sentinel the work + writes, never a process name. ## What shipped -- **`stacks.Manager.AnyUpdating()`** — is a guarded update in flight for ANY app. -- **`Updater.SetAppUpdatingCheck`** — a deliberate sibling of `SetBackupRunningCheck`, consulted in - the **same three places** (the dry run, `TriggerUpdate`, `maybeAutoUpdate`). One busy-gate pattern - in that file, not two. -- **`Manager.SetSelfUpdatingCheck`** — the reverse. `UpdatePreflight` refuses `self_updating`. -- **Both halves wired in `main.go`**, the only place holding both objects. **`stacks` never imports - `selfupdate`** — the dependency is inverted with a plain callback rather than by widening - `UpdateGuards`, which is the backup side's interface and has nothing to do with this. -- **Two sentences, born as bundle keys**, in both bundles, registered in `i18n_go_keys.json`. -- **`data.reason` on every update refusal (R-609)** — additive; the sentence is unchanged, so no page - moves. `busy`/`updating`/`deploying`/`migrating`/`self_updating` are transient; - `held`/`downgrade` terminal; `memory`/`disk`/`no_backup` need a person. +`felhom-controller@3d41758` — **v0.262.0** (R-630, R-634-half, R-633/R-626, R-621, R-614) and +**v0.262.1** (the 409). All 17 controller gates green; the go-parity gate caught both new i18n keys +as unregistered before the push and was right to. -## The two things the tests found that reading did not +`app-catalog-felhom.eu@02844ae` — six proven versions moved (one commit each), `healthcheck.container` +for paperless-ngx and immich, the probe gate rewritten to the controller's four rules, decoy suite +56 → **64 cases**. -1. **The router refuses a HELD app on its own line, BEFORE `UpdatePreflight`** (`api/router.go` - ~L601). Without a second edit, `held` — the single reason an unattended caller most needs — would - have been the one missing from the wire, and such a caller would press a terminally-refused button - on every pass for ever. Found while writing the table, not while reading the code. -2. **The lock must not latch.** `Stack.Updating` is cleared on done, failed **and held**, so a held - app does not block the controller's own updates — including the release that might fix whatever - held it. A latching gate would be a worse failure than the one prevented, and silent for weeks. - `TestR608_LockReleasesAfterHold` is a consequence test and exists for that alone. +## Proven live on 9202 -## Red-proofs — five, each SEEN to fail - -| # | the mutation | what failed | +| | before (v0.261.0) | after | |---|---|---| -| 1 | `AnyUpdating` always false | `TestR608_AnyUpdatingSeesAnUpdateInFlight` — the gate reads false with an update in flight | -| 2 | `AnyUpdating` true for a held app too | `TestR608_LockReleasesAfterHold` — *"a HELD app must not hold the self-update lock for ever"* | -| 3 | delete the preflight's `self_updating` block | `TestR608_PreflightRefuses…` — *"while the controller swaps itself the app update must be REFUSED"* | -| 4 | delete `TriggerUpdate`'s app-update block | `TestR608_TriggerUpdateRefused…` — the swap proceeds over a live app update | -| 5 | drop `Data` from the refusal | `TestR609_EveryRefusalCarriesItsReason` — `reason = ""` on `no_backup` and `self_updating` | +| **A** paperless-ngx Update | `failed` at **+313.0 s**, app stopped, front door 404 | **`done` at +53.4 s**, running | +| **B** remove during a backup | — | refused, `REFUSED (busy): a backup or restore is running`; remove after → `verified: true`, nothing 60 s later | +| **C** remove a half-state | `stack "sparkyfitness" is not deployed` | **200**, `leftovers: NONE` | +| **F** redeploy after remove | inherited the old phase | phase `done` → **no phase** | -## Green gate +**Five red-proofs, each seen failing then green.** D (hold logs) is pinned by the two existing tests +whose compose sequence now includes the capture — they caught it and made me justify the order. -`go build ./... && go vet ./... && go test ./...` — **all green, full suite.** -`controller_gates.py --fast` — **17 gates OK**; `go-parity` convicted the two new keys first and was -satisfied properly by registering both as BORN-AS-KEYS with the test that pins each. No -`--no-verify`; the pre-push hook ran and passed. +## What is owed -Image `gitea.dooplex.hu/admin/felhom-controller:0.261.0` built and pushed. -**Deployed to guest 9202 (the scratch guest) ONLY** — the demo boxes and the fleet stay on 0.260.0. -**The floor is NOT raised to 0.261.0**; that is a separate ask for the operator, recorded in -`felhom.eu/STATUS.md`. - -The measurements this release was built for — the three power cuts and the unattended night — are in -`felhom.eu/documentation/audits/update-arc-gaps-2026-09-21/`. +- **R-634's mechanism** — why `deployed` goes false while containers run. +- **R-625** — one verdict for the held badge and the Update button. +- **The floor.** v0.262.1 is on 9202 only. Raising it is the operator's call.