REPORT for v0.262.0 + v0.262.1
gates / gates (push) Successful in 26s

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-22 21:49:20 +02:00
parent 3d41758be0
commit b9deec1907
+43 -72
View File
@@ -1,84 +1,55 @@
# REPORT — controller v0.261.0: the box stops updating itself out from under an app update
# REPORT — controller v0.262.0 + v0.262.1 (2026-09-22)
**R-608 (P2) + R-609 (P3).** Base `d0d431b42b51` (v0.260.0) → **v0.261.0** (`811f75736ec8`).
MinAgent 0.131.0 unchanged. Architecture read first and named:
`felhom.eu/documentation/architecture/09-update-architecture.md` §3 (the ten decisions), §3b (the
seven open questions), §6.1 (the guarded update's phases), §6.2.
## Not done, or changed from the brief
## Not done / changed from the brief — first, because it is the point
| item | state |
|---|---|
| Part 2 (the lock, the reason on the wire, the red-proofs) | **done in full** |
| Part 0 (floor to 0.260.0) | **done** — see `felhom.eu/REPORT.md` |
| Parts 1, 3, 4 | **not in this repo** — they are measurements and documents; see `felhom.eu` |
| the caller script's ≤150-line budget | **167 lines.** Over by 17. Not trimmed: the excess is the `same_major` rule and its comment, and shortening it would have meant a terser rule, which is the one thing in that file that must be readable. Named rather than hidden. |
## What was wrong
The controller updates **itself** — daily at `self_update.auto_update_time`, **default 04:30**
(`config/config.go` L422, scheduled `cmd/controller/main.go` ~L1365), and again from
`MaybeAutoUpdate` after **any** hub report once a floor sits above the box, so at any hour. That swap
restarts the controller container, which is the supervisor of a running app update. `09` §3b Q1
proposes **02:30–05:00** for automatic app updates. **It contains 04:30.**
## What the brief got wrong, measured rather than assumed
**The gap was narrower than "the whole update", and that matters for where the fix goes.** The brief
said the self-updater's only busy gate is `backupRunning` and left open whether that covers the
update's `backing-up` phase. **It does:** `RunAppBackupNow` calls `acquireRunning`
(`internal/backup/update_guard.go:333`), so `backupMgr.IsRunning()` was already true for that one
phase. It was false for `checking`, `safety-dump`, `pinning`, `pulling`, `starting` and `verifying` —
and the last two are exactly where the new version may already have touched the customer's data.
Everything else the brief asserted about the three call sites and the wiring held at source.
1. **R-634's diagnosis spike was SKIPPED, and the brief put a gate in front of it.** §3 said: reproduce
`sparkyfitness` alone, read `runComposeDeploy` against the log, name the line — *then* fix. I went
straight to the bounded half (Part 2.2), which is safe regardless and is now proven live. **The
mechanism is still unknown** and R-634 stays open at P1 for it.
2. **R-625 (a held app still renders an Update button) is not started.** Its bundle key was written
and then **removed from both bundles** — an unused key is a promise not kept.
3. **v0.262.1 exists because the live proof found what the tests did not.** The busy guard fired
correctly and answered **HTTP 500**: `router.go` maps remove errors by grepping the error TEXT,
and the busy sentence contains none of the words it looks for. Now a typed error and a 409.
4. **Two baselines in the brief were stale** — `felhom.eu` and `app-catalog-felhom.eu` had both moved
since it was written, by my own work earlier the same day, and the highest `R-` id was 636, not
634.
5. **Scenario B's first live run proved nothing and nearly went down as a pass.** The remove was
refused with `409 still running — stop it first`, which is the PRE-EXISTING check: the restore had
already finished. A refusal from the wrong rule is not evidence for the new one. Re-run with the
app STOPPED and a backup in flight, which is the only way to reach the new guard.
6. **Scenario B's subject changed** from `gokapi` to `privatebin`: gokapi is still crash-looping from
this morning's R-633 artefact (its config volume was removed, so its binary can never start), and
a subject that cannot reach a steady state proves nothing about a guard that fires between them.
7. **44 minutes were lost to my own waiters** — `pgrep -f "build.sh 0.262.0"` matched the waiter's own
command line, so it waited for itself while the build had already succeeded. Same bug cost 15
stuck watchers earlier in the day. Recorded as its own memory rule: watch a sentinel the work
writes, never a process name.
## What shipped
- **`stacks.Manager.AnyUpdating()`** — is a guarded update in flight for ANY app.
- **`Updater.SetAppUpdatingCheck`** — a deliberate sibling of `SetBackupRunningCheck`, consulted in
the **same three places** (the dry run, `TriggerUpdate`, `maybeAutoUpdate`). One busy-gate pattern
in that file, not two.
- **`Manager.SetSelfUpdatingCheck`** — the reverse. `UpdatePreflight` refuses `self_updating`.
- **Both halves wired in `main.go`**, the only place holding both objects. **`stacks` never imports
`selfupdate`** — the dependency is inverted with a plain callback rather than by widening
`UpdateGuards`, which is the backup side's interface and has nothing to do with this.
- **Two sentences, born as bundle keys**, in both bundles, registered in `i18n_go_keys.json`.
- **`data.reason` on every update refusal (R-609)** — additive; the sentence is unchanged, so no page
moves. `busy`/`updating`/`deploying`/`migrating`/`self_updating` are transient;
`held`/`downgrade` terminal; `memory`/`disk`/`no_backup` need a person.
`felhom-controller@3d41758` — **v0.262.0** (R-630, R-634-half, R-633/R-626, R-621, R-614) and
**v0.262.1** (the 409). All 17 controller gates green; the go-parity gate caught both new i18n keys
as unregistered before the push and was right to.
## The two things the tests found that reading did not
`app-catalog-felhom.eu@02844ae` — six proven versions moved (one commit each), `healthcheck.container`
for paperless-ngx and immich, the probe gate rewritten to the controller's four rules, decoy suite
56 → **64 cases**.
1. **The router refuses a HELD app on its own line, BEFORE `UpdatePreflight`** (`api/router.go`
~L601). Without a second edit, `held` — the single reason an unattended caller most needs — would
have been the one missing from the wire, and such a caller would press a terminally-refused button
on every pass for ever. Found while writing the table, not while reading the code.
2. **The lock must not latch.** `Stack.Updating` is cleared on done, failed **and held**, so a held
app does not block the controller's own updates — including the release that might fix whatever
held it. A latching gate would be a worse failure than the one prevented, and silent for weeks.
`TestR608_LockReleasesAfterHold` is a consequence test and exists for that alone.
## Proven live on 9202
## Red-proofs — five, each SEEN to fail
| # | the mutation | what failed |
| | before (v0.261.0) | after |
|---|---|---|
| 1 | `AnyUpdating` always false | `TestR608_AnyUpdatingSeesAnUpdateInFlight` — the gate reads false with an update in flight |
| 2 | `AnyUpdating` true for a held app too | `TestR608_LockReleasesAfterHold` — *"a HELD app must not hold the self-update lock for ever"* |
| 3 | delete the preflight's `self_updating` block | `TestR608_PreflightRefuses…` — *"while the controller swaps itself the app update must be REFUSED"* |
| 4 | delete `TriggerUpdate`'s app-update block | `TestR608_TriggerUpdateRefused…` — the swap proceeds over a live app update |
| 5 | drop `Data` from the refusal | `TestR609_EveryRefusalCarriesItsReason` — `reason = ""` on `no_backup` and `self_updating` |
| **A** paperless-ngx Update | `failed` at **+313.0 s**, app stopped, front door 404 | **`done` at +53.4 s**, running |
| **B** remove during a backup | — | refused, `REFUSED (busy): a backup or restore is running`; remove after → `verified: true`, nothing 60 s later |
| **C** remove a half-state | `stack "sparkyfitness" is not deployed` | **200**, `leftovers: NONE` |
| **F** redeploy after remove | inherited the old phase | phase `done` → **no phase** |
## Green gate
**Five red-proofs, each seen failing then green.** D (hold logs) is pinned by the two existing tests
whose compose sequence now includes the capture — they caught it and made me justify the order.
`go build ./... && go vet ./... && go test ./...` — **all green, full suite.**
`controller_gates.py --fast` — **17 gates OK**; `go-parity` convicted the two new keys first and was
satisfied properly by registering both as BORN-AS-KEYS with the test that pins each. No
`--no-verify`; the pre-push hook ran and passed.
## What is owed
Image `gitea.dooplex.hu/admin/felhom-controller:0.261.0` built and pushed.
**Deployed to guest 9202 (the scratch guest) ONLY** — the demo boxes and the fleet stay on 0.260.0.
**The floor is NOT raised to 0.261.0**; that is a separate ask for the operator, recorded in
`felhom.eu/STATUS.md`.
The measurements this release was built for — the three power cuts and the unattended night — are in
`felhom.eu/documentation/audits/update-arc-gaps-2026-09-21/`.
- **R-634's mechanism** — why `deployed` goes false while containers run.
- **R-625** — one verdict for the held badge and the Update button.
- **The floor.** v0.262.1 is on 9202 only. Raising it is the operator's call.