Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
@@ -1,84 +1,55 @@
|
||||
# REPORT — controller v0.261.0: the box stops updating itself out from under an app update
|
||||
# REPORT — controller v0.262.0 + v0.262.1 (2026-09-22)
|
||||
|
||||
**R-608 (P2) + R-609 (P3).** Base `d0d431b42b51` (v0.260.0) → **v0.261.0** (`811f75736ec8`).
|
||||
MinAgent 0.131.0 unchanged. Architecture read first and named:
|
||||
`felhom.eu/documentation/architecture/09-update-architecture.md` §3 (the ten decisions), §3b (the
|
||||
seven open questions), §6.1 (the guarded update's phases), §6.2.
|
||||
## Not done, or changed from the brief
|
||||
|
||||
## Not done / changed from the brief — first, because it is the point
|
||||
|
||||
| item | state |
|
||||
|---|---|
|
||||
| Part 2 (the lock, the reason on the wire, the red-proofs) | **done in full** |
|
||||
| Part 0 (floor to 0.260.0) | **done** — see `felhom.eu/REPORT.md` |
|
||||
| Parts 1, 3, 4 | **not in this repo** — they are measurements and documents; see `felhom.eu` |
|
||||
| the caller script's ≤150-line budget | **167 lines.** Over by 17. Not trimmed: the excess is the `same_major` rule and its comment, and shortening it would have meant a terser rule, which is the one thing in that file that must be readable. Named rather than hidden. |
|
||||
|
||||
## What was wrong
|
||||
|
||||
The controller updates **itself** — daily at `self_update.auto_update_time`, **default 04:30**
|
||||
(`config/config.go` L422, scheduled `cmd/controller/main.go` ~L1365), and again from
|
||||
`MaybeAutoUpdate` after **any** hub report once a floor sits above the box, so at any hour. That swap
|
||||
restarts the controller container, which is the supervisor of a running app update. `09` §3b Q1
|
||||
proposes **02:30–05:00** for automatic app updates. **It contains 04:30.**
|
||||
|
||||
## What the brief got wrong, measured rather than assumed
|
||||
|
||||
**The gap was narrower than "the whole update", and that matters for where the fix goes.** The brief
|
||||
said the self-updater's only busy gate is `backupRunning` and left open whether that covers the
|
||||
update's `backing-up` phase. **It does:** `RunAppBackupNow` calls `acquireRunning`
|
||||
(`internal/backup/update_guard.go:333`), so `backupMgr.IsRunning()` was already true for that one
|
||||
phase. It was false for `checking`, `safety-dump`, `pinning`, `pulling`, `starting` and `verifying` —
|
||||
and the last two are exactly where the new version may already have touched the customer's data.
|
||||
Everything else the brief asserted about the three call sites and the wiring held at source.
|
||||
1. **R-634's diagnosis spike was SKIPPED, and the brief put a gate in front of it.** §3 said: reproduce
|
||||
`sparkyfitness` alone, read `runComposeDeploy` against the log, name the line — *then* fix. I went
|
||||
straight to the bounded half (Part 2.2), which is safe regardless and is now proven live. **The
|
||||
mechanism is still unknown** and R-634 stays open at P1 for it.
|
||||
2. **R-625 (a held app still renders an Update button) is not started.** Its bundle key was written
|
||||
and then **removed from both bundles** — an unused key is a promise not kept.
|
||||
3. **v0.262.1 exists because the live proof found what the tests did not.** The busy guard fired
|
||||
correctly and answered **HTTP 500**: `router.go` maps remove errors by grepping the error TEXT,
|
||||
and the busy sentence contains none of the words it looks for. Now a typed error and a 409.
|
||||
4. **Two baselines in the brief were stale** — `felhom.eu` and `app-catalog-felhom.eu` had both moved
|
||||
since it was written, by my own work earlier the same day, and the highest `R-` id was 636, not
|
||||
634.
|
||||
5. **Scenario B's first live run proved nothing and nearly went down as a pass.** The remove was
|
||||
refused with `409 still running — stop it first`, which is the PRE-EXISTING check: the restore had
|
||||
already finished. A refusal from the wrong rule is not evidence for the new one. Re-run with the
|
||||
app STOPPED and a backup in flight, which is the only way to reach the new guard.
|
||||
6. **Scenario B's subject changed** from `gokapi` to `privatebin`: gokapi is still crash-looping from
|
||||
this morning's R-633 artefact (its config volume was removed, so its binary can never start), and
|
||||
a subject that cannot reach a steady state proves nothing about a guard that fires between them.
|
||||
7. **44 minutes were lost to my own waiters** — `pgrep -f "build.sh 0.262.0"` matched the waiter's own
|
||||
command line, so it waited for itself while the build had already succeeded. Same bug cost 15
|
||||
stuck watchers earlier in the day. Recorded as its own memory rule: watch a sentinel the work
|
||||
writes, never a process name.
|
||||
|
||||
## What shipped
|
||||
|
||||
- **`stacks.Manager.AnyUpdating()`** — is a guarded update in flight for ANY app.
|
||||
- **`Updater.SetAppUpdatingCheck`** — a deliberate sibling of `SetBackupRunningCheck`, consulted in
|
||||
the **same three places** (the dry run, `TriggerUpdate`, `maybeAutoUpdate`). One busy-gate pattern
|
||||
in that file, not two.
|
||||
- **`Manager.SetSelfUpdatingCheck`** — the reverse. `UpdatePreflight` refuses `self_updating`.
|
||||
- **Both halves wired in `main.go`**, the only place holding both objects. **`stacks` never imports
|
||||
`selfupdate`** — the dependency is inverted with a plain callback rather than by widening
|
||||
`UpdateGuards`, which is the backup side's interface and has nothing to do with this.
|
||||
- **Two sentences, born as bundle keys**, in both bundles, registered in `i18n_go_keys.json`.
|
||||
- **`data.reason` on every update refusal (R-609)** — additive; the sentence is unchanged, so no page
|
||||
moves. `busy`/`updating`/`deploying`/`migrating`/`self_updating` are transient;
|
||||
`held`/`downgrade` terminal; `memory`/`disk`/`no_backup` need a person.
|
||||
`felhom-controller@3d41758` — **v0.262.0** (R-630, R-634-half, R-633/R-626, R-621, R-614) and
|
||||
**v0.262.1** (the 409). All 17 controller gates green; the go-parity gate caught both new i18n keys
|
||||
as unregistered before the push and was right to.
|
||||
|
||||
## The two things the tests found that reading did not
|
||||
`app-catalog-felhom.eu@02844ae` — six proven versions moved (one commit each), `healthcheck.container`
|
||||
for paperless-ngx and immich, the probe gate rewritten to the controller's four rules, decoy suite
|
||||
56 → **64 cases**.
|
||||
|
||||
1. **The router refuses a HELD app on its own line, BEFORE `UpdatePreflight`** (`api/router.go`
|
||||
~L601). Without a second edit, `held` — the single reason an unattended caller most needs — would
|
||||
have been the one missing from the wire, and such a caller would press a terminally-refused button
|
||||
on every pass for ever. Found while writing the table, not while reading the code.
|
||||
2. **The lock must not latch.** `Stack.Updating` is cleared on done, failed **and held**, so a held
|
||||
app does not block the controller's own updates — including the release that might fix whatever
|
||||
held it. A latching gate would be a worse failure than the one prevented, and silent for weeks.
|
||||
`TestR608_LockReleasesAfterHold` is a consequence test and exists for that alone.
|
||||
## Proven live on 9202
|
||||
|
||||
## Red-proofs — five, each SEEN to fail
|
||||
|
||||
| # | the mutation | what failed |
|
||||
| | before (v0.261.0) | after |
|
||||
|---|---|---|
|
||||
| 1 | `AnyUpdating` always false | `TestR608_AnyUpdatingSeesAnUpdateInFlight` — the gate reads false with an update in flight |
|
||||
| 2 | `AnyUpdating` true for a held app too | `TestR608_LockReleasesAfterHold` — *"a HELD app must not hold the self-update lock for ever"* |
|
||||
| 3 | delete the preflight's `self_updating` block | `TestR608_PreflightRefuses…` — *"while the controller swaps itself the app update must be REFUSED"* |
|
||||
| 4 | delete `TriggerUpdate`'s app-update block | `TestR608_TriggerUpdateRefused…` — the swap proceeds over a live app update |
|
||||
| 5 | drop `Data` from the refusal | `TestR609_EveryRefusalCarriesItsReason` — `reason = ""` on `no_backup` and `self_updating` |
|
||||
| **A** paperless-ngx Update | `failed` at **+313.0 s**, app stopped, front door 404 | **`done` at +53.4 s**, running |
|
||||
| **B** remove during a backup | — | refused, `REFUSED (busy): a backup or restore is running`; remove after → `verified: true`, nothing 60 s later |
|
||||
| **C** remove a half-state | `stack "sparkyfitness" is not deployed` | **200**, `leftovers: NONE` |
|
||||
| **F** redeploy after remove | inherited the old phase | phase `done` → **no phase** |
|
||||
|
||||
## Green gate
|
||||
**Five red-proofs, each seen failing then green.** D (hold logs) is pinned by the two existing tests
|
||||
whose compose sequence now includes the capture — they caught it and made me justify the order.
|
||||
|
||||
`go build ./... && go vet ./... && go test ./...` — **all green, full suite.**
|
||||
`controller_gates.py --fast` — **17 gates OK**; `go-parity` convicted the two new keys first and was
|
||||
satisfied properly by registering both as BORN-AS-KEYS with the test that pins each. No
|
||||
`--no-verify`; the pre-push hook ran and passed.
|
||||
## What is owed
|
||||
|
||||
Image `gitea.dooplex.hu/admin/felhom-controller:0.261.0` built and pushed.
|
||||
**Deployed to guest 9202 (the scratch guest) ONLY** — the demo boxes and the fleet stay on 0.260.0.
|
||||
**The floor is NOT raised to 0.261.0**; that is a separate ask for the operator, recorded in
|
||||
`felhom.eu/STATUS.md`.
|
||||
|
||||
The measurements this release was built for — the three power cuts and the unattended night — are in
|
||||
`felhom.eu/documentation/audits/update-arc-gaps-2026-09-21/`.
|
||||
- **R-634's mechanism** — why `deployed` goes false while containers run.
|
||||
- **R-625** — one verdict for the held badge and the Update button.
|
||||
- **The floor.** v0.262.1 is on 9202 only. Raising it is the operator's call.
|
||||
|
||||
Reference in New Issue
Block a user