night 2026-09-24: Part C spike (09 §6.4.2 build brief), demo-hp pool incident evidence, rows R-668..R-680
gates / gates (push) Successful in 41s

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS
This commit is contained in:
2026-09-24 13:30:30 +02:00
parent 1cedf94d6b
commit 6ebe95c13e
25 changed files with 2541 additions and 1 deletions
@@ -1115,7 +1115,7 @@ what the part can do to a household's data if it is wrong, not how likely that i
| **4** | **SHIPPED — catalog `6db08a5`, night 2026-09-23** (`audits/DRILL-night-2026-09-23.md`): `update_ladder:` in `.felhom.yml`, one JSON entry per line (spiked on controller v0.266.0 and v0.267.0 first — the controller ignores the key); `check-test-record.py` (static, CI) + `check-test-record-move.py` (history + the registry for moved refs only, decision 23); the only writer `upgrade-test.py --write-ladder`; the 21 moves of 2026-09-22 backfilled from their records (21 proven). Not built: `steps/<to>.yml` (part 5 needs it once an app has two steps); `CompareImageRefs` did NOT move to the gate — the gate asks for a proven test instead, which is decision 13's own test. **The test record + the catalog gate + the memory check.** The harness writes the ladder entry (below) from its verdict record, including the memory watch's peak and marks; the gate refuses an image move with no entry, an entry with a `failed` verdict, or one with no memory watch; `CompareImageRefs`' rule moves here as the push-time safety net. **Backfill:** one entry per current pin — the 21 proven moves from their records, every other pin `needs_person: "never tested"`, which is honest and keeps them manual. A version move re-checks `mem_limit` against the watch's peak (the RomM follow-up: gate, not checklist, because the watch now produces the number). | 13, R-635 follow-up | **2.5** | the memory watch (shipped 2026-09-23) | none on a box — catalog-side only |
| **5** | **SHIPPED — controller v0.268.0 (`206b035`) + catalog `5ed599c`, proven live on 9202 2026-09-24** (`audits/ladder-2026-09-24/partD/`): romm 5.3.0/11.4 → 5.3.1/11.4 (its app step, from `steps/90dd9d68258286ef.yml`) → 5.3.1/11.8 (the engine step; `mariadb-upgrade` ran) in two presses, the seeded account read back after each, the page's „Hátralévő frissítési lépések" 2 → 1 → none. Step files: `templates/<app>/steps/<StepKey(to)>.yml` (sha256 of `to` as canonical JSON, 16 hex; `stacks.StepKey` = `ladder.step_key`), refused absent or wrong by `check-test-record.py` rule 4, written by `--write-ladder` when a step is superseded, 8 backfilled from the NEWEST commit naming each step's images (romm's from `f4eb94f`, not the OOM-looping `15f9ebf`). The box reads them from its `--depth 1` clone (the whole tree is there — measured). One press = one step; a missing step file refuses before anything moves; an installed version matching no entry jumps, logged by name (measured live: vikunja 2.5.0). Limitation: a step has no `.felhom.yml` of its own (R-664). **The ladder on the box.** Read `update_ladder:` from the clone, find the installed step, apply ONE step with its OWN definition (`steps/<to>.yml`, the last step the current template), repeat next night; a failed step stops the ladder for that app. `CatalogOrder` compares refs with the digest stripped (see part 7). | 14 | **2.5** | 1, 4 | medium — each step is the guarded update + undo; the new risk is rendering the wrong step's definition, pinned by a test per step shape |
| **6** | **CATALOG HALF SHIPPED — night 2026-09-23:** every ladder entry carries the digest per `to` ref (`scripts/image_digest.py`, stdlib; equals Docker's `RepoDigests` on a box), and the move gate refuses a digest the registry no longer serves. **The box half (compare, render `name:tag@sha256`) is not built.** **Digests.** The catalog records `sha256` per pin at push time (`check-image-resolvable.py` already resolves it); the box compares it for the badge and renders `name:tag@sha256:…` when present. **Measured 2026-09-23 on 9202:** Docker and Compose both pull and run `redis:7-alpine@sha256:858f…`, and refuse a digest that does not exist (`audits/update-rulings-2026-09-23/70-…`). **Build trap, read from source:** `splitImageRef` returns "unorderable" for ANY ref containing `@` (`updateorder.go:134`), so the digest must be split off before ordering or every digest-pinned app reads Unknown. A digest gone upstream fails the PULL — Scenario E, pin back, nothing ran. | 17, R-446 | **2** | 4 (the entry carries the digest) | low |
| **7** | **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has FINISHED (legs are clock-scheduled today, not chained — a completion signal is new), one app at a time (there is no single-flight, §3b Q4), `app_update.unattended` default ON, `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it. **See the one open point below.** | 11, 12 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe |
| **7** | **SPIKED 2026-09-24 (night, Part C) — the measured build brief is §6.4.2 below.** **The update leg in the chain + the automatic caller + the switch.** A leg that starts when the off-site leg has ENDED (on every exit path, skips included), one app at a time, one step per press, `app_update.unattended` default ON (decision 12), `stacks.update_window` removed, reads `UpdateRefusal.Reason`, remembers a failed step so it never re-presses it, and the full-system backup's gate waits for it until W+5h (decision 20). | 11, 12, 14, 15, 20 | **3** | 1, 2, 5 | medium — the only part that acts with nobody watching; everything above is what makes it safe |
| **8** | **SHIPPED — controller v0.265.0 + hub v0.121.0, proven live on 9202 2026-09-23** (`audits/cleanup-2026-09-23/`; a kernel `oom_kill` counter, not the sticky flag — `08` §6.2). **R-636** — the same OOM key re-firing escalates instead of staying one `warning` for six hours. | R-636 | **1** | — | none |
| **9** | **SHIPPED — controller v0.265.0, proven live on 9202 2026-09-23** (badge „Megállítva — visszaállítás szükséges" / "Stopped — restore needed", no Update button, 409 `held` unchanged). **R-625** — a held app stops offering an Update it will refuse. With the undo, holds become rarer; the lie on the page does not go away by itself. | R-625 | **0.5** | 1 | none |
| **10** | **PostgreSQL majors converted by the box.** A guarded-update step: `pg_dumpall` from the old engine, a NEW datadir (the old one kept aside, never deleted, until the check passes), load, check; then each of the eleven apps proven on the bench before its catalog move. | 16, R-463 | **2 + 3** | 1 (the same load discipline), 4 | **HIGH** — it rebuilds the datadir; bounded by keeping the old datadir aside |
@@ -1140,6 +1140,66 @@ step takes ~1 min when it works and ~2–6 min when it fails and is undone.
**Recommendation: the first.** It keeps the ruling's order and its promise that the full-system backup
is never skipped for an update; only the start time inside its existing window moves.
### 6.4.2 Part 7 — the build brief, from measurement (night 2026-09-24, Part C)
Evidence: `audits/night-2026-09-24/C/` and `C-00…C-07`. Spike only — no product code for the caller was
written. Read with decisions 11, 12, 14, 15 and 20.
**1. The chain, as it runs today (source + a real night).** The legs are clock-scheduled from one window
start W and NOTHING waits on the leg before it: db-dump at W, Tier 2 at W+60m, off-site at W+105m
(`backupwindow.go:16-21`, `cmd/controller/main.go:1048/1134/1267`); the full-system backup's gate is a
poll loop that opens at W+2h for 4 h (`quiesce.go:656-657`, window check `quiesce.go:238-245`).
**Measured on demo-hp 9201's real night (W = 02:30):** the off-site leg ran 04:15:05 → 04:18:33 CEST
(3 min 28 s, `ok`); the gate opened at 04:30. **So a leg placed after the off-site leg has ~11 minutes
before the gate** — decision 20's wait is required, not optional.
**2. The signal that the off-site leg has ended.** There is no event. The nearest thing is the
persisted `offbox` status in `settings.json` — `LastStatus` (`running` at start, `ok`/`incomplete`/
`error` at the end) with `LastRun` beside it, written by `UpdateOffboxStatus` (`offbox.go:1003-1125`).
The status travels with the timestamp (presence-is-not-success holds), **but it is not a reliable
"finished" signal:** five early returns (`offbox.go:863-897` — not configured, escrow pending, orphaned,
migration, another backup running) and the scheduler's own skip (`main.go:1268-1271`) write nothing, and
a box with no off-site target (9202 tonight) never runs the leg at all. **Build: CHAIN, do not poll.**
The `offbox-backup` job function calls the update leg after `RunOffboxBackup` returns, on EVERY path
including every skip — one call site, no new signal to go stale. A box with no off-site target runs
the update leg at W+105m.
**3. The gate's interlock (decision 20).** Insert after the window check in `quiesce.runOnce`, as an
`Options` func beside `WindowStartFn` (`quiesce.go:75-95`): *update leg active and now < W+5h → defer*
(log and return, exactly like the window deferral; the loop re-polls every 5 min). The leg itself stops
STARTING steps at W+5h; a step already running finishes (≤ health timeout + undo). Manual whole-box runs
(`TriggerNow`) bypass it, as today.
**4. One simulated night (9202, v0.269.1; the caller = `tools/unattended_caller_ladder.py`, which presses
only the public Update).** Four apps: wishlist 1 step behind, navidrome 2, romm 3, vikunja 1 with a
failing step. Measured per step: navidrome 10.6 s and 8.5 s (no database); wishlist 55.8 s (its first
update: the backup ran); romm 77.4 s, 71.2 s, 93.9 s (database + volume dump each step); vikunja's
failing step **undone in 104 s** with the drill's 90 s health timeout — **with the default 5 min it is
≈ 6–11 min** (verify up to 5 min, the undo's own verify up to 5 min). The whole leg: 7 min 18 s for
four steps and one undo. The household's pages: badges „Naprakész" for the three that climbed, and for
vikunja the undone sentence in both languages („…A doboz automatikusan visszaállította az előző
változatot és az adatokat — semmi nem veszett el." / "…put back the previous version and its data
automatically — nothing was lost."); event `app_update_undone` (dropped on 9202 — no hub, as designed).
During the db-dump leg every press was refused `busy` (transient, retried next pass).
**5. What the spike found that the build must handle** (rows filed):
- **`ladder_steps_left` and the badge are STALE after a step ends `done`** until the next scan — the first
caller run re-pressed a current app four times (R-678). The leg rescans after every step, and the
update's own finish should refresh them.
- **An Update pressed on an app already at the head runs the whole guarded update** — backup, pull,
restart, `done` — with no refusal (R-679: navidrome restarted four times for nothing). The leg never
presses an app whose pin equals the head; the preflight should refuse with reason `current`.
- **The product does not remember a failed step.** After vikunja's undo the badge still offers the
same step; only the caller's memory kept it from a re-press. Decision 15 needs it on the box: the
failed `to` is recorded per app and the leg skips it until the catalog's ladder changes.
**6. The build, in order (≈ 3 evenings, unchanged):** (a) the leg as a function called from the
off-site job on every path; one app at a time, one step per press, rescan between steps, stop at W+5h;
(b) the failed-step record + `current` refusal (R-679); (c) the gate's `UpdateLegActiveFn`; (d) the
per-box switch `app_update.unattended`, **default ON (decision 12)**, and `stacks.update_window`
removed; (e) the live proof: this night's four apps, with the window moved forward through the
product's own form (`POST /backups/window` — measured: the three legs reschedule at once).
### 6.4.1 (record) The update night — the drill brief that preceded the rulings, costed and re-costed
**The ruling is decision 6: all 53 apps, through the nightly rotation.** This is an ORDER inside that