From 6b2176e48026ea8bfb261b0528c12b514ba69521 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Fri, 25 Sep 2026 04:54:52 +0200 Subject: [PATCH] night 2026-09-25: DRILL record, Part D real night, STATUS morning note, register 341 -> 338 Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 17 ++ REPORT.md | 45 +--- STATUS.md | 35 ++-- .../architecture/09-update-architecture.md | 2 + .../audits/DRILL-night-2026-09-25.md | 193 ++++++++++++++++++ .../D/D4-demo-felhom-agent.log | 0 .../D/D4-demo-felhom-controller.log | 28 +++ .../night-2026-09-25/D/D4-demo-hp-agent.log | 0 .../D/D4-demo-hp-controller.log | 16 ++ .../D/D5-demo-felhom-opengist-after.txt | 14 ++ .../D/D6-demo-felhom-agent-night.log | 0 .../D/D6-demo-felhom-night-full.log | 30 +++ .../D/D6-demo-hp-agent-night.log | 0 .../D/D6-demo-hp-night-full.log | 10 + .../D/D7-hub-report-update-leg.txt | 2 + .../night-2026-09-25/G/G4-drill-reset.txt | 2 + .../night-2026-09-25/G/G5-host-hub-after.txt | 40 ++++ .../audits/night-2026-09-25/PROGRESS.md | 4 + .../night-2026-09-25/REPORT-DRAFT-AB.md | 31 --- .../audits/night-2026-09-25/REPORT-DRAFT-C.md | 26 --- documentation/backlog/CLOSED-ITEMS.md | 1 + documentation/backlog/OPEN-ITEMS.md | 5 +- 22 files changed, 389 insertions(+), 112 deletions(-) create mode 100644 documentation/audits/DRILL-night-2026-09-25.md create mode 100644 documentation/audits/night-2026-09-25/D/D4-demo-felhom-agent.log create mode 100644 documentation/audits/night-2026-09-25/D/D4-demo-felhom-controller.log create mode 100644 documentation/audits/night-2026-09-25/D/D4-demo-hp-agent.log create mode 100644 documentation/audits/night-2026-09-25/D/D4-demo-hp-controller.log create mode 100644 documentation/audits/night-2026-09-25/D/D5-demo-felhom-opengist-after.txt create mode 100644 documentation/audits/night-2026-09-25/D/D6-demo-felhom-agent-night.log create mode 100644 documentation/audits/night-2026-09-25/D/D6-demo-felhom-night-full.log create mode 100644 documentation/audits/night-2026-09-25/D/D6-demo-hp-agent-night.log create mode 100644 documentation/audits/night-2026-09-25/D/D6-demo-hp-night-full.log create mode 100644 documentation/audits/night-2026-09-25/D/D7-hub-report-update-leg.txt create mode 100644 documentation/audits/night-2026-09-25/G/G5-host-hub-after.txt delete mode 100644 documentation/audits/night-2026-09-25/REPORT-DRAFT-AB.md delete mode 100644 documentation/audits/night-2026-09-25/REPORT-DRAFT-C.md diff --git a/CONTEXT.md b/CONTEXT.md index 4746c584..a64b3dfb 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -14,6 +14,23 @@ > language, one screen, no identifiers in the prose. Same subjects, different readers; merging them > would make one of the two audiences stop reading. `STATUS.md` is also a **view of `OPEN-ITEMS.md`** > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +## 2026-09-25 (night shift) — agents 0.133.0 + 0.134.0 delivered, demo-hp backs up, automatic updates shipped (controller v0.271.0) + +- **Agent v0.133.0 and v0.134.0 delivered to both demo boxes** by CC-signed `agent_update` jobs (ruling 1); restore + test back ON. v0.134.0 = R-685 agent half (backup space preflight; free space from `GET /nodes//storage` — + `GET /storage` has no usage; a build that read it started a real vzdump on demo-hp in its live test, aborted). +- **demo-hp: operator option A applied** — `local_backup_retention: 1`, two old archives pruned by PVE's prune, one + product-triggered backup fitted (8.18 GB). R-684 closed. +- **Controller v0.271.0 = `09` §6.4 part 7**, floor 0.271.0 (MinAgent 0.131.0). Decisions taken by CC unattended, + operator may reverse (`09` §3): **31** after W+5h the full-system gate waits only for a step already in flight, cap + W+5h30m; **32** the controller's self-update waits for the whole leg; **33** one step per app per night (the brief's + own words). The demo boxes' first real night: demo-felhom opengist 1.13 → 1.15 done in 20 s, demo-hp nothing to do; + both summaries in the hub's stored reports (`update_leg`, allowlisted in `wire_contract_gate.py` until a hub surface + exists). +- **Catalog:** n8n 2.41.2 and mealie v3.28.0 moved (bench + box proven). +- **Open for the operator:** Peti's box before it returns (STATUS); R-686 (resume the leg after a restart?). + Record: `documentation/audits/DRILL-night-2026-09-25.md`. + ## 2026-09-24 (evening) — the restore test off on the demo hosts, 9201 repaired, agent v0.133.0 + hub v0.124.0 (R-672, R-673) **Operator rulings (recorded in `03` §8).** (1) The scheduled restore-test is OFF on both demo hosts until an diff --git a/REPORT.md b/REPORT.md index 984e8e49..7193760a 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,39 +1,8 @@ -# REPORT — 2026-09-24: the fleet to 0.267.0 and 0.268.0; the undo after a restore; option A; the ladder +# REPORT — night shift 2026-09-24/25 (felhom.eu side) -Full record (not done first, claims checked, red-proofs, live proofs, teardown): -`documentation/audits/ladder-2026-09-24/README.md`. - -## Not done, or changed - -Nothing left undone. Changed: the operator's English sentence without „please"; the hold names one (the newest) -whole copy; Part A's second failing update was not a failing edge (R-665); R-653/R-656 unit-tested only. - -## Baselines → ends - -| repo | start | end | -|---|---|---| -| felhom-controller | `80e6ad8c4772` v0.267.0 | `206b035` v0.268.0 | -| felhom.eu | `3e58c184f62e` hub v0.121.0 | hub v0.122.0 (`e4d45a8`, manifest `500488c`) + docs | -| app-catalog-felhom.eu | `585a7cba22cf` | `5ed599c` | -| felhom-agent | `d9864a94bf62` | untouched | - -## Hub v0.122.0 - -`app_hold_no_whole_copy` in `allowedEventTypes` + `operatorOnlyEvents` + `perAppCooldownEvents`; test pinning both -registers, red-proofed. Rolled out by ArgoCD (Synced, image read back from the pod, startup line `felhom-hub 0.122.0 -starting`). - -## Floors - -0.267.0 (05:12:14Z) and 0.268.0 (06:47:13Z), each with declared MinAgent 0.131.0, read back from the hub; both demo -boxes healthy within 30 s. `drill-r50` stays held (agent 0.129.0). - -## Documents - -`09` §3 decisions 24–25 + §6.4 part 5 SHIPPED + §6.1a note; `07` §6 whole-copy truth table; `08` §5 fourth -suppression + the crash-loop measurement; capability map row; `CONTEXT.md`; `STATUS.md`. - -## Rows - -Closed R-658, R-659, R-660, R-651, R-653, R-656, R-40, R-663. Opened R-661, R-662, R-664, R-665, R-666, R-667. -Narrowed R-450. Register 336 → 335 rows, 683,233 → 679,393 bytes. +The full record is `documentation/audits/DRILL-night-2026-09-25.md` (opens with "Not done, or changed"). This repo +carried: the evidence tree `documentation/audits/night-2026-09-25/`; `09` (part 7 shipped, decisions 31–33), `03` +(delivery + R-685), `07` (the fourth leg; demo-hp retention), `08` (what decision 28's suppression covers), the +capability map; the register (341 → 338 rows); `STATUS.md`; the hub's global floor raised to 0.271.0 (MinAgent +0.131.0) through the operator UI endpoint; `scripts/wire_contract_gate.py` allowlists the report's new `update_leg` +with its reason. No hub release. diff --git a/STATUS.md b/STATUS.md index 949056aa..2cf37594 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,25 +1,32 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-09-24 (evening). The demo boxes are on controller 0.270.0. The HP customer box is repaired. A fixed host agent (0.133.0) is released and waits for your signature. Until then the automatic restore test is off on both demo boxes.** +**Updated 2026-09-25 (morning). Both demo boxes run controller 0.271.0 and host agent 0.134.0. Automatic app updates are built and ran on their own last night. The HP box can back up again.** **Decisions I took on my own (you may reverse them).** -- **The space check uses the real, uncompressed size, not the backup file size.** The rule in your brief ("file size × 1.2 + 5 GB") would NOT have stopped last night's accident. The HP box's backup file is 7 GB, but restoring it writes 22.6 GB. -- **To switch the test off I used the value −1, not 0.** In this setting, 0 means "every 6 hours". Only a negative number turns it off. -- **The hub got a small release too (0.124.0).** A thin disk pool now counts as urgent at 90 % (data or its bookkeeping part), with one alarm per pool per 6 hours. Before, the alarm came only at 95 %, up to 15 minutes late, and one full pool could silence another. +- **The box takes one tested step per app per night.** Your brief said "one tested step per app". An app three steps behind now needs three nights. +- **After the 5-hour mark, the full-system backup waits only for an update step that is already running, and never more than 30 minutes.** Starting the backup in the middle of a step would stop the app while it is being checked and undo a good update. +- **The controller does not update itself while the night's app updates run.** A controller restart in the middle would stop the rest of that night's updates. **What I did, and it worked.** -- **The HP customer box is repaired.** I stopped it, checked both disks (small damage fixed, second check clean), and started it. All apps came back, it took the current version by itself, and the hub shows it online. Two apps' cache files (Redis) had been cut off by the full disk; with your yes I kept copies and repaired them. Both apps run. -- **The box's own crash-loop stop worked for real:** while the cache was broken, the box stopped those two apps and told the hub, by itself. -- **The new agent never starts a restore test that does not fit.** Tested on the HP box: it refused, said why, and created nothing. -- **Four controller fixes (0.270.0):** Update on an app that is already current is refused, with no restart. An install cut off by a restart is now cleaned up, reported and shown on the page. After a restore, the box keeps the right health check. One log line is corrected. +- **The fixed host agent reached both demo boxes.** I signed it for each box. The restore test is on again. On the N100 box it passed (85 seconds). On the HP box it said "not enough space" and created nothing, which is correct. +- **The HP box backs up again (your option A).** It now keeps one old whole-box backup instead of three. I removed the two oldest. A backup started with the backup page's own button fitted: 8 GB, and the disk is now 60 % full. +- **Automatic updates are built.** Each night, after the off-site copy, the box updates its apps by itself: one app at a time, one tested step each, the previous version ready to put back. There is a switch on the settings page, in both languages, on by default. A successful update sends no mail; the app page shows a line. +- **Proven on the scratch box first, over six simulated nights.** A failing update was put back and not tried again until the catalog re-tested it. A step marked "needs a person" was never taken. A power cut during an update: the box finished it after the restart. The controller killed during an update: the update was cleanly put back. Switch off: nothing happened. The app data read back after every night. +- **The demo boxes' first real automatic night.** The N100 box updated opengist (1.13 → 1.15) by itself at 04:15, in 20 seconds. The HP box had nothing to update, and said so. Both reported to the hub. +- **A new host agent (0.134.0) skips a whole-box backup that cannot fit**, and says why, before it starts. It is on both demo boxes. +- **Two more apps moved in the catalog** (n8n, mealie), each tested twice before it moved. **What broke, or is not done.** -- **The HP box's own backups have been failing every day since yesterday.** Its host disk has 4 GB free, but one backup needs 7 GB, and three old ones are kept. Nothing fixes this by itself. -- **No full restore test fits on the HP box today** (it needs 30 GB free, 22 GB is free). The new agent will refuse it every time, correctly. So a passing test on the HP box is not proven yet. +- **My own test started a real whole-box backup on the HP box for 5 minutes.** I stopped it. Nothing was left behind, and the apps kept running. The cause was a bug in the new agent, and I fixed it before the release. +- **If the controller restarts in the night, the rest of that night's updates wait for the next night.** Your decision is below. +- **The backup-page sentence for "backup does not fit" is not built yet.** The agent half is done. The page half needs the next controller release. +- **Last night no whole-box backup was due on either box**, so the "backup waits for the updates" rule did not happen for real yet. The tests prove it. -**Rows.** 1 opened, 4 closed, 2 updated. The list went from 344 to 341. +**Rows.** 3 opened, 6 closed, 2 updated. The list went from 341 to 338. **What needs you.** -1. **Sign the agent update (0.133.0) for demo-hp and demo-felhom.** Version 0.133.0, sha256 `3aa303452b8c6be58573d00af01a0ab4a0d97e4f885ffecd0144a18ac24e69b6`, one `felhom-opsign -op agent_update` job per box, the same way as 0.131.0. If you do nothing, the boxes keep the old agent and the automatic restore test stays off. -2. **After the new agent is on a box, turn the restore test back on.** On that host, remove the key `restore_test_eval_interval_seconds` from `backup` in `/etc/felhom-agent/agent.json` (the saved copy `agent.json.pre-r672` has it absent), then `systemctl restart felhom-agent`. If you do nothing, no backup is proven restorable on that box. -3. **Decide what to do with the HP box's backup storage:** keep fewer old backups, back up somewhere else, or add disk. If you do nothing, its whole-box backup keeps failing every night. +1. **Peti's box: what should its automatic updates be before it comes back online?** It has been silent since 15 July, on a very old version. The switch is on by default. If it comes back and takes the new version, it updates its apps by itself from its first night. + - **A (my choice): hold Peti's box on its current version until someone looks at it after it returns.** Cost: it stays behind until then, and the hold must be removed later by hand. + - **B: leave it.** Cost: on its first night online it jumps many versions and updates its apps by itself, with nobody watching. + - **If you do nothing, B happens.** +2. **If the controller restarts in the night, should the box continue that night's updates?** (A) Yes: the box remembers the night and continues until the 5-hour mark. (B) No: it waits a day, as now. I would choose A. If you do nothing, B stays. diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 8824eb6a..c873798b 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -1256,6 +1256,8 @@ product's own form (`POST /backups/window` — measured: the three legs reschedu (c) the gate interlock is `quiesce.Options.UpdateLegFn` with decision 31's in-flight grace; (d) the switch is `settings.json` `app_update.unattended`, absent = ON, a card on the settings page; `stacks.update_window` removed; (e) the live proof: `audits/DRILL-night-2026-09-25.md` Part C (four nights on 9202) and Part D (the demo boxes). +**The demo boxes' first real automatic night (2026-09-25):** demo-felhom's leg ran at 04:15:46 after the off-site +copy and took opengist 1.13 → 1.15 in 20 s; demo-hp's found nothing to take. Both summaries reached the hub. **Found live, not in the brief:** the leg presses the next app in the same second the previous step ends, so "a controller killed BETWEEN two apps" lands inside the next step's early phases — which the guarded update puts back (Scenario G); the rest of the night is lost (R-686). diff --git a/documentation/audits/DRILL-night-2026-09-25.md b/documentation/audits/DRILL-night-2026-09-25.md new file mode 100644 index 00000000..0f9194e5 --- /dev/null +++ b/documentation/audits/DRILL-night-2026-09-25.md @@ -0,0 +1,193 @@ +# DRILL — night 2026-09-25: the fixed agent delivered, the HP box backs up again, automatic updates built, proven, and watched through the demo boxes' first real night + +Brief: "NIGHT SHIFT 2026-09-24/25". Architecture read before any claim: `09-update-architecture.md` §3 decisions +11–30 and §6.4.2 (the spec for Part B), `03-host-agent.md`, `07-backup-architecture.md`, `08-alarm-ladder.md`, +`04-control-plane-authorization.md` §2–§3.1. Evidence: `audits/night-2026-09-25/` (A, B, C, D, E, F, G, tools). + +## Not done, or changed + +- **An unplanned whole-box backup of demo-hp 9201 ran for 5 min 16 s (22:40:21–22:45:37 CEST) and was aborted by + me.** It came from my own live test of agent v0.134.0 (R-685), which was meant to prove a REFUSAL: the first build + read free space from `GET /storage`, which carries no usage, so the check failed open and a real vzdump started. + I stopped the task; vzdump removed its temporary files; no archive, no lock; root disk back to 60 %; the guest was + suspended 2 s (snapshot mode), its apps kept running. This breached the brief's fence "no whole-box backup on a + demo box except Part A4". Fixed before release (free space from `GET /nodes//storage`), red-proofed twice, + and re-proven with safe builds that stop before any vzdump (`F/F1-*`, `F/F2-*`). +- **The floor was raised at 22:59 CEST, not "around 02:00".** Part C had passed; raising early gave 3.5 hours to + see both boxes arrive and read their switch and window before the night. Both arrived in ~15 s. +- **Part F ran BEFORE Part D, and its controller half is not done.** Agent v0.134.0 (the space check) was released, + signed and delivered to both demo boxes (Part F's live proof passed). The controller's backup-page line for a space + skip needs a controller release, and v0.271.0 was already the night's one — R-685 stays open for that half. + **R-671, R-670, R-677 were not started** — each needs a controller release (one per repo per night). +- **Both agent_update jobs of 0.134.0 were queued within a minute**, not one box after the other's commit (each is + per-box by signature; the 0.132.0 delivery did the same). 0.133.0 went strictly one after the other. +- **Part C's accidents: two ran live, two could not** on a scratch box without an off-site tier and an agent — the + off-site leg FAILING and W+5h reached with steps left are unit-proven only (R-687). "The controller killed BETWEEN + two apps" landed INSIDE the next app's step: the leg presses the next app in the same second (see Part C). +- **Part D: no whole-box backup followed the leg tonight, so the backup gate never had to wait.** demo-hp's last + whole-box backup was Part A4's (21:59), so it is not due until this evening; demo-felhom's is due ~07:27, after the + session. Both legs ended by 04:19, before the gate opened at 04:30. The gate's wait is unit-proven and red-proofed + only (R-687 updated). +- **One step per app per night** (decision 33) follows the brief's own words; §6.4.2 (a)'s "rescan between steps" + read as climbing several — recorded as a decision the operator may widen. + +**Interventions: 0** on product behaviour during the real night (nothing was touched on the demo boxes between the +floor raise and the end of the night except to read). **The demo boxes' real night: demo-felhom 1 step done +(opengist 1.13 → 1.15, 20 s), demo-hp nothing to do (every app current); 0 undone, 0 held, 0 skipped.** **The one +result that matters most:** the automatic leg ran unattended on both demo boxes exactly where the chain puts it — +after the off-site copy — pressed only the one tested step waiting, and reported itself to the hub. + + +## Part B — controller v0.271.0 (`9cf13a3`, CI job 974 success) + +Built exactly §6.4.2 point 6 (a)–(d) plus R-680, R-678, decision 13's marks, the older-than-ladder rule, the switch +and the events rule. As built: `stacks/unattended.go` (`RunUpdateLeg`), `chainUpdateLeg` in `cmd/controller/main.go` +(the `offbox-backup` job), `quiesce.Options.UpdateLegFn`, `backupwindow.UpdateLegStopOffsetMin`, `settings.json` +`app_update.unattended` (absent = ON) with its card on /settings, app.yaml `failed_update_step` and +`last_auto_update`, the report's `update_leg`, `backup.FreshWholeCopy`; `stacks.update_window` removed. Decisions +31–33 taken unattended (`09` §3). **Red-proofs — 11, each seen failing, tree restored after each** +(`night-2026-09-25/B/redproofs/`): the leg on every off-site path (a deferred call: error and panic cases fail without +it); the gate deferring while the leg runs and not at W+5h; the leg starting nothing at W+5h; one step per app a +night; the failed step skipped; `needs_person` skipped; `files_may_change` without a whole copy skipped; no crash-loop +verdict during a step (the first version of this test passed under the mutation — its container was set before the +leg's rescan rebuilt the stack; fixed and re-seen red); the steps-left count fresh at `done`; the switch ON by +default; the leg wired at startup. Gates green; one parity snapshot regenerated for the new card (diff: additions only). + +## Part A — the agent, the restore test, the HP backup + +**A1 — agent v0.133.0 signed and delivered, one box after the other** (ruling 1, 2026-09-16). Package sha256 +`3aa30345…e69b6` verified against the published package before signing (`A/A1-*.txt`). + +| box | signed (UTC) | box took it | AUTHORIZED → COMPLETED | committed | hub reads | +|---|---|---|---|---|---| +| demo-felhom-8363b5 | 19:25:30 | 21:34:05 CEST | job `c4b592ff7c9ff307` | 21:35:09 | 0.133.0 | +| demo-hp-bb76ea | 19:35:55 | 21:48:58 CEST | job `36aab1ed8ad00979` | 21:50:02 | 0.133.0 | + +Peti's box: never signed for. + +**A2 — the restore test back on, both boxes.** `agent.json.pre-r672` copied back; the diff was only +`restore_test_eval_interval_seconds: -1` (and the trailing newline). The `-1` config is kept as +`agent.json.night-0925-off`. Start-up line on both: `restore-test scheduler starting (per-archive due-check) +eval_interval=6h0m0s settle=24h0m0s`. + +**A3 — one on-demand restore test.** demo-felhom: **PASS** — scratch 990000 restored, booted, verified, +torn down in 1 min 25 s; pool 3.13 % → peak 5.68 % → 3.13 %; preflight "required 12.8 GB, avail 362.8 GB". +demo-hp: **REFUSED, correctly** — "restoring 21.1 GiB (vzdump log: total bytes written) needs 30.3 GiB free, has +22.1 GiB", exit 4, pool 59.01 % before and after, nothing created. + +**A4 — HP whole-box backups, option A.** Before: root disk (`local`) 89 % used, 4.3 GB free, three 9201 archives +(6.2, 6.3, 7.4 GB), retention 3. Changed: `local_backup_retention` 3 → 1 (saved copy +`agent.json.pre-a4-retention`), agent restarted, the two oldest archives removed with PVE's own +`pvesm prune-backups local --vmid 9201 --keep-last 1` (dry-run shown first) → 58 % used, 16.8 GB free. **Why the +prune by hand:** PVE prunes AFTER a successful backup, so lowering the retention alone would not have let the next +backup fit. One whole-box backup by the product's own button (`POST /api/guest-backup/trigger`, the Mentések +page) — 9201's apps were stopped for it (snapshot mode, resumed at "snapshotted"): local archive **8.18 GB** in +7 min, then the PBS tier (the button covers every tier) 23.2 GB in 4.5 min; the new prune kept 1 archive; root +60 %; all 20 app containers healthy. The product row is **R-685**. +## Part C — the proof on 9202: six simulated nights + +Controller v0.271.0 on 9202 only (`C/C0-deploy-9202.txt`), drill catalog (`C/C0-repoint-drill.txt`, health timeout +90 s). Four apps installed through the product at the ladder's first `from`, then the drill put back at the head: +wishlist 1 step behind, navidrome 2, romm 3, vikunja 1 with a failing step (its head probe on port 8999). Marks +set in the drill: navidrome step 1 `needs_person`, wishlist's step `files_may_change`. Data seeded through each +app's own front door and read back after EVERY night (`C/data-readback.txt`: 8 of 8 rounds, 4 of 4 apps, each +fixture with its negative control). The window was moved through the product's own form (`POST /backups/window`) +so the off-site job — and the leg chained to it — fired two minutes later. + +| night | what it proved | leg summary (the box's own line) | pages (hu + en) | +|---|---|---|---| +| 1 | the chain on a box with NO off-site target ("the off-site leg does nothing; the update leg runs now"); one step per app; `needs_person` skipped; `files_may_change` taken WITH a whole copy (wishlist); a failing step undone | `done=2 undone=1 held=0 failed=0 skipped=1 in 3m6s [skipped: navidrome=needs_person]` — romm 60.1 s, vikunja undone 105.2 s, wishlist 20.1 s | „Automatikus frissítés 2026-09-24 22:31-kor — sikeres." / "Automatic update at … — done."; vikunja's undone sentence in both | +| 2 | R-680 live: the undone step is NOT pressed again; navidrome's `files_may_change` step taken (its music folder is `excluded`, so its own unit is whole) | `done=2 undone=0 … skipped=1 in 1m10s [skipped: vikunja=failed_before]` | badges current / behind as expected | +| 3 — accident | controller killed (kill -9) the moment navidrome's step ended | the leg had pressed romm IN THE SAME SECOND; the kill landed before romm's `up` → pin and definition put back, romm runs its previous version, page: „A frissítés megszakadt, mert a vezérlő újraindult…" / "The update was interrupted…"; the leg was not resumed (R-686) | all data read back | +| 4 — accident | power cut (`pct stop`/`pct start` of 9202) while romm's step was `verifying` | box up in 13 s; "resumed 1 interrupted update(s)"; romm healthy after 30 s, DONE in 1 m 32 s; no automatic-update line on the page (the leg died — R-686) | all data read back | +| 5 | the switch OFF (`POST /settings/app-update`, page reads unchecked) | `done=0 … stopped=switched_off in 0s` — nothing pressed | — | +| 6 | the switch ON again; the catalog RE-TESTED vikunja's step (probe fixed, new `tested_at`) → the ladder print changed → R-680's record no longer binds | `done=1 … in 10s` — vikunja done | „Automatikus frissítés … 22:58-kor — sikeres." | + +**Not run live, and why (R-687):** W+5h reached with steps left (the leg starts at W+105m; needs a 3-hour leg — unit +test `TestLeg_NoStepAtOrAfterW5h`); the off-site leg FAILING (9202 has no off-site tier — unit test +`TestChainUpdateLeg_EveryPath` covers error and panic); a `files_may_change` step WITHOUT a whole copy (every app +given the mark was whole on 9202); the full-system gate's deferral (9202 has no agent, so no quiesce loop). + +**Verdict: Part C passed** — every accident that could run on 9202 ended with the app healthy, the data read back, +and a true sentence on its page; the four not-runnable cases are named, unit-proven and in the register. + +## Part D — the demo boxes' first real automatic night + +Floor 0.271.0 (declared MinAgent 0.131.0) saved 20:59:04Z, read back; both boxes on 0.271.0 ~15 s later +(`D/D1`, `D/D2`). Before the night, per box (`D/D3`, `D/D0`): switch ON (key absent = the default; demo-hp's +settings page reads it checked in both languages), window 02:30 (off-site + update leg at 04:15, gate [04:30, 08:30)). +Tested steps waiting: **demo-felhom — opengist 1.13 → 1.15; demo-hp — none** (all 10 apps at the catalog head). +Agents: 0.134.0 on both (Part F). Nothing was touched on either box during the night except to read. + +**demo-felhom (N100), leg by leg** (`D/D6-demo-felhom-night-full.log`): db-dump 02:30:00 (0.8 s) → Tier 2 03:30:00 +→ off-site 04:15:00–04:15:46 (1 app, 11 snapshots, 42 s) → **update leg 04:15:46–04:16:06: opengist pressed; safety +dump, undo copy (0.2 MiB), pin, pull, start, verify healthy after 10 s, DONE in 17 s; `done=1 undone=0 held=0 failed=0 +skipped=0 in 20s`**. After: container `opengist:1.15` healthy; app.yaml `last_auto_update: done 1.13 → 1.15` (the page +line). One WARN worth keeping: the volume "carries no compose label (recreated by a restore before v0.268.0 — R-658) +— copied by name" — the R-658 fallback working. No mail (a successful step sends none). + +**demo-hp** (`D/D6-demo-hp-night-full.log`): db-dump 02:30:00–02:31:59 → Tier 2 03:30:00–03:30:24 → off-site +04:15:00–04:18:41 (9 apps, 90 snapshots, 3 min 35 s, OK) → **update leg 04:18:41: nothing to press, `done=0 … in +0s`**. + +**The hub** holds both summaries in its stored reports (received 02:44Z; `D/D7-hub-report-update-leg.txt`, read from a +copy of the hub DB taken with its -wal, then deleted). **Whole-box backups:** none ran tonight on either box — not +due (see "Not done"); both agents alive (627 / 1,069 journal lines since 02:25, the 10-minute janitor on schedule). +**No app ended held or stopped.** + +## Part E — two apps moved + +Venue: bench LXC **9401** on demo-hp (created and destroyed tonight — no permission check refused either; the +Debian template it needed was removed again), harness v3; box walk on 9202 through the real guarded Update. +**Negative control: `n8n=alpine:3.20` → `failed`.** A bench setup miss is recorded: the first queue ran without +the `traefik-public` network and returned `inconclusive — FROM deploy failed` (kept as `*-noNetwork`); re-run. + +| app | step | bench | box | memory peak (own) | catalog commit (CI) | +|---|---|---|---|---|---| +| n8n | 2.41.1 → 2.41.2 | proven | proven | 22.4 % | `f14a608` (job 980) | +| mealie | v3.27.0 → v3.28.0 | proven | proven | 23.1 % | `b996218` (job 981) | + +Published at 02:31–02:32, after the demo boxes' night began and on neither box's app list. Not moved, with the +reason for each: catalog `REPORT.md` (drift `E/E0-drift.log`). + +## Part F — agent v0.134.0 (R-685 agent half) + +Released by `scripts/release-agent.sh` (tag `v0.134.0` at `0722b2c`, sha256 `7593bebe…81c72d`, verified by download; +CI jobs 975–977). A vzdump to a local target needs free ≥ newest archive × 1.25 + 1 GiB; a shortfall is a named skip +before anything starts; fail-open on PBS / first backup / unknown usage. Live, safe builds (a hard stop before vzdump): +×10 → "local has 14.9 GiB free; … needs about 77.2 GiB" refused; ×1.25 → "space preflight passed" need 11.3 GB, +avail 16.0 GB. Signed to both demo boxes (`F/F3-*`): demo-felhom committed 22:52:45, demo-hp 22:58:53; hub reads 0.134.0. +The incident above belongs to this part. + +## Teardown — three layers + +- **Machine:** 9202 back on the live catalog (`catalog-cache` at `c802509` at the time; the controller follows `main`), + window 02:30 again, the six drill apps removed through the product (0 containers, 0 volumes, 0 undo copies); left by + name on the scratch drive: `userdata/navidrome`, `userdata/romm` (the remove kept drive data — R-442's fail-closed + path on 9202). The switch key on 9202 is now explicit `true` (was absent = ON). 9202 stays on controller 0.271.0 (the + floor). Bench 9401 destroyed (hostname checked first), template removed. +- **Host:** demo-hp `pct list` 9201 + 9202, local-lvm 60.05 %, root 60 %; demo-felhom 9201 only, local-lvm 3.16 %. + Config changes kept on purpose: restore test ON (both), `local_backup_retention: 1` (demo-hp). No test binaries left + in /tmp. +- **Hub:** global floor 0.271.0 (MinAgent 0.131.0); both demo boxes 0.271.0 / agent 0.134.0. Drill repo reset to live + `main` (`dd9c6ad`), private, Actions off. + +## Claims in the brief that turned out wrong (or right) + +1. *The off-site job function is the single place the leg can be chained from on every path* — **TRUE for every path + of the job, with one exception outside it:** a box with backups disabled registers no off-site job at all, so no + leg (the update itself would refuse `no_backup` there anyway). The job's own early returns and the scheduler's "not + configured" skip all return into `chainUpdateLeg`. A controller restart after the job fired loses the rest of the + night (R-686). +2. *A demo box has a tested step waiting tonight* — **TRUE for demo-felhom (opengist), FALSE for demo-hp.** +3. *The whole-box backup lands on demo-hp's root disk* — **TRUE** (`local` = `/var/lib/vz`, root LV). +4. *Decision 28's suppressions cover a deploy's first start* — **WRONG.** `Deploying` clears when `compose up -d` + returns; a first start that restarts ≥ 6 times in 10 min is stopped (R-676 updated). An automatic step, its verify + and its undo ARE covered (pinned by a test). + +## Register + +Before: **341 rows / 685,148 B**. After: **338 rows / 682,053 B**. Opened R-685 (backup-fit warning; agent half +shipped, controller page half open), R-686 (no resume of the leg after a restart), R-687 (live-proof gaps). Closed +R-672, R-673 (delivered), R-684 (option A), R-680, R-678, R-643. Updated R-676 (deploy first start), R-450 (narrowed +to part 10). diff --git a/documentation/audits/night-2026-09-25/D/D4-demo-felhom-agent.log b/documentation/audits/night-2026-09-25/D/D4-demo-felhom-agent.log new file mode 100644 index 00000000..e69de29b diff --git a/documentation/audits/night-2026-09-25/D/D4-demo-felhom-controller.log b/documentation/audits/night-2026-09-25/D/D4-demo-felhom-controller.log new file mode 100644 index 00000000..34ed904b --- /dev/null +++ b/documentation/audits/night-2026-09-25/D/D4-demo-felhom-controller.log @@ -0,0 +1,28 @@ +2026/09/25 02:15:00 [INFO] [scheduler] Running job: offbox-backup +2026/09/25 02:15:00 [INFO] [offbox] backup run started (1 app(s) toggled) +2026/09/25 02:15:00 [INFO] [offbox] pre-push dump leg completed in 795ms — snapshot pair is coherent +2026/09/25 02:15:08 [INFO] [offbox] backed up opengist (/mnt/sys_drive/felhom-data/backups/primary/opengist, 0 mandatory path(s)) +2026/09/25 02:15:46 [INFO] [offbox] backup OK: 1 app(s) backed up, 11 snapshot(s), 42s +2026/09/25 02:15:46 [INFO] [update-leg] started (after-offsite): window 02:30, no step starts at or after 07:30 +2026/09/25 02:15:46 [INFO] [stacks] update opengist: accepted — guarded update started +2026/09/25 02:15:46 [INFO] [update-leg] opengist: step pressed opengist=ghcr.io/thomiceli/opengist:1.13 → opengist=ghcr.io/thomiceli/opengist:1.15 +2026/09/25 02:15:46 [INFO] [stacks] update opengist: phase checking +2026/09/25 02:15:46 [INFO] [stacks] update opengist: ladder — the last step (1 of 1) — the catalog's current definition +2026/09/25 02:15:46 [INFO] [stacks] update opengist: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T02:15:00Z (1m0s old, limit 24h0m0s) +2026/09/25 02:15:46 [INFO] [stacks] update opengist: phase safety-dump +2026/09/25 02:15:46 [INFO] [stacks] update opengist: safety dump done (0 file(s)) [] +2026/09/25 02:15:47 [WARN] [stacks] update opengist: volume(s) [opengist_opengist_data] carry no compose label (recreated by a restore before v0.268.0 — R-658) — copied by name +2026/09/25 02:15:47 [INFO] [stacks] update opengist: the undo copy will hold 1 named volume(s), 0.2 MiB +2026/09/25 02:15:47 [INFO] [stacks] update opengist: phase pinning +2026/09/25 02:15:47 [INFO] [stacks] update opengist: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/opengist/docker-compose.yml (opengist=ghcr.io/thomiceli/opengist:1.15) +2026/09/25 02:15:47 [INFO] [stacks] update opengist: phase pulling +2026/09/25 02:15:53 [INFO] [stacks] update opengist: phase copying +2026/09/25 02:15:53 [INFO] [stacks] update opengist: phase copying +2026/09/25 02:15:53 [INFO] [stacks] update opengist: copied opengist_opengist_data → opengist_opengist_data.pre-update-20260925T021553Z in 293ms +2026/09/25 02:15:53 [INFO] [stacks] update opengist: phase starting +2026/09/25 02:15:54 [INFO] [stacks] update opengist: phase verifying +2026/09/25 02:16:04 [INFO] [stacks] update opengist: healthy after 10s (the app's health check passed) +2026/09/25 02:16:04 [INFO] [stacks] update opengist: DONE in 17s +2026/09/25 02:16:06 [INFO] [update-leg] opengist: step ended done after 20.1 s +2026/09/25 02:16:06 [INFO] [update-leg] update leg (after-offsite): done=1 undone=0 held=0 failed=0 skipped=0 in 20s [skipped: ] +2026/09/25 02:16:06 [INFO] [scheduler] Job offbox-backup completed (took 1m6.957s) diff --git a/documentation/audits/night-2026-09-25/D/D4-demo-hp-agent.log b/documentation/audits/night-2026-09-25/D/D4-demo-hp-agent.log new file mode 100644 index 00000000..e69de29b diff --git a/documentation/audits/night-2026-09-25/D/D4-demo-hp-controller.log b/documentation/audits/night-2026-09-25/D/D4-demo-hp-controller.log new file mode 100644 index 00000000..e1868bd7 --- /dev/null +++ b/documentation/audits/night-2026-09-25/D/D4-demo-hp-controller.log @@ -0,0 +1,16 @@ +2026/09/25 02:15:00 scheduler.go:346: [INFO] [scheduler] Running job: offbox-backup +2026/09/25 02:15:00 offbox.go:922: [INFO] [offbox] backup run started (9 app(s) toggled) +2026/09/25 02:16:54 offbox.go:969: [INFO] [offbox] pre-push dump leg completed in 1m54.507s — snapshot pair is coherent +2026/09/25 02:17:12 offbox.go:1388: [INFO] [offbox] backed up kimai (/mnt/sys_drive/felhom-data/backups/primary/kimai, 0 mandatory path(s)) +2026/09/25 02:17:15 offbox.go:1388: [INFO] [offbox] backed up opengist (/mnt/sys_drive/felhom-data/backups/primary/opengist, 0 mandatory path(s)) +2026/09/25 02:17:26 offbox.go:1388: [INFO] [offbox] backed up paperless-ngx (/mnt/felhom-drives/hdd_1/backups/primary/paperless-ngx, 1 mandatory path(s)) +2026/09/25 02:17:35 offbox.go:1388: [INFO] [offbox] backed up romm (/mnt/felhom-drives/hdd_1/backups/primary/romm, 0 mandatory path(s)) +2026/09/25 02:17:39 offbox.go:1385: [WARN] [offbox] backed up bentopdf (/mnt/sys_drive/felhom-data/backups/primary/bentopdf, 0 mandatory path(s)) — but the recovery unit carried NO database dump and NO volume tar, so this snapshot holds none of the app's data; the next run with a dump leg will replace it +2026/09/25 02:17:45 offbox.go:1388: [INFO] [offbox] backed up bookstack (/mnt/sys_drive/felhom-data/backups/primary/bookstack, 0 mandatory path(s)) +2026/09/25 02:17:50 offbox.go:1388: [INFO] [offbox] backed up calibre-web (/mnt/felhom-drives/hdd_1/backups/primary/calibre-web, 1 mandatory path(s)) +2026/09/25 02:17:54 offbox.go:1388: [INFO] [offbox] backed up privatebin (/mnt/sys_drive/felhom-data/backups/primary/privatebin, 0 mandatory path(s)) +2026/09/25 02:18:00 offbox.go:1388: [INFO] [offbox] backed up docmost (/mnt/sys_drive/felhom-data/backups/primary/docmost, 0 mandatory path(s)) +2026/09/25 02:18:41 offbox.go:1158: [INFO] [offbox] backup OK: 9 app(s) backed up, 90 snapshot(s), 3m35s +2026/09/25 02:18:41 unattended.go:254: [INFO] [update-leg] started (after-offsite): window 02:30, no step starts at or after 07:30 +2026/09/25 02:18:41 unattended.go:235: [INFO] [update-leg] update leg (after-offsite): done=0 undone=0 held=0 failed=0 skipped=0 in 0s [skipped: ] +2026/09/25 02:18:41 scheduler.go:363: [INFO] [scheduler] Job offbox-backup completed (took 3m41.126s) diff --git a/documentation/audits/night-2026-09-25/D/D5-demo-felhom-opengist-after.txt b/documentation/audits/night-2026-09-25/D/D5-demo-felhom-opengist-after.txt new file mode 100644 index 00000000..6ac463bc --- /dev/null +++ b/documentation/audits/night-2026-09-25/D/D5-demo-felhom-opengist-after.txt @@ -0,0 +1,14 @@ +ghcr.io/thomiceli/opengist:1.15 Up 4 minutes (healthy) +last_auto_update: +2026/09/25 02:15:00 [INFO] [scheduler] Running job: offbox-backup +2026/09/25 02:15:00 [INFO] [offbox] backup run started (1 app(s) toggled) +2026/09/25 02:15:00 [INFO] [offbox] pre-push dump leg completed in 795ms — snapshot pair is coherent +2026/09/25 02:15:08 [INFO] [offbox] backed up opengist (/mnt/sys_drive/felhom-data/backups/primary/opengist, 0 mandatory path(s)) +2026/09/25 02:15:46 [INFO] [offbox] backup OK: 1 app(s) backed up, 11 snapshot(s), 42s +last_auto_update: + at: "2026-09-25T02:15:46Z" + outcome: done + from: + opengist: ghcr.io/thomiceli/opengist:1.13 + to: + opengist: ghcr.io/thomiceli/opengist:1.15 diff --git a/documentation/audits/night-2026-09-25/D/D6-demo-felhom-agent-night.log b/documentation/audits/night-2026-09-25/D/D6-demo-felhom-agent-night.log new file mode 100644 index 00000000..e69de29b diff --git a/documentation/audits/night-2026-09-25/D/D6-demo-felhom-night-full.log b/documentation/audits/night-2026-09-25/D/D6-demo-felhom-night-full.log new file mode 100644 index 00000000..3d2df45d --- /dev/null +++ b/documentation/audits/night-2026-09-25/D/D6-demo-felhom-night-full.log @@ -0,0 +1,30 @@ +2026/09/25 00:30:00 [INFO] [scheduler] Running job: db-dump +2026/09/25 00:30:00 [INFO] [scheduler] Job db-dump completed (took 804ms) +2026/09/25 01:30:00 [INFO] [scheduler] Running job: tier2-backup +2026/09/25 01:30:00 [INFO] [scheduler] Job tier2-backup completed (took 5ms) +2026/09/25 02:15:00 [INFO] [scheduler] Running job: offbox-backup +2026/09/25 02:15:00 [INFO] [offbox] backup run started (1 app(s) toggled) +2026/09/25 02:15:46 [INFO] [offbox] backup OK: 1 app(s) backed up, 11 snapshot(s), 42s +2026/09/25 02:15:46 [INFO] [update-leg] started (after-offsite): window 02:30, no step starts at or after 07:30 +2026/09/25 02:15:46 [INFO] [stacks] update opengist: accepted — guarded update started +2026/09/25 02:15:46 [INFO] [update-leg] opengist: step pressed opengist=ghcr.io/thomiceli/opengist:1.13 → opengist=ghcr.io/thomiceli/opengist:1.15 +2026/09/25 02:15:46 [INFO] [stacks] update opengist: phase checking +2026/09/25 02:15:46 [INFO] [stacks] update opengist: ladder — the last step (1 of 1) — the catalog's current definition +2026/09/25 02:15:46 [INFO] [stacks] update opengist: precondition met — Tier 1 (own recovery unit) copy from 2026-09-25T02:15:00Z (1m0s old, limit 24h0m0s) +2026/09/25 02:15:46 [INFO] [stacks] update opengist: phase safety-dump +2026/09/25 02:15:46 [INFO] [stacks] update opengist: safety dump done (0 file(s)) [] +2026/09/25 02:15:47 [WARN] [stacks] update opengist: volume(s) [opengist_opengist_data] carry no compose label (recreated by a restore before v0.268.0 — R-658) — copied by name +2026/09/25 02:15:47 [INFO] [stacks] update opengist: the undo copy will hold 1 named volume(s), 0.2 MiB +2026/09/25 02:15:47 [INFO] [stacks] update opengist: phase pinning +2026/09/25 02:15:47 [INFO] [stacks] update opengist: pin advanced to /opt/docker/felhom-controller/data/catalog-cache/templates/opengist/docker-compose.yml (opengist=ghcr.io/thomiceli/opengist:1.15) +2026/09/25 02:15:47 [INFO] [stacks] update opengist: phase pulling +2026/09/25 02:15:53 [INFO] [stacks] update opengist: phase copying +2026/09/25 02:15:53 [INFO] [stacks] update opengist: phase copying +2026/09/25 02:15:53 [INFO] [stacks] update opengist: copied opengist_opengist_data → opengist_opengist_data.pre-update-20260925T021553Z in 293ms +2026/09/25 02:15:53 [INFO] [stacks] update opengist: phase starting +2026/09/25 02:15:54 [INFO] [stacks] update opengist: phase verifying +2026/09/25 02:16:04 [INFO] [stacks] update opengist: healthy after 10s (the app's health check passed) +2026/09/25 02:16:04 [INFO] [stacks] update opengist: DONE in 17s +2026/09/25 02:16:06 [INFO] [update-leg] opengist: step ended done after 20.1 s +2026/09/25 02:16:06 [INFO] [update-leg] update leg (after-offsite): done=1 undone=0 held=0 failed=0 skipped=0 in 20s [skipped: ] +2026/09/25 02:16:06 [INFO] [scheduler] Job offbox-backup completed (took 1m6.957s) diff --git a/documentation/audits/night-2026-09-25/D/D6-demo-hp-agent-night.log b/documentation/audits/night-2026-09-25/D/D6-demo-hp-agent-night.log new file mode 100644 index 00000000..e69de29b diff --git a/documentation/audits/night-2026-09-25/D/D6-demo-hp-night-full.log b/documentation/audits/night-2026-09-25/D/D6-demo-hp-night-full.log new file mode 100644 index 00000000..aad5f610 --- /dev/null +++ b/documentation/audits/night-2026-09-25/D/D6-demo-hp-night-full.log @@ -0,0 +1,10 @@ +2026/09/25 00:30:00 scheduler.go:346: [INFO] [scheduler] Running job: db-dump +2026/09/25 00:31:59 scheduler.go:363: [INFO] [scheduler] Job db-dump completed (took 1m59.507s) +2026/09/25 01:30:00 scheduler.go:346: [INFO] [scheduler] Running job: tier2-backup +2026/09/25 01:30:24 scheduler.go:363: [INFO] [scheduler] Job tier2-backup completed (took 24.096s) +2026/09/25 02:15:00 scheduler.go:346: [INFO] [scheduler] Running job: offbox-backup +2026/09/25 02:15:00 offbox.go:922: [INFO] [offbox] backup run started (9 app(s) toggled) +2026/09/25 02:18:41 offbox.go:1158: [INFO] [offbox] backup OK: 9 app(s) backed up, 90 snapshot(s), 3m35s +2026/09/25 02:18:41 unattended.go:254: [INFO] [update-leg] started (after-offsite): window 02:30, no step starts at or after 07:30 +2026/09/25 02:18:41 unattended.go:235: [INFO] [update-leg] update leg (after-offsite): done=0 undone=0 held=0 failed=0 skipped=0 in 0s [skipped: ] +2026/09/25 02:18:41 scheduler.go:363: [INFO] [scheduler] Job offbox-backup completed (took 3m41.126s) diff --git a/documentation/audits/night-2026-09-25/D/D7-hub-report-update-leg.txt b/documentation/audits/night-2026-09-25/D/D7-hub-report-update-leg.txt new file mode 100644 index 00000000..c2a07176 --- /dev/null +++ b/documentation/audits/night-2026-09-25/D/D7-hub-report-update-leg.txt @@ -0,0 +1,2 @@ +demo-felhom report received 2026-09-25 02:44:15 UTC; update_leg = {"trigger": "after-offsite", "started_at": "2026-09-25T02:15:46.825918487Z", "ended_at": "2026-09-25T02:16:06.961273956Z", "deadline": "2026-09-25T07:30:00+02:00", "enabled": true, "done": 1, "undone": 0, "held": 0, "failed": 0, "skipped": 0, "steps": [{"app": "opengist", "outcome": "done", "from": {"opengist": "ghcr.io/thomiceli/opengist:1.13"}, "to": {"opengist": "ghcr.io/thomiceli/opengist:1.15"}, "seconds": 20.1}]} +demo-hp report received 2026-09-25 02:44:17 UTC; update_leg = {"trigger": "after-offsite", "started_at": "2026-09-25T02:18:41.05475777Z", "ended_at": "2026-09-25T02:18:41.126639901Z", "deadline": "2026-09-25T07:30:00+02:00", "enabled": true, "done": 0, "undone": 0, "held": 0, "failed": 0, "skipped": 0, "steps": null} diff --git a/documentation/audits/night-2026-09-25/G/G4-drill-reset.txt b/documentation/audits/night-2026-09-25/G/G4-drill-reset.txt index 16d97c98..588cd6e3 100644 --- a/documentation/audits/night-2026-09-25/G/G4-drill-reset.txt +++ b/documentation/audits/night-2026-09-25/G/G4-drill-reset.txt @@ -1,2 +1,4 @@ drill=c8025093d23c live=c8025093d23c — 2026-09-24T21:50:07+00:00 has_actions: False private: True +drill=b99621897639 live=b99621897639 — 2026-09-25T00:32:29+00:00 (after the Part E moves) +drill=dd9c6ad57b07 live=dd9c6ad57b07 — 2026-09-25T00:32:50+00:00 (final) diff --git a/documentation/audits/night-2026-09-25/G/G5-host-hub-after.txt b/documentation/audits/night-2026-09-25/G/G5-host-hub-after.txt new file mode 100644 index 00000000..f7b76fd0 --- /dev/null +++ b/documentation/audits/night-2026-09-25/G/G5-host-hub-after.txt @@ -0,0 +1,40 @@ +# G5 host + hub layers after — 2026-09-25T02:52:16+00:00 +== felhom-pve +felhom-agent 0.134.0 +VMID Status Lock Name +9201 running demo-felhom +Name Type Status Total (KiB) Used (KiB) Available (KiB) % +felhom-backup dir active 960303848 11072284 900377052 1.15% +felhom-pbs pbs active 0 0 0 0.00% +local dir active 98497780 28725888 64722344 29.16% +local-lvm lvmthin active 365760512 11558032 354202479 3.16% + data 3.16 0.58 +/dev/mapper/pve-root 94G 28G 62G 31% / +agent.json +agent.json.bak-r191 +agent.json.campaign8-before +agent.json.campaign9-before +agent.json.campaign9-prev +agent.json.night-0925-off +agent.json.pre-e-target-move +agent.json.pre-prunegate.bak +agent.json.pre-r672 +== hp +felhom-agent 0.134.0 +VMID Status Lock Name +9201 running demo-hp +9202 running demo-hp-scratch +Name Type Status Total (KiB) Used (KiB) Available (KiB) % +felhom-pbs pbs active 0 0 0 0.00% +local dir active 40453376 22725576 15640684 56.18% +local-lvm lvmthin active 56487936 33921005 22566930 60.05% +nvme-scratch dir active 983379700 74952044 858401044 7.62% + data 60.05 2.67 +/dev/mapper/pve-root 39G 22G 15G 60% / +agent.json +agent.json.night-0925-off +agent.json.pre-a4-retention +agent.json.pre-r672 +== hub +Effective floor: v0.271.0 — source: DB (hub_settings) ; env +Customer Status Events Last Seen CPU Memory Disk Containers Last Backup Version Demo Ügyfél OK — 8 min ago 1% 16% 2% 1/1 – 0.271.0 Demo HP OK 2 4 4 8 min ago 10% 19% 33%/8% 20/20 – 0.271.0 Peti Proxmox DOWN — 71d ago 9% 66% 7% 2/2 – 0.115.0 Tester 1 DOWN 3 8d ago 13% 60% 29%/1% 22/22 – 0.245.0 Auto-refreshes every 60 seconds · Felhom Hub 0.124.0 diff --git a/documentation/audits/night-2026-09-25/PROGRESS.md b/documentation/audits/night-2026-09-25/PROGRESS.md index 445d0aab..fccb3ee1 100644 --- a/documentation/audits/night-2026-09-25/PROGRESS.md +++ b/documentation/audits/night-2026-09-25/PROGRESS.md @@ -24,3 +24,7 @@ Capacity: phase0-capacity.txt. demo-felhom pool 3.13 %, root 31 %. demo-hp pool - [x] E bench 9401 created+destroyed (template removed); n8n 2.41.1→2.41.2 and mealie v3.27.0→v3.28.0 PROVEN on bench AND box; control failed (correct). TODO after 02:30: --write-ladder both, commit each, push, CI. Then reset drill again. - [x] G partial: 9202 apps removed via product, window 02:30, live catalog c802509 (switch key now explicit true), drill reset to live - [ ] 02:31 publish E; 04:10 start D watch (d_watch.py) on both boxes +- [x] E published: n8n f14a608, mealie b996218 (CI 980/981), catalog CHANGELOG/REPORT dd9c6ad (982); drill reset to dd9c6ad +- [x] D real night: demo-felhom opengist done 20 s; demo-hp 0 steps; hub has both update_leg; no whole-box backup due on either (gate never waited) +- [x] G all three layers recorded (G1..G5); report DRILL-night-2026-09-25.md; STATUS; register 338 rows / 682,053 B +- unproven.py: NOT WALKED 35 of 55 (unchanged) diff --git a/documentation/audits/night-2026-09-25/REPORT-DRAFT-AB.md b/documentation/audits/night-2026-09-25/REPORT-DRAFT-AB.md deleted file mode 100644 index 94afeafb..00000000 --- a/documentation/audits/night-2026-09-25/REPORT-DRAFT-AB.md +++ /dev/null @@ -1,31 +0,0 @@ -## Part A — the agent, the restore test, the HP backup - -**A1 — agent v0.133.0 signed and delivered, one box after the other** (ruling 1, 2026-09-16). Package sha256 -`3aa30345…e69b6` verified against the published package before signing (`A/A1-*.txt`). - -| box | signed (UTC) | box took it | AUTHORIZED → COMPLETED | committed | hub reads | -|---|---|---|---|---|---| -| demo-felhom-8363b5 | 19:25:30 | 21:34:05 CEST | job `c4b592ff7c9ff307` | 21:35:09 | 0.133.0 | -| demo-hp-bb76ea | 19:35:55 | 21:48:58 CEST | job `36aab1ed8ad00979` | 21:50:02 | 0.133.0 | - -Peti's box: never signed for. - -**A2 — the restore test back on, both boxes.** `agent.json.pre-r672` copied back; the diff was only -`restore_test_eval_interval_seconds: -1` (and the trailing newline). The `-1` config is kept as -`agent.json.night-0925-off`. Start-up line on both: `restore-test scheduler starting (per-archive due-check) -eval_interval=6h0m0s settle=24h0m0s`. - -**A3 — one on-demand restore test.** demo-felhom: **PASS** — scratch 990000 restored, booted, verified, -torn down in 1 min 25 s; pool 3.13 % → peak 5.68 % → 3.13 %; preflight "required 12.8 GB, avail 362.8 GB". -demo-hp: **REFUSED, correctly** — "restoring 21.1 GiB (vzdump log: total bytes written) needs 30.3 GiB free, has -22.1 GiB", exit 4, pool 59.01 % before and after, nothing created. - -**A4 — HP whole-box backups, option A.** Before: root disk (`local`) 89 % used, 4.3 GB free, three 9201 archives -(6.2, 6.3, 7.4 GB), retention 3. Changed: `local_backup_retention` 3 → 1 (saved copy -`agent.json.pre-a4-retention`), agent restarted, the two oldest archives removed with PVE's own -`pvesm prune-backups local --vmid 9201 --keep-last 1` (dry-run shown first) → 58 % used, 16.8 GB free. **Why the -prune by hand:** PVE prunes AFTER a successful backup, so lowering the retention alone would not have let the next -backup fit. One whole-box backup by the product's own button (`POST /api/guest-backup/trigger`, the Mentések -page) — 9201's apps were stopped for it (snapshot mode, resumed at "snapshotted"): local archive **8.18 GB** in -7 min, then the PBS tier (the button covers every tier) 23.2 GB in 4.5 min; the new prune kept 1 archive; root -60 %; all 20 app containers healthy. The product row is **R-685**. diff --git a/documentation/audits/night-2026-09-25/REPORT-DRAFT-C.md b/documentation/audits/night-2026-09-25/REPORT-DRAFT-C.md deleted file mode 100644 index d3ecd7b3..00000000 --- a/documentation/audits/night-2026-09-25/REPORT-DRAFT-C.md +++ /dev/null @@ -1,26 +0,0 @@ -## Part C — the proof on 9202: six simulated nights - -Controller v0.271.0 on 9202 only (`C/C0-deploy-9202.txt`), drill catalog (`C/C0-repoint-drill.txt`, health timeout -90 s). Four apps installed through the product at the ladder's first `from`, then the drill put back at the head: -wishlist 1 step behind, navidrome 2, romm 3, vikunja 1 with a failing step (its head probe on port 8999). Marks -set in the drill: navidrome step 1 `needs_person`, wishlist's step `files_may_change`. Data seeded through each -app's own front door and read back after EVERY night (`C/data-readback.txt`: 8 of 8 rounds, 4 of 4 apps, each -fixture with its negative control). The window was moved through the product's own form (`POST /backups/window`) -so the off-site job — and the leg chained to it — fired two minutes later. - -| night | what it proved | leg summary (the box's own line) | pages (hu + en) | -|---|---|---|---| -| 1 | the chain on a box with NO off-site target ("the off-site leg does nothing; the update leg runs now"); one step per app; `needs_person` skipped; `files_may_change` taken WITH a whole copy (wishlist); a failing step undone | `done=2 undone=1 held=0 failed=0 skipped=1 in 3m6s [skipped: navidrome=needs_person]` — romm 60.1 s, vikunja undone 105.2 s, wishlist 20.1 s | „Automatikus frissítés 2026-09-24 22:31-kor — sikeres." / "Automatic update at … — done."; vikunja's undone sentence in both | -| 2 | R-680 live: the undone step is NOT pressed again; navidrome's `files_may_change` step taken (its music folder is `excluded`, so its own unit is whole) | `done=2 undone=0 … skipped=1 in 1m10s [skipped: vikunja=failed_before]` | badges current / behind as expected | -| 3 — accident | controller killed (kill -9) the moment navidrome's step ended | the leg had pressed romm IN THE SAME SECOND; the kill landed before romm's `up` → pin and definition put back, romm runs its previous version, page: „A frissítés megszakadt, mert a vezérlő újraindult…" / "The update was interrupted…"; the leg was not resumed (R-686) | all data read back | -| 4 — accident | power cut (`pct stop`/`pct start` of 9202) while romm's step was `verifying` | box up in 13 s; "resumed 1 interrupted update(s)"; romm healthy after 30 s, DONE in 1 m 32 s; no automatic-update line on the page (the leg died — R-686) | all data read back | -| 5 | the switch OFF (`POST /settings/app-update`, page reads unchecked) | `done=0 … stopped=switched_off in 0s` — nothing pressed | — | -| 6 | the switch ON again; the catalog RE-TESTED vikunja's step (probe fixed, new `tested_at`) → the ladder print changed → R-680's record no longer binds | `done=1 … in 10s` — vikunja done | „Automatikus frissítés … 22:58-kor — sikeres." | - -**Not run live, and why (R-687):** W+5h reached with steps left (the leg starts at W+105m; needs a 3-hour leg — unit -test `TestLeg_NoStepAtOrAfterW5h`); the off-site leg FAILING (9202 has no off-site tier — unit test -`TestChainUpdateLeg_EveryPath` covers error and panic); a `files_may_change` step WITHOUT a whole copy (every app -given the mark was whole on 9202); the full-system gate's deferral (9202 has no agent, so no quiesce loop). - -**Verdict: Part C passed** — every accident that could run on 9202 ended with the app healthy, the data read back, -and a true sentence on its page; the four not-runnable cases are named, unit-proven and in the register. diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 2e423690..06b6d46c 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -391,3 +391,4 @@ Compressed here to title, shipping version, evidence, and the sentences that sta | **R-684** | **demo-hp's whole-box backups could not fit on its root-disk target (P2).** Operator ruling 2026-09-24 evening, option A: retention 3 → 1 (`local_backup_retention`), the two oldest archives pruned by PVE's own prune, and one backup by the product's trigger fitted (8.18 GB; root 89 % → 60 %). The product lesson — warn BEFORE the night — is R-685 (agent v0.134.0 skips with a reason). | ruling 2026-09-24; applied 2026-09-24 night | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/A/A4-hp-backup-space.txt`, `A4-hp-backup-run.txt` | | **R-680** | **The box did not remember a failed update step (P2).** Controller v0.271.0: an undone or held step is recorded in app.yaml (`failed_update_step`, tied to the ladder's print); the automatic leg skips it until the catalog's ladder changes; a person can still press. Live on 9202: vikunja undone night 1, skipped `failed_before` night 2, re-tried after the catalog re-tested it. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/C/`; `B/redproofs/R680-*` | | **R-678** | **After a step ended `done`, steps-left and the badge stayed stale (P3).** Controller v0.271.0: the update re-reads the app's catalog fields BEFORE it says done (and after an undo). Live on 9202: every automatic step's page read current at the leg's end. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `B/redproofs/R678-*` | +| **R-643** | **The ruled chain left the automatic update leg at most 15 minutes a night (P2).** Decision 20, built in controller v0.271.0: the full-system backup's gate defers while the leg runs, until W+5h (then only for a step in flight, cap W+5h30m — decision 31); the leg starts no step at or after W+5h; one shared constant. Unit + red-proof (`TestD20_GateWaitsForTheLeg`); live on the demo boxes: see the night record Part D. | v0.271.0, 2026-09-25 | `git show 75ff264:documentation/backlog/OPEN-ITEMS.md`; `audits/night-2026-09-25/B/redproofs/D20-*` | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index fe3e13fa..d4b4d2fd 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -666,7 +666,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-444** | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | | **R-445** | **[P3-LOW] Hub app telemetry survives the app's removal, so a 15-minute throwaway now sets a FLEET-WIDE memory recommendation.** MEASURED 2026-09-01: this spike's Phase 6 Nextcloud existed for ~15 minutes on demo-hp, spent part of it crash-looping, and was then removed with all volumes. The hub's `/apps/nextcloud` page still reports `Deployments`, `Avg Memory 208 MB`, `P95 Memory 280 MB` and **`Suggested Limit (P95x1.2) = 352 MB`**, plus three MariaDB `io_uring` rows under Known Issues attributed to demo-hp. **The suggested limit is an operator-facing recommendation derived from a sample that no longer exists anywhere** — and Nextcloud is a real catalog app whose limit someone may act on. **RETAINED DELIBERATELY BY THIS RUN, NOT CLEARED, and the reason is part of the row:** the hub offers `POST /apps/nextcloud/reset-telemetry` whose own confirm reads *"Delete all telemetry data for nextcloud? This cannot be undone."* — an irreversible write on the operator's surface, and the operator authorised Phase 6, not this. **The one-line command is recorded in the audit doc so it is a decision, not a task.** The general question is the row: should telemetry for an app with zero live deployments age out, or be excluded from the suggestion? Owner: **VIKTOR rules, CC implements.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: VIKTOR rules, CC implements** | | **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** Slice 2 (controller v0.233.0, 2026-09-02) compares the RECORDED image reference per compose service against the reference the current template pins, and **queries no registry** — deliberately: a customer's box must not depend on reaching eight upstream registries to render a page (`felhom-controller/controller/internal/web/updatebadge.go`, `compareInstalledToTemplate`). **For the 23 floating pins that comparison is blind by construction:** `postgres:16-alpine`, `mariadb:11.6` and 21 others can carry an identical reference over an image that has moved. **MEASURED, not theorised — spike §5 found `mariadb:11.4` and `mariadb:12.3` had BOTH already moved upstream while two fully-pinned CONTROLS held.** So `romm` and `bookstack` on demo-hp would read „Naprakész" over a database engine build that is not the one the catalog now resolves to. **This is a KNOWN LIMITATION OF A SHIPPED FEATURE, filed the same session rather than left implicit**, and it is stated in the same words in `architecture/09-update-architecture.md` §8.1 and in the controller's `README.md`. The close is a digest comparison against the registry, which needs a network call, a cache and a failure posture — it is not a one-liner and it is not slice 2's job. **Depends on R-440**, whose fix (stop floating) would remove the problem instead of measuring it — take that route first if it is available. `architecture/09-update-architecture.md` **MEASURED 2026-09-21, and the blind spot is not one or two pins.** `audits/UPDATE-ARC-STATE-2026-09-21.md` §3.3: the catalog carries **10 floating pins of 66** (recounted — the old "23" was stale), and **6 of the 7 measurable engine pins have been repushed upstream since the catalog set them** — `postgres:16-alpine` (8 apps), `postgres:15-alpine`, `redis:7-alpine` (6 apps), `mariadb:11.4`, `mariadb:12.3`, `postgis:16-3.5-alpine`; only `mariadb:11.6` has not. The 8th (immich's own ghcr build) is UNMEASURED — ghcr exposes no anonymous last-modified timestamp. **So on demo-hp today four apps read „Naprakész" over a database engine image that has demonstrably moved.** The fix does NOT need the box to query a registry: the catalog can record each pin's digest at push time (`check-image-resolvable.py` already resolves it) and the box compares digests. Put to the operator as `09` §3b **Q6**, recommended YES — the cheapest real improvement on the arc's list. **— UPDATE NIGHT 2026-09-21:** **MEASURED ON A BOX 2026-09-21 (update night, leg B8), and it REFINES the row in two ways rather than merely confirming it.** §8.1's numbers came from a registry sweep on DooPlex; this is the same question asked of a customer-shaped box, where the badge actually renders. On guest 9202, `docmost`'s two floating pins were read as `installed_images` records them and compared against the upstream digests measured the same night: `postgres:16-alpine` → **`sha256:721873c34ceb9…` on the box and `sha256:721873c34ceb9…` upstream**, and `redis:7-alpine` → **`sha256:858f009f9709c…` both sides**. **Identical. So the badge „Naprakész" is TRUE for this box**, and the app reads correctly. **(1) The defect's size is set by INSTALL AGE, not by the catalog.** A floating pin is wrong only for a box that pulled BEFORE the tag moved; a box deployed after the repush holds the current image and its badge is right. R-446's "six repushed pins" measured the tag against the date the CATALOG set it, which is the right measure for *the catalog* and not for *a box*. **(2) The producer Q6 needs ALREADY EXISTS on the box.** `installed_images` records a real `digest` per service (`installed.go` §7.1) — the box knows exactly what it is running. What it cannot do is COMPARE, because the catalog carries no digest to compare against. That is Q6's proposal, and this is a concrete confirmation that only the catalog half is missing. Evidence: `audits/update-night-2026-09-21/23-B8-floating-pin.txt`. **-- RULED 2026-09-23 (`09` §3 decision 17):** YES — the catalog records the image digest of every pin at push time; the box compares against it and, where the catalog carries one, pulls **that exact image**, which makes a floating tag reproducible, not only the badge honest. *Pull-by-digest while the definition names a tag is a claim to verify in the build, not a ruling on mechanism.* **-- 2026-09-23:** pull-by-digest MEASURED on 9202 — Docker and Compose both pull and run `redis:7-alpine@sha256:858f…` and refuse a digest that does not exist. **Build trap, read from source:** `splitImageRef` returns "unorderable" for any ref containing `@` (`stacks/updateorder.go:134`), so a digest-carrying pin must have its digest split off before ordering or every such app reads Unknown. `09` §6.4 part 6. **— NIGHT 2026-09-23:** the CATALOG half of the close shipped: every ladder entry records the digest the registry served for each `to` ref (`scripts/image_digest.py`), and the move gate refuses a digest the registry no longer serves. The box does not compare it yet (`09` §6.4 part 6, box half). | **READY TO BUILD — owner: CC; `09` §6.4** | -| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. **-- RULED 2026-09-23 (operator, `09` §3 decisions 11–15):** Q1–Q4 answered. The update is a leg of the backup chain after off-site and before the full-system backup (11); automatic with a per-box switch ON by default (12); **the TEST decides, not the tag** — the box applies every step the catalog holds because the catalog holds only tested steps, and `CompareImageRefs` moves to the catalog gate (13, REPLACES "never across a major"); a box behind climbs **one tested step at a time** (14); **the box UNDOES a failed update itself** — old definition + the pre-pin safety dump + health check again, HOLD only if the undo fails (15, REPLACES §6.1's no-auto-undo). The undo and the ladder were SPIKED the same day before any build (`audits/update-rulings-2026-09-23/`); build order and costs in `09` §6.4. **-- SPIKED 2026-09-23 (`audits/update-rulings-2026-09-23/`):** the undo works by hand on three real migrating edges and needs eight product additions (R-637..R-642); **the ladder is measured absent** — one press on a box two steps behind jumped vikunja 2.3.0 → 2.5.0 in 9.5 s and 2.4.0 never ran, and the box cannot see intermediate steps at all because its catalog clone is `--depth 1` (`sync.go:283`/`:300`, one commit visible on both demo guests). The ladder's recommended format is an `update_ladder:` list in `.felhom.yml` with each intermediate step's own definition, NOT the git history (romm's image-moving commit is the definition that OOM-looped). The chain's update leg has ≤15 min as ruled (R-643). Build order `09` §6.4. **— NIGHT 2026-09-23: `09` §6.4 part 4 SHIPPED** (catalog `6db08a5`): the test record `update_ladder:` + two gates + the only writer + the 21-move backfill; 12 more steps published with records. Parts 5 (the box climbs), 6-box-half and 7 remain; the romm press on demo-hp showed today's jump live — 5.3.0 → 5.3.1 AND mariadb 11.4 → 11.8 in one press (both tested steps; `done`). **— 2026-09-24: `09` §6.4 PART 5 SHIPPED** (controller v0.268.0 `206b035`, catalog `5ed599c`): one press = one tested step, each step's own definition at `templates//steps/.yml`, proven live on 9202 (romm 5.3.0/11.4 → 5.3.1/11.4 → 5.3.1/11.8 in two presses, `audits/ladder-2026-09-24/partD/`). Left in this row: part 6's box half, part 7 (the automatic leg), part 10 (PostgreSQL majors). | **READY TO BUILD — owner: CC; `09` §6.4 part by part, each part returns to the operator for go/no-go** | +| **R-450** | **[P2-MEDIUM] UPDATE ARC SLICE 6 — a version sequence: automatic WITHIN a major, never ACROSS one, and an engine change gets its OWN edge.** The first half is an operator ruling of 2026-09-02 and its justification is R-449's measurement: a cross-major jump can be refused by the app itself and cannot be undone. **The second half is a rule recorded now, while it is cheap:** an engine change must never be bundled with an app version bump. `bookstack`'s `0b73e5e` moved the application 25.02.2 → 26.05.2 **and** MariaDB 11.6 → 12.3 in one commit — **two migrations behind one edge**, and an unreadable failure when it breaks. Needs a catalog-side convention and, eventually, a gate. `architecture/09-update-architecture.md` §6 **HALF SHIPPED 2026-09-21 (catalog `5ff36d098cbc`): the second half — an engine change gets its OWN edge — is now ENFORCED** by `check-engine-major.py`, which refuses a commit moving a MariaDB major together with any other image move in that template, naming what it was bundled with. The FIRST half (automatic within a major) is Slice 6 and needs four operator answers — `09` §3b **Q1–Q4**, with the shape it would take in `09` §6.2. **The urgency is now measured:** 46 of the catalog's 58 exact pins are behind upstream and **39 of those are within a major** — the population the 2026-09-02 ruling already says may move without a human. **-- RULED 2026-09-23 (operator, `09` §3 decisions 11–15):** Q1–Q4 answered. The update is a leg of the backup chain after off-site and before the full-system backup (11); automatic with a per-box switch ON by default (12); **the TEST decides, not the tag** — the box applies every step the catalog holds because the catalog holds only tested steps, and `CompareImageRefs` moves to the catalog gate (13, REPLACES "never across a major"); a box behind climbs **one tested step at a time** (14); **the box UNDOES a failed update itself** — old definition + the pre-pin safety dump + health check again, HOLD only if the undo fails (15, REPLACES §6.1's no-auto-undo). The undo and the ladder were SPIKED the same day before any build (`audits/update-rulings-2026-09-23/`); build order and costs in `09` §6.4. **-- SPIKED 2026-09-23 (`audits/update-rulings-2026-09-23/`):** the undo works by hand on three real migrating edges and needs eight product additions (R-637..R-642); **the ladder is measured absent** — one press on a box two steps behind jumped vikunja 2.3.0 → 2.5.0 in 9.5 s and 2.4.0 never ran, and the box cannot see intermediate steps at all because its catalog clone is `--depth 1` (`sync.go:283`/`:300`, one commit visible on both demo guests). The ladder's recommended format is an `update_ladder:` list in `.felhom.yml` with each intermediate step's own definition, NOT the git history (romm's image-moving commit is the definition that OOM-looped). The chain's update leg has ≤15 min as ruled (R-643). Build order `09` §6.4. **— NIGHT 2026-09-23: `09` §6.4 part 4 SHIPPED** (catalog `6db08a5`): the test record `update_ladder:` + two gates + the only writer + the 21-move backfill; 12 more steps published with records. Parts 5 (the box climbs), 6-box-half and 7 remain; the romm press on demo-hp showed today's jump live — 5.3.0 → 5.3.1 AND mariadb 11.4 → 11.8 in one press (both tested steps; `done`). **— 2026-09-24: `09` §6.4 PART 5 SHIPPED** (controller v0.268.0 `206b035`, catalog `5ed599c`): one press = one tested step, each step's own definition at `templates//steps/.yml`, proven live on 9202 (romm 5.3.0/11.4 → 5.3.1/11.4 → 5.3.1/11.8 in two presses, `audits/ladder-2026-09-24/partD/`). Left in this row: part 6's box half, part 7 (the automatic leg), part 10 (PostgreSQL majors). **— 2026-09-25 night: part 6's box half shipped in v0.269.0/v0.269.1; PART 7 SHIPPED as controller v0.271.0** (the automatic leg, proven over six simulated nights on 9202 and watched through the demo boxes' first real night, `audits/DRILL-night-2026-09-25.md`). **Left: part 10 only (PostgreSQL majors, R-463).** | **NARROWED to `09` §6.4 part 10 — owner: CC; returns to the operator for go/no-go** | | **R-451** | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 **BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies.** The controller's payload carries no image (`internal/report/types.go` L98–103) and the hub's `Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts** — but **the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in `09` §6.2–6.3; the payload question is `09` §3b **Q7**. **-- RULED 2026-09-23 (`09` §3 decision 18):** the report carries, per compose service (database included), the installed reference, the catalog reference and the badge state. **Built later, when the fleet grows** — Q7's recommendation, confirmed. | **RULED — build deferred until the fleet grows; owner: CC** | | **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** | | **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** | @@ -799,7 +799,6 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-633** | **[P2-MEDIUM] A remove sent while a restore is still running reports success, deletes the app's record, and leaves a container restarting forever with a live public route.** MEASURED 2026-09-22 on guest 9202, controller v0.261.0, during the twenty-eight walk. `gokapi` was restored from its own local copy at 11:34:07 and removed at 11:34:22. `POST /backup/restore` answers **302 and works in the background**; the remove tore down what existed and the restore's own `compose up` then RE-CREATED the container at **11:34:24**. **Both calls returned success.** Twenty-five minutes later: `GET /api/stacks/gokapi` reads **`deployed: false`**, and `docker ps -a` shows `gokapi` **`Restarting (1)`** with `RestartCount` climbing, carrying its full traefik label set — including `traefik.http.routers.gokapi.rule: Host(`.enkisfelhom.hu`)`, **a rule with an empty subdomain**, because the deploy values that filled it were deleted with the app. Its own log loops *„Salt for admin password invalid, generating new salt… password does not appear to be a SHA-1 hash"* — the volume holding its config was removed correctly, so the binary can never start. **WHY THIS IS A ROW AND NOT A HARNESS ARTEFACT:** a household can press exactly these two buttons in exactly this order, the product accepted both, and **the remove reported success while leaving the orphan**. Presence of a success message is not evidence of a result. **WHAT IT COSTS:** an app the household believes is gone keeps a container in a restart loop, keeps a route registered on the public reverse proxy, and is invisible to every product surface because the controller no longer records the stack. Nothing in the alarm ladder fires: `08` §4 keys on stacks the controller KNOWS about. **This is R-626's class with the mechanism finally visible** — that row recorded a removed `navidrome` coming back and could not diagnose it because the controller had restarted; here the window is 17 seconds and both halves are in the evidence. **Needs:** the remove path to refuse, or to wait, while a restore for the same stack is in flight — and, either way, to verify the teardown rather than report success without looking. Evidence: `audits/the-28-2026-09-22/apps/gokapi/came-back-evidence.txt`, `apps/gokapi/log.txt`. **FIXED in v0.262.0, both halves.** **(a) The guard the product already had everywhere else:** `RemoveStack` now consults `UpdateGuards.Busy` (backup single-flight, restore status, app-data holds) and `IsUpdating`, and refuses with the app's own sentence in both languages. The product refused exactly this clash for `update` and for `restore` — the restore refusal even NAMES the blocking app — and `remove` was the one door with no lock on it. **(b) `down` returning 0 is a request, not a result:** the compose project is now watched for 25 s after `down`, anything carrying its label is removed BY NAME with its labels logged, and the API answer carries `verified: true/false`. That is the instrument R-626 asked for, in the product rather than in a drill script. **PROVEN LIVE ON 9202, and the live proof found a defect a test had not.** The guard fires: with `privatebin` STOPPED (so the pre-existing "still running" check could not answer first) and a box-wide backup in flight, the remove was refused and the controller said `RemoveStack privatebin REFUSED (busy): a backup or restore is running (single-flight held)`, with the household's own sentence *„Az alkalmazáson mentés vagy visszaállítás fut. Várd meg, amíg befejeződik.”* **But it answered HTTP 500.** `router.go`'s status mapping greps the error TEXT for `not deployed` / `still running` / `not found` / `protected`, and the busy sentence contains none of them, so it fell through to the default. **A 500 tells the UI something broke; this is “wait a moment”.** Fixed in **v0.262.1**: a typed `*stacks.RemoveBusyError` matched with `errors.As`, answered **409**, carrying both the Hungarian bytes and the key — and its test asserts the sentence contains none of the words the text mapping greps for, so the type is load-bearing rather than decorative. **The verification half is proven too:** a permitted remove answered `verified: true` with no reappearance, and no container carrying the project label existed 60 s later. **AND A FIRST ATTEMPT THAT PROVED NOTHING, recorded because it nearly went down as a pass:** the first run refused the remove with `409 still running — stop it first`, which is the PRE-EXISTING running check, not this guard — the restore had already finished by then. A refusal from the wrong rule is not evidence for the new one. | **CLOSED 2026-09-22 — v0.262.0 + v0.262.1; refusal and verification both proven live, and the live proof corrected the status code — **re-proven on v0.262.1: HTTP 409 with the household's sentence**, `RemoveStack privatebin REFUSED (busy)`** | | **R-635** | **[P1-HIGH] `romm 5.3.0` does not fit the memory the template gives it, and the guarded Update called that a success — the app has been OOM-crash-looping on demo-hp for six hours at ~500% CPU.** FOUND 2026-09-22 17:37 **because the operator heard the fans**, which is the only reason it was found at all. `romm` was promoted `5.0.0 -> 5.3.0` on the live catalog that morning (`audits/PROBE-FIX-2026-09-22.md`) after the edge was PROVEN on scratch guest 9202, and the guarded Update was then pressed on demo-hp guest 9201 at **09:08:22Z**, reaching **`done` in 74.8 s** with the app `running`. **It ran clean for two hours.** The first worker kill is at **11:09:20Z**; by 15:38Z there had been **4,530** of them — `Worker (pid:…) was sent SIGKILL! Perhaps out of memory?` — with `docker inspect` reading **`OOMKilled: true`**, the container pinned at **457 MiB of its 512 MiB limit**, and `docker stats` showing **499.51% CPU**. The host's load average sat at **5.2 while otherwise idle**. `rq_cron` is killed and restarted every few seconds in a permanent storm. **THREE THINGS THIS ESTABLISHES, and the third is the one that changes how promotions are judged.** (1) The template's own comment says *`RAM: ~300MB (mem_limit: 1024M total — romm 512M + mariadb 384M + redis 128M)`*; **5.3.0 needs more than 512M and the template was not re-sized when the version moved.** A version move is not only an `image:` line. (2) **The update reported `done` and the app reads `running`**, because nginx answers `GET /` with 200 while the gunicorn workers behind it are being killed — a THIRD variant of the R-618/R-630 theme: the probe is right, the port is right, and the answer is still a false green. (3) **THE ALARM DID FIRE, AND MY FIRST WRITE-UP OF THIS ROW SAID IT DID NOT — CORRECTED 2026-09-22 BY THE OPERATOR, WHO PRODUCED THE MAILS.** The controller HAS an OOM detector (`main.go:821`, `[deadapp] romm: container romm was OOM-killed`, evaluated every 30 s), it emits **`app_oom`** (`notifier.go:726`), the hub allow-lists it (`dispatcher.go:636`) and dispatched it to the OPERATOR channel as **SENT** — `admin@felhom.eu` received **`[Felhom] ⚠️ demo-hp: app_oom`** at **11:09 CEST** and again at **17:48 CEST**, each carrying the app name, the Hungarian sentence and the dashboard link. The CUSTOMER channel is **SKIPPED**, which is correct. **I asserted an absence without looking at the instrument** — the hub's own Events and Notifications tabs show all of it — and I did it by reasoning from a memory note (`lxc-docker-oom-signals-unreliable`, R-528) instead of reading the hub. **That is R-628's shape exactly, four days on, and from the same hand.** **WHAT IS ACTUALLY WRONG, and it is narrower and real:** `notifier.go:715-726` keys the alarm on `container|startedAt` in an `oomSeen` map and emits **once per container lifetime**. So **six hours of continuous thrashing — 4,530 worker kills — produced exactly ONE mail**, severity `warning`, never escalating, while the app went on reading `running`. **A six-hour storm is indistinguishable from a single transient kill.** The hub's App Telemetry panel did carry the magnitude (RomM: **5,023 errors, 632 warnings**, peak 855 MB) but nothing turns that into a second, louder signal. So the fix worth having is not a detector — there is one — but an ESCALATION: a warning that repeats for hours should stop looking like a warning that happened once. **AND THE LESSON FOR R-462's METHOD, which is the real cost:** every `proven` verdict in the update night and in the twenty-eight measures the app for the **minutes of the walk**, not for a day of running. `romm` passed its walk, was seeded, read back and restored — and broke two hours later. **`proven` currently means "the update applied and the data survived", NOT "the new version runs".** Needs: decide between rolling the catalog back to 5.0.0 and raising romm's `mem_limit` (measured, not guessed); and a soak longer than a walk before any future promotion. **FIXED AND MEASURED 2026-09-22, in two steps, and the FIRST step was still a guess.** *Step 1 (operator's choice):* the limit was raised 512M → **768M** (`app-catalog-felhom.eu@886956d`). It slowed the kills from ~12/min to ~7/min and **stopped nothing** — 37 SIGKILLs in five minutes, `OOMKilled` still true, and the cgroup's own `memory.peak` read **exactly 768 MiB**: it hit the new ceiling and died there. *Step 2, from a MEASUREMENT instead:* the per-process RSS inside the container reads **~216 MiB per warm uvicorn worker**, so the image's default of four workers plus the master needs **~882 MiB** before nginx and the job runner — more than any sensible limit for this box. **The lever was in the image all along:** `/init:143` runs `--workers "${WEB_SERVER_CONCURRENCY:-4}"`. **Four workers is a SERVER default on an appliance serving one household.** Setting **`WEB_SERVER_CONCURRENCY=2`** (`app-catalog-felhom.eu@f4eb94f`, limit left at 768M) and applying it through the product's own Update button fixed it. **PROVEN UNDER LOAD, not just at idle:** 6 concurrent callers driven at romm through the household's own route for 300 s — **26,645 requests** (9,687 × 200, 16,958 × 401 on the auth-gated endpoints), CPU a steady **~200%** (exactly two workers saturated, by design), memory oscillating **416–614 MiB against the 768 MiB limit and trending DOWN**, and **zero** SIGKILLs, `OOMKilled: false`, `RestartCount: 0` throughout. At idle afterwards: **1.64% CPU**, 610 MiB, host load falling from 5.2 to 2.4. **AND THE FIRST SOAK MEASURED NOTHING, which is worth more than the second one:** it was pointed at `arcade.enkisfelhom.hu` — the scratch-guest fixture's default subdomain — while this box deployed romm at `jatek`. Every request 404'd at traefik in 9 ms, romm sat idle at 0.64% CPU, and the counter cheerfully reported **14,026 successful requests**. It was caught only because 0.64% CPU under load is not believable. **A positive AND a negative control are now asserted before any load is driven** (`jatek` must not 404; a nonsense host must). R-96 rule 3, in a new surface. **WHAT STAYS OPEN, and it is the part that outlives romm:** 610 MiB of 768 MiB is **79%** — it works with ~158 MiB of headroom and the soak never exceeded 614 MiB, but it is not generous, and nothing watches it. **And the method lesson for R-462:** a version move is not only an `image:` line — the new version's SHAPE (worker counts, per-worker footprint) has to be measured too, and a walk lasting minutes cannot see a ceiling reached in two hours. Every `proven` verdict in the update night and in the twenty-eight means *"the update applied and the data survived"*, **not** *"the new version runs"*. Evidence: `audits/probe-fix-2026-09-22/romm-soak.json`, `romm-soak.out`. | **CLOSED 2026-09-22 — two workers, 768M, proven under 26,645 requests; the 79% headroom and the `proven`-means-minutes lesson are carried into R-462** | | **R-638** | **[P2-MEDIUM] The product's database loader cannot replay a copy over a NEWER schema: on PostgreSQL it FAILS, on MariaDB it leaves the newer version's tables behind.** MEASURED 2026-09-23 on 9202. `ImportDump` (`appbackup/dbdump.go:719`, `psql -v ON_ERROR_STOP=1 --single-transaction`) replays a `pg_dump --clean --if-exists` file over the live database. After docmost 0.95.0 → 0.96.0 migrated, the replay of the pre-update copy was refused in 0.40 s, rc 3: *cannot drop constraint workspaces_pkey on table public.workspaces because other objects depend on it / DETAIL: constraint oauth_clients_workspace_id_fkey …* — the new version created six tables whose foreign keys point at old ones, and `--clean` only drops what the dump knows. Database unchanged (the transaction rolled back). On MariaDB (`mariadb-dump`, `FOREIGN_KEY_CHECKS=0`) the same replay after romm 5.0.0 → 5.3.0 returned rc 0 in 1.25 s and left **12 base tables** of the new version behind; RomM 5.0.0 happened to ignore them. **What worked:** `DROP SCHEMA public CASCADE; CREATE SCHEMA public;` + the dump in ONE transaction — rc 0 in 1.38 s, every table, index and extension back. **Why this is a row of its own and not only part of R-637:** the SAME loader backs shipped paths — `rollbackSafetyDump` (off-site restore's undo) and the dump replay of the restores — so **any restore of a copy taken BEFORE an update that migrated, replayed over the migrated database, may fail the same way. NOT MEASURED:** whether the unit restore the hold sentence names does this (it also carries the data VOLUME tar, which may make the replay moot). That is the measurement owed, on 9202, before anyone relies on it. Evidence: `audits/update-rulings-2026-09-23/README.md` Part 1, `docmost-45`, `romm-44`. **-- NARROWED 2026-09-23:** the undo no longer touches this loader — it copies folders (decision 19, controller v0.263.0). **What stays open is the part about SHIPPED paths:** `rollbackSafetyDump` and the restores' dump replay still replay over whatever schema is live, and whether the unit restore the hold sentence names works after a real schema migration is STILL UNMEASURED. | **OPEN — P2, narrowed to the restore paths; owner: CC; measure the named restore after a real schema migration first** | -| **R-643** | **[P2-MEDIUM] The ruled chain leaves the automatic update leg AT MOST 15 MINUTES a night.** FOUND 2026-09-23 while writing the build plan for `09` §3 decision 11 (*updates after the off-site copy, before the full-system backup*). The off-site leg starts at W+105m (`cmd/controller/main.go:961`) and the full-system backup's gate opens at W+2h (`quiesce/quiesce.go:656`, span to W+6h); the legs are clock-scheduled, not chained. One step takes ~1 min when it works and ~2–6 min when it fails and is undone. Options and the recommendation (the full-system backup waits for the leg inside its own window; the leg stops starting steps at W+5h) are in `09` §6.4. **-- RULED 2026-09-23 (`09` §3 decision 20):** the full-system backup waits for the update leg inside its own window; the leg stops starting new steps at W+5h. Built with `09` §6.4 part 7. | **RULED — build with §6.4 part 7; owner: CC** | | **R-644** | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** | | **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** | | **R-652** | **[P3-LOW] The memory watch counted the kernel's file cache as the app's memory.** MEASURED 2026-09-23 night on the bench: nextcloud 34.0.4 read **100 %** of its 1 GiB and immich's PostgreSQL **100 %**, each with **0** kernel `oom_kill`s — `memory.peak` includes page cache, which the kernel drops before it kills anything. Under the watch as built (R-635 follow-up) both would be marked `memory_tight`, and the gate would demand a raised `mem_limit` — a customer-box capacity figure — for cache. **Done the same night (09 §3 decision 22, CC — operator may reverse):** the watch samples the app's own memory (`anon` of `memory.stat`) every 15 s; the mark and the ladder's `memory_peak_pct` read it where measured; the cgroup peak stays beside it (`memory_cgroup_peak_pct`). **Open:** romm's backfilled entry carries M1's 80.9 % cgroup peak (measured before the anon sample existed) — re-measure it on its next step; and decide whether an app whose anon is low but whose cgroup stays pinned at its limit (cache thrash) should be marked at all. Evidence: `audits/night-2026-09-23/apps/nextcloud/bench-1024M/`, `apps/immich/bench-noanon/`. | **READY — P3; owner: CC (catalog harness)** | @@ -815,7 +814,7 @@ class (an image `VOLUME` at an unmounted path) is still live — `immich-server` | **R-683** | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | | **R-685** | **[P2-MEDIUM] A box whose whole-box backup cannot fit must say so BEFORE the night — an operator event and a line on the backup page — never only a nightly failure a log shows.** The product lesson of R-684 (operator ruling 2026-09-24 evening, option A for demo-hp): demo-hp's root-disk target held three 6–7 GB archives with ~4 GB free and failed `No space left on device` every night from 2026-09-23 while nothing but the vzdump log said why. Measured 2026-09-24 night: a 9201 archive is 8.18 GB (22.6 GB uncompressed); PVE prunes AFTER a successful backup, so a target must hold keep-last + 1 archives during the run — lowering the retention alone does not un-stick a full target. **Fix direction:** the agent predicts the archive size (last archive × margin) against the target's free space before a whole-box backup, skips with a reason reported to the hub (operator event), and the controller's backup page shows the sentence. `audits/night-2026-09-25/A/A4-hp-backup-space.txt` | **READY — P2; owner: CC (agent + controller)** | | **R-686** | **[P3-LOW] The automatic update leg is not resumed after a controller restart during the night — the apps it had not reached wait a whole day.** MEASURED 2026-09-24 night on 9202 (v0.271.0, night 3): the controller was killed (kill -9) during romm's step; the step itself was put back correctly (pin back, the page says the update was interrupted), but vikunja — next in line, re-tested and ready — was not pressed that night. **Night 4 (power cut during romm's verify) showed a second effect:** the interrupted step was RESUMED after the boot and ended `done` (data read back), but the leg that pressed it was gone, so `last_auto_update` was never written and romm's page carries no „Automatikus frissítés … — sikeres" line for a step the box did take by itself. By design today: the leg is one call of the `offbox-backup` job, and a Daily job that already fired does not fire again. The controller's own self-update now waits for the leg (decision 32), so the common restart cause is excluded; a crash or a power cut still ends the night's leg. **Fix direction (needs a ruling):** persist "leg started, not finished, window W" and resume at boot while before W+5h — or accept one night's delay. `audits/night-2026-09-25/C/night3-kill/` | **OPEN — P3; owner: CC (operator ruling on resume)** | -| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent) — it is Part D's to show on the demo boxes. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` | **OPEN — P3; owner: CC** | +| **R-687** | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` | **OPEN — P3; owner: CC** |