From bfdea832046c296e35ec8fa5deac2a24eb4a622b Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 6 Oct 2026 12:24:43 +0200 Subject: [PATCH] ten answers: golden 0.300.0 baked (pinned), delivery evidence; R-645 R-856 R-747 R-774 R-734 R-624 R-502 R-99 closed; R-890/R-891 filed; 03 gains FELHOM_FSTRIM and GET /host/crash-guard (149 -> 143) CI 1423 (929e59e8) was red on golden-currency: it read controller v0.300.0 before this commit recorded its golden. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- documentation/architecture/03-host-agent.md | 5 + .../delivery/controller-and-bundle.txt | 6 + .../delivery/hub-deploy.txt | 9 + .../delivery/r444-sudo-check.txt | 5 + .../delivery/sign-bundle.txt | 14 + .../delivery/vouch-and-sign-agent.txt | 17 + .../delivery/vouch-golden-floors.txt | 11 + documentation/backlog/CLOSED-ITEMS.md | 15 + documentation/backlog/OPEN-ITEMS.md | 10 +- .../02-round-trip.txt | 3 + .../tests/golden-0.300.0-2026-10-06/README.md | 29 ++ .../tests/golden-0.300.0-2026-10-06/bake.log | 337 ++++++++++++++++++ 12 files changed, 453 insertions(+), 8 deletions(-) create mode 100644 documentation/audits/ten-answers-2026-10-06/delivery/controller-and-bundle.txt create mode 100644 documentation/audits/ten-answers-2026-10-06/delivery/hub-deploy.txt create mode 100644 documentation/audits/ten-answers-2026-10-06/delivery/r444-sudo-check.txt create mode 100644 documentation/audits/ten-answers-2026-10-06/delivery/sign-bundle.txt create mode 100644 documentation/audits/ten-answers-2026-10-06/delivery/vouch-and-sign-agent.txt create mode 100644 documentation/audits/ten-answers-2026-10-06/delivery/vouch-golden-floors.txt create mode 100644 documentation/tests/golden-0.300.0-2026-10-06/02-round-trip.txt create mode 100644 documentation/tests/golden-0.300.0-2026-10-06/README.md create mode 100644 documentation/tests/golden-0.300.0-2026-10-06/bake.log diff --git a/documentation/architecture/03-host-agent.md b/documentation/architecture/03-host-agent.md index 84aa8dff..4e48ed57 100644 --- a/documentation/architecture/03-host-agent.md +++ b/documentation/architecture/03-host-agent.md @@ -118,6 +118,7 @@ fixed file, delivered by the signed config bundle): | `FELHOM_OOB` | the OOB firewall sets | `add element` takes exactly `{ [/n] }` or `{ }` — no chained command | — | | `FELHOM_PBSDR` / `FELHOM_BACKUPTARGET` | PBS DR entry; whole-system backup target | unchanged: the arguments stay coarse, and the root wrappers (`felhom-pbs-apply`, `felhom-backup-target-apply`) are the gate (fixed verbs, own validation) | coarse argv into a checking wrapper | | `FELHOM_ESCROW` | the recovery-code ceremony (runs the agent binary as root) | the binary is only ever an operator-signed one (`FELHOM_SELFUPDATE`); as root it pins the PVE secret dir and the WG state dir, refuses a storage id that is a path, and reads its two staged files by walking the path with `openat(O_NOFOLLOW)` (no symlink anywhere) | **by design the agent relays R**, so a compromised agent can still learn this box's PBS key through the ceremony — not root, but the backup key | +| `FELHOM_FSTRIM` | the weekly disk trim of each customer guest (R-444, `09` §3 decision 139; agent v0.149.0) | ONE regex-anchored rule `/usr/sbin/pct ^fstrim [0-9]+$` — no option, no second vmid, no chained command (pinned by `TestSudoersFstrimRuleIsExact`; read live on demo-hp and demo-felhom 2026-10-06: `pct fstrim 9201` allowed, `--ignore-mountpoints`, `;x`, `9201 9202`, `pct destroy` refused) | — | | `FELHOM_SELFHEAL` / `FELHOM_GUESTNET` / `FELHOM_OSAPPLY` | networking restart; guest DHCP watchdog; OS updates | exact; `felhom-os-apply --plan …` stays the glob line on purpose — the bundle's own self-check reads that exact text, and the wrapper refuses any other plan path (R1) | — | **What this does not change.** The operator key (`/etc/felhom/operator-signers`, root-owned, never a bundle path) stays @@ -229,6 +230,10 @@ The controller (in its LXC) reaches the agent (on the host) over the local bridg - `POST /backup` — request a backup-now of *this* guest (enqueued; non-destructive). - `GET /backup/due` — whether a policy-scheduled backup is due for *this* guest, so the controller can quiesce then call `POST /backup` (the app-consistent path, §8). - `GET /backup/status`, `GET /restore-test/status` — read-only status for the controller's UI. + - **Crash-boot fact (R-856, agent v0.149.0):** `GET /host/crash-guard` — the host crash guard's last-boot record + (`present`, `last_boot_at`, `last_boot_unclean`, `tripped`) read from `/var/lib/felhom-crash-guard/state.json`; a missing + or garbled file answers 200 `present:false`, an older agent 404 — both read as a normal boot. The controller waits + ~15 min with app mails after a crash boot (`09` §3 decision 143). - **Host metrics (slice 9):** `GET /host/metrics` — **host-wide** health for the customer's monitoring view: cpu%/mem/load/uptime, **CPU/chassis temperature** (`cpu_temp_c`, nullable — "n/a" when the hardware exposes no sensor), and per-storage capacity (total/used/fraction, diff --git a/documentation/audits/ten-answers-2026-10-06/delivery/controller-and-bundle.txt b/documentation/audits/ten-answers-2026-10-06/delivery/controller-and-bundle.txt new file mode 100644 index 00000000..4c5d3901 --- /dev/null +++ b/documentation/audits/ten-answers-2026-10-06/delivery/controller-and-bundle.txt @@ -0,0 +1,6 @@ +== controller delivery 2026-10-06T10:23:27Z +demo-hp + demo-felhom felhom-controller:0.300.0 (healthy) 10:22:07Z; tester-1 hub page 'Controller elindult (0.300.0)' +R-856 live read (both NORMAL branches, through the real agent route): + demo-hp: [deadapp] boot grace 1m30s: the host's last unclean boot (2026-10-05T07:56:41Z) is not the one this start followed (R-856) + demo-felhom: [deadapp] boot grace 1m30s: the host's last boot was clean (R-856) +bundle e182c82d on demo-hp (BUNDLE DONE 12:06:31 local, capability probe 68/68) and demo-felhom (config-bundle.json); sudo -l on demo-felhom: fstrim allowed, destroy refused diff --git a/documentation/audits/ten-answers-2026-10-06/delivery/hub-deploy.txt b/documentation/audits/ten-answers-2026-10-06/delivery/hub-deploy.txt new file mode 100644 index 00000000..ecb389dc --- /dev/null +++ b/documentation/audits/ten-answers-2026-10-06/delivery/hub-deploy.txt @@ -0,0 +1,9 @@ +== 2026-10-06T09:55:58Z +healthz 200 +system 200 +image=gitea.dooplex.hu/admin/felhom-hub:0.139.0 +sync=Synced health=Healthy rev=929e59e8f4fdf1163313729113ba29f145854597 +2026/10/06 11:55:10 [INFO] felhom-hub 0.139.0 starting +2026/10/06 11:55:10 [INFO] off-site secrets sealed at rest (0 legacy plaintext row(s) sealed now) +2026/10/06 11:55:10 [INFO] console passwords sealed at rest (0 legacy plaintext row(s) sealed now) +2026/10/06 11:55:10 [INFO] box secrets sealed at rest (0 legacy plaintext value(s) sealed now) diff --git a/documentation/audits/ten-answers-2026-10-06/delivery/r444-sudo-check.txt b/documentation/audits/ten-answers-2026-10-06/delivery/r444-sudo-check.txt new file mode 100644 index 00000000..cb829e77 --- /dev/null +++ b/documentation/audits/ten-answers-2026-10-06/delivery/r444-sudo-check.txt @@ -0,0 +1,5 @@ +pct fstrim 9201: ALLOWED +pct fstrim 9201 --ignore-mountpoints: refused +pct fstrim 9201;x: refused +pct fstrim 9201 9202: refused +pct destroy 9201: refused diff --git a/documentation/audits/ten-answers-2026-10-06/delivery/sign-bundle.txt b/documentation/audits/ten-answers-2026-10-06/delivery/sign-bundle.txt new file mode 100644 index 00000000..c766c7d7 --- /dev/null +++ b/documentation/audits/ten-answers-2026-10-06/delivery/sign-bundle.txt @@ -0,0 +1,14 @@ +== System/hosts 2026-10-06T09:56:47Z: demo-hp, demo-felhom, tester-1 agent 0.149.0; Tester-2 0.142.0 (nothing sent) +== agent_config_update 0.149.0 (bundle e182c82d…), 2026-10-06T09:56:47Z +-- demo-hp-bb76ea +signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=58ab3d41a9318bfac142372f6362c2c0 expires=2026-10-06T10:41:47Z +wrote envelope to /env-demo-hp-bb76ea-acu.json +uploaded signed op to the hub jobs queue +-- demo-felhom-8363b5 +signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=0fb920820f8499349e7f098f3008295f expires=2026-10-06T10:41:47Z +wrote envelope to /env-demo-felhom-8363b5-acu.json +uploaded signed op to the hub jobs queue +-- tester-1-d70be4 +signed: op=agent_config_update host=tester-1-d70be4 guest="" key_id=felhom-op-1 nonce=8033f0f0d70cf5a5f7b3eb74aa0efc01 expires=2026-10-06T10:41:47Z +wrote envelope to /env-tester-1-d70be4-acu.json +uploaded signed op to the hub jobs queue diff --git a/documentation/audits/ten-answers-2026-10-06/delivery/vouch-and-sign-agent.txt b/documentation/audits/ten-answers-2026-10-06/delivery/vouch-and-sign-agent.txt new file mode 100644 index 00000000..c165b6af --- /dev/null +++ b/documentation/audits/ten-answers-2026-10-06/delivery/vouch-and-sign-agent.txt @@ -0,0 +1,17 @@ +== vouch 2026-10-06T09:50:38Z: agent 0.149.0, golden 0.299.0, min_agent 0.131.0 +HTTP/1.1 303 See Other +Location: /configuration?flash=artifacts_set +2026/10/06 11:51:08 [INFO] Artifact manifest set: agent=0.149.0 golden=0.299.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="e182c82dcf4a67faa3bcb74dbe4ffa7b06e0b27dc8451cb7574d6339ce91ad66" +== agent_update 0.149.0 (sha 6bcae9c2…) signed with felhom-op-1, ttl 45m, 2026-10-06T09:51:11Z +-- demo-hp-bb76ea +signed: op=agent_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=07c59d5479139b7bc3807b977eb37729 expires=2026-10-06T10:36:11Z +wrote envelope to /env-demo-hp-bb76ea-au.json +uploaded signed op to the hub jobs queue +-- demo-felhom-8363b5 +signed: op=agent_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=9ba373bce19019bc3a76ecb3b429c666 expires=2026-10-06T10:36:11Z +wrote envelope to /env-demo-felhom-8363b5-au.json +uploaded signed op to the hub jobs queue +-- tester-1-d70be4 +signed: op=agent_update host=tester-1-d70be4 guest="" key_id=felhom-op-1 nonce=7ca22f2ed83309082d779840275822da expires=2026-10-06T10:36:11Z +wrote envelope to /env-tester-1-d70be4-au.json +uploaded signed op to the hub jobs queue diff --git a/documentation/audits/ten-answers-2026-10-06/delivery/vouch-golden-floors.txt b/documentation/audits/ten-answers-2026-10-06/delivery/vouch-golden-floors.txt new file mode 100644 index 00000000..5339063c --- /dev/null +++ b/documentation/audits/ten-answers-2026-10-06/delivery/vouch-golden-floors.txt @@ -0,0 +1,11 @@ +== vouch 2026-10-06T10:21:01Z: agent 0.149.0, golden 0.300.0, min_agent 0.131.0 +HTTP/1.1 303 See Other +Location: /configuration?flash=artifacts_set +== floors 2026-10-06T10:21:30Z: 0.300.0 / min_agent 0.131.0 +demo-hp: Location: /customers/demo-hp?flash=floor_set +demo-felhom: Location: /customers/demo-felhom?flash=floor_set +tester-1: Location: /customers/tester-1?flash=floor_set +2026/10/06 12:21:30 [INFO] Artifact manifest set: agent=0.149.0 golden=0.300.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="e182c82dcf4a67faa3bcb74dbe4ffa7b06e0b27dc8451cb7574d6339ce91ad66" +2026/10/06 12:21:31 [INFO] Customer demo-hp controller-version floor override set to "0.300.0" (declared MinAgent "0.131.0") +2026/10/06 12:21:31 [INFO] Customer demo-felhom controller-version floor override set to "0.300.0" (declared MinAgent "0.131.0") +2026/10/06 12:21:31 [INFO] Customer tester-1 controller-version floor override set to "0.300.0" (declared MinAgent "0.131.0") diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 58363b96..830ade8e 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,21 @@ --- +## 2026-10-06 (midday) — the operator's ten answers built + +The full text of every row below: `git show 929e59e8:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-774** | **[P3-LOW] Two things the new apps' pages do not show yet: Karakeep's mail-ON path is unproven, and its official phone app reports crashes to its makers.** (P3) | CLOSED 2026-10-06 — BUILT AND DELIVERED (operator ruling, `09` §3 decision 148): Karakeep's page says its phone app sends crash reports | catalog `ec72c9d` (live `1938921`): hu + en. The row's other half — one password-reset mail from Karakeep on a hub-enabled box — is NOT covered by this ruling; it is not built. | +| **R-734** | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** (P4) | CLOSED 2026-10-06 — BUILT AND DELIVERED (operator ruling, `09` §3 decision 145): the update test ignores listed marker files, each with a reason | catalog `b0939cf`: `scripts/upgrade-test.py` per-app list, immich's six `.immich` markers first; a listed file is ignored only when changed/added and ≤ 64 bytes; verdict records `files_ignored`; harness v5. `MarkerIgnore` tests on the measured immich lists; red-proof: an ignore-all mutant fails two tests. | +| **R-645** | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** (P3) | CLOSED 2026-10-06 — BUILT AND DELIVERED (operator ruling, `09` §3 decision 142): the night backup skips an app that runs another version than it saved | controller `2d63714` (v0.300.0): every night leg (DB dump, volume dump, unit capture, Tier 2) skips an app whose pin is not what it runs; one amber line on the backups page (`backup.status.version_skip`, hu + en); unknown never skips. `TestR645_HandLiftedHoldKeepsTheGoodUnit` runs the night + Tier 2 on the hand-lift shape (unit checksum unchanged); red-proof: without the skip the unit was rewritten with `docmost:0.96.0`. Delivered 10:22Z to the three boxes. **Residual, stated:** after the lift the boot reconciler may START the app on the new version, and then pin = running — the ruling's predicate protects the window between the lift and that start (the measured overwrite came 3 s after the restart); covering the started-new-version case is a design question, not built. | +| **R-99** | Server-side prune **never removes** a phantom snapshot. Confirmed it does NOT count them toward `keep-last` (dry-run kept 2 real + the phantom) so there is **no retention/data-loss bug** — but one acc (P4) | CLOSED 2026-10-06 — BUILT AND DELIVERED (operator ruling, `09` §3 decision 140): phantom leftovers are deleted by a runbook — none exist today | `runbooks/pbs-phantom-cleanup.md` + `pbs-phantom-list.py`; read-only listing of ep0 2026-10-06: 9 snapshots in 5 namespaces, all ≥ 369,808,250 B, all verification `ok`, every directory has its manifest — **no phantom, nothing deleted**, real counts unchanged (`audits/ten-answers-2026-10-06/r99-ep0-listing.txt`). The agent's WARN for a phantom now names the runbook (agent `be398f9`, v0.149.0, `TestRejectedArchiveWarnNamesTheCleanupRunbook`). | +| **R-747** | **[P3-LOW] A stranger can lock the household out of mealie with five wrong logins.** (P3) | CLOSED 2026-10-06 — BUILT AND DELIVERED (operator ruling, `09` §3 decision 144): mealie's page says five wrong logins lock the account for 1–2 hours | catalog `ec72c9d` (live `1938921`): a new last first step, hu + en, informal. The 1–2 h is true where mealie runs with the one-hour lock setting; an already-installed mealie gets it when its compose is rendered again (still unmeasured, as the row said). | +| **R-856** | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** (P4) | CLOSED 2026-10-06 — BUILT AND DELIVERED (operator ruling, `09` §3 decision 143): app mails wait ~15 minutes after a crash boot | controller `c393d85` (v0.300.0, `internal/crashboot`) + agent `f277e61` route `GET /host/crash-guard` (v0.149.0). Tests + red-proofs both halves (crash boot holds the mails at +3m30s; a normal boot keeps 90 s; an agent 404 = normal). **Live, through the real route (`audits/ten-answers-2026-10-06/delivery/controller-and-bundle.txt`):** demo-hp logged „boot grace 1m30s: the host's last unclean boot (2026-10-05T07:56:41Z) is not the one this start followed", demo-felhom „… the host's last boot was clean". The crash branch itself was not shown live (no crash allowed by the brief). | +| **R-502** | **[P3-LOW] The bootstrap regression harness is run by NO gate and NO CI — and it had never exercised the pairing banner.** (P4) | CLOSED 2026-10-06 — BUILT AND DELIVERED (operator ruling, `09` §3 decision 147): the ISO first-boot test is a gate, full runs only | felhom.eu `9d39faab`: `scripts/iso_bootstrap_gate.py`, fast=False (CI and the pre-push hook call --fast, so never), NOT CHECKED (exit 2) without docker or the image; a built-in decoy every full run; 9 docker-free decoy tests. **First real full run on DooPlex 2026-10-06** (after building `felhom-iso-assistant:trixie`): 73 harness checks green, the built-in decoy convicted (`audits/ten-answers-2026-10-06/r502-first-full-run.txt`). | +| **R-624** | **[P3-LOW] Three of the catalog's apps cannot be seeded by ANY headless route, and for two of them that is a deliberate security decision — so the upgrade harness has a permanent ceiling nobody has written down.** (P4) | CLOSED 2026-10-06 — BUILT AND DELIVERED (operator ruling, `09` §3 decision 146): the bench may seed vaultwarden through its admin route, bench only | catalog `aed80ee`. **Proven on the recreated bench 9401 (2026-10-06):** admin sign-in 200, invite 200, invited registration 200, seed read back before AND after the move; `.env` shredded, 64-hex grep 0 (control 1), `VW_ADMIN=` 0; WITHOUT the run flag the admin route is not tried and the verdict is `inconclusive` (`audits/ten-answers-2026-10-06/r624-bench/`). The move used (1.36.0-alpine → 1.36.0) failed on health — the non-alpine image fails the alpine health check; that is the test target, not the seed. Zipline needs no held secret (its `/api/setup`); its redaction is covered by `SecretHygiene`. **New question filed as R-890:** the ladder writer needs both venues, so a vaultwarden step still cannot be written. | + ## 2026-10-06 (midday) — the ten answers The full text of every row below: `git show 8c65ff0c:documentation/backlog/OPEN-ITEMS.md`. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 9fc5f07b..c60601b3 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -130,7 +130,6 @@ stopping line that lies. | **R-612** | Apps & catalog | P3 | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** **— FIXED 2026-09-23 night (catalog `a5a729a`): 512M.** Measured on the bench: the seed peaks at 312–345M and was OOM-killed at 128M on every run (kernel `oom_kill` 3); running, the app sits at ~115M (90 % of the old limit alone). On 9202 at 128M the kill came AFTER the Role/Group rows this time, so the sign-up lie did NOT reproduce — the kill is timing-dependent, the fault is memory (the brief's claim held). At 512M the seed completes (`The seed command has been executed`), and wishlist then moved v0.66.0 → v0.67.1 with its test record. **Still open:** a failed first-boot seed is invisible to the box — nothing reads the seed's exit. | **READY — P3, narrowed to making a failed seed visible; owner: CC (catalog)** **Re-ranked 2026-10-03: P1->P3: the memory fix shipped (catalog a5a729a); only the silent-seed detection remains.** | — | — | CC | | **R-676** | Apps & catalog | P3 | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. | **READY — rank P2-MEDIUM; owner: CC (catalog)** **Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again.** | — | — | CC | -| **R-774** | Apps & catalog | P3 | **[P3-LOW] Two things the new apps' pages do not show yet: Karakeep's mail-ON path is unproven, and its official phone app reports crashes to its makers.** READ/MEASURED 2026-10-01: Karakeep's `smtp_mapping` (plaintext :2526, as Cal.com) was proven only with mail OFF (a fresh install boots) — 9202 has no hub, so the relay cannot be exercised there; the official mobile app ships Sentry crash reporting with a hard-coded DSN (`apps/mobile/app/_layout.tsx`, FIT.md). **Needs:** one password-reset mail from Karakeep on a hub-enabled box (demo-hp 9201, a throwaway install); and a sentence on the page about the phone app (copy, freeze). `audits/new-apps-2026-10-01/bench/karakeep-mail-off-boot.txt` | **READY — rank P3-LOW; owner: CC (catalog)** **RULED 2026-10-06 10:41 (`09` §3 decision 148): A — one sentence on Karakeep's page about the phone app's crash reports.** Being built. | — | — | CC | | **R-76** | Apps & catalog | P4 | **FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state `idea (surfaced by the R-75 spike, 2026-07-26)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Two related findings from `audits/SPIKE-catalog-data-paths-2026-07-26.md` P3/P5, both **pre-existing** and deliberately left alone by that spike. **(a)** FileBrowser Quantum 1.3.3 creates files `0644` and folders `0755` and does **not** propagate the setgid bit — even though the entrypoint wrapper's `umask 002` really is in effect (`/proc/1/status` `Umask: 0002`). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's *group* half holds and only its *mode* half is lost. The consequence is proven with a control: inside a UI-created `0755` folder a gid-1000 process's file landed group **1000**, while the identical write into the 2775 parent landed group **100**. So **any folder a customer creates through FileBrowser breaks the shared-group chain one level down.** Latent today — every userdata-touching catalog app that declares an identity declares uid/gid **1000**, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at `infra/infra.go:156` is right that the image ignores `-e UMASK` but does not say t | CC | | **R-577** | Apps & catalog | P4 | **[P3-LOW] A guest SHARE visitor has no way to pick a language, and the household's setting is the wrong default for them.** FOUND 2026-09-18 by localisation slice 2 release C (R-557, controller v0.254.0): every other page a person can reach now carries a language globe — the dashboard (the household's setting), and the sign-in and claim pages (the visitor's own cookie). The two guest share pages (`launcher_shared`, `launcher_share_password`) deliberately do NOT, and `TestGuestSharePagesHaveNoGlobe` pins that so it stays a decision rather than an oversight. **Why it is the operator's and not CC's:** a share visitor is a stranger the household sent a link to, and what language they are shown is a promise the SHARE FEATURE makes, not an implementation detail. The `felhom_lang` cookie already built would fit them exactly (display-only, their own browser, never the household's setting). **Fix shape, if the operator says yes:** add `{{template "lang_globe" .}}` to both shells with the anonymous form, and one render case per page per language. | **READY - rank P3-LOW; owner: operator (the decision), CC (the change)** **Re-ranked 2026-10-03: P3->P4: feature decision for the operator; Hungarian default works today.** | — | — | operator | | **R-644** | Apps & catalog | P4 | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** **Re-ranked 2026-10-03: P3→P4: seen only on a scratch test guest; the row itself calls it not a customer fact.** | — | — | CC | @@ -149,7 +148,7 @@ stopping line that lies. | **R-469** | App updates | P3 | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. **-- NARROWED 2026-09-25 (evening), catalog `6a4a5f0`, `09` §3 decision 35:** the PostgreSQL half now passes ONE app at a time — only a template whose ladder entry for the step is proven on BOTH venues and carries `engine_conversion` (the box converts it, controller v0.273.0), as the only image move in its commit. Every other PostgreSQL app stays refused; the postgis family is judged now (it was not). CLAUDE.md rule text updated the same commit. Decoys + red-proof: `audits/night-2026-09-26/C/`. | **NARROWED** — **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** | — | — | CC | | **R-683** | App updates | P3 | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-785** | App updates | P3 | **[P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24).** READ 2026-10-01 (`audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt`). **Needs:** an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. | **OPEN — rank P3-LOW; owner: CC (after R-784)** | — | — | CC | -| **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** **RULED 2026-10-06 10:41 (`09` §3 decision 145): A — a per-app list of marker files the update test ignores, each with a reason.** Being built. | — | — | CC + operator | +| **R-890** | App updates | P4 | **A vaultwarden update step still cannot be written to the ladder: the ladder writer needs the step proven on BOTH the bench and the box, and decision 146 keeps the admin seed on the bench only.** Found 2026-10-06 building R-624: the bench now seeds vaultwarden through its admin invite (proven on bench 9401), so the bench answers abort and memory; the box walk (scratch 9202) still cannot seed it, so no vaultwarden step is ever written. | **OPEN — filed 2026-10-06** | — | Operator: (A) the box walk may use the admin token on scratch 9202, or (B) the writer accepts a vaultwarden entry proven on the bench alone. If nothing: vaultwarden updates stay unwritten (by hand only) | operator | ## Backup & restore — 34 rows (P2 8, P3 11, P4 15) @@ -172,10 +171,8 @@ stopping line that lies. | **R-433** | Backup & restore | P3 | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01.** The one question that can move this is drafted and ready to send: **Question 1** of `documentation/runbooks/provider-questions-2026-09-01.md` — *can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?* **Neither answer leaves this row where it is:** "yes" makes per-file recovery real but permanently operator-only, and closes this; "no, full restore only" means the snapshots do not bound a single customer's exposure at all, because using them costs every other customer on the box their newer snapshots — **and R-95 becomes urgent.** Nothing here can progress without it, and it is not CC's to send (§11-D). Tracked by a dated check in the DUE-CHECKS block. **⚠ A TRAP ON THE WAY TO THIS ANSWER, found by the operator 2026-09-01 and armed against in the questions file: HETZNER HAS TWO SIMILARLY-NAMED PRODUCTS WHOSE DOCS SAY OPPOSITE THINGS.** **Storage SHARE** (a managed Nextcloud, NOT us) documents *"Currently, we only support restores for the full backup ZFS snapshot to a specific point in time"* (`docs.hetzner.com/storage/storage-share/faq/backup-snapshot/`). **Storage BOX** (ours) documents the opposite — *"You can download individual files or entire directories as usual"* (`docs.hetzner.com/storage/storage-box/snapshots/`). **A web search for the obvious phrasing surfaces the SHARE page, and it reads like a definitive NO.** Tell them apart by the giveaways: the Share page talks about *Nextcloud's data cache*, a *database dump* and the *konsoleH* interface, and never mentions Storage Box. **Why this is filed and not left as trivia: a support agent could answer Question 1 from the wrong page, and that answer would push this row to the top of the register for no reason.** Question 1 now names the product, quotes the Storage Box line, and says explicitly that we know what the Share FAQ says. **If an answer cites full-snapshot-only, check WHICH PRODUCT it is about before acting on it.** **Nothing in this repository ever leaned on the Share claim** — verified by grep at the time; the only vendor line we cite is the Storage Box one. **RE-DATED 2026-09-15 → 2026-09-22 (due-checks gate):** the operator mailbox read through the Gmail connector (`(from:hetzner OR subject:hetzner OR "Storage Box") after:2026/09/01`) holds no Hetzner reply — one match, our own `offsite_snapshots_dropped` alarm. Whether the two tickets were ever sent is not visible to CC; if they were not, sending them is the operator's act. **-- ANSWERED. Hetzner replied on ticket #2026090103040671; the operator supplied the thread on 2026-09-22 after this session searched for it and WRONGLY reported no answer (see the instrument note below and R-628).** **Q1 - can the MAIN account read individual files and directories out of a specific snapshot, without restoring the whole box?** Hetzner: *"With the main account it should be possible to download files and directories from the snapshot as it is stated in the documentation"*, citing `docs.hetzner.com/storage/storage-box/snapshots#access-to-snapshots`; and separately *"A restore of a snapshot will revert the Storage Box completely to the state it was in when the snapshot was taken."* **So the Storage BOX documentation governs, not the Storage Share FAQ** - which is exactly the trap `provider-questions-2026-09-01.md` warned the reader about, and the answer came back on the right side of it. **File-level snapshot access is a MAIN-ACCOUNT capability; the sub-account cannot do it**, which matches R-433's own measured finding (777,600 names swept from a sub-account, zero hits). **HEDGE, kept because it is in the reply: "should be possible" is not "is", and nobody has yet read a file out of a snapshot from the main account on `u629488`. That is now the measurement this row needs, and it is cheap.** **Q2 - is `--append-only` enforced by Hetzner, or taken from what the client sends?** Hetzner: *"You can use --append-only and it would look like this: `command="rclone serve restic --stdio --append-only path/to/repo" `"*, with `fluix.one/blog/hetzner-restic-append-only/`. **So it is enforced by US, by pinning a FORCED COMMAND to the SSH key in the Storage Box's `authorized_keys`** - not by Hetzner globally, and not by anything the client asks for. A key without that prefix still gets a deleting server. **This is the answer R-95 has been blocked on** and it says append-only IS expressible on a Storage Box, which R-95 records as impossible over SFTP - see R-95 and R-436. **NOT YET MEASURED, and that distinction is the whole of what this row is worth: this is a support answer, not a proof.** What is owed is a real forced-command key on `u629488`, a restic `forget --prune` through it that is REFUSED, and a backup through it that still succeeds. | **BLOCKED** — **BLOCKED-ON-PROVIDER — Question 1 of `runbooks/provider-questions-2026-09-01.md`** | — | — | CC + operator | | **R-540** | Backup & restore | P3 | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** **2026-10-05 (burn-down night): NEEDS A DESIGN (and money).** A selection rule needs a multi-box configuration and a second box. Next: an operator pick of the rule and its trigger (e.g. add box 2 at 70 %). | — | — | CC | | **R-548** | Backup & restore | P3 | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. **-- 2026-09-30, the class on a real box: demo-hp's whole-guest LOCAL tier was refused by the space preflight (R-685) 10 times since 2026-09-27** („local has 12.1 GiB free; the last archive of guest 9201 was 8.9 GiB, so a new one needs about 12.1 GiB") — its newest local archive was 2026-09-29 04:42, 30 h old at the day's read (the weekly PBS tier, last 2026-09-24, is its whole copy meanwhile). The refusal is correct and named; what is missing is that nothing gives the box the room back (the host's `local` holds the 8.9 GiB archive of the only guest plus templates on a 39 GB root). `audits/pg-last-six-2026-09-30/C/`. | **READY — rank P3-LOW; owner: CC** | — | — | CC | -| **R-645** | Backup & restore | P3 | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** **2026-10-05 (burn-down night): NEEDS THE OPERATOR.** Three shapes, none chosen (the CLI refuses to lift an UPDATE hold; the lift also restores the pin; the capture skips an app whose pin is not what it runs) — each changes the operator's tool or the nightly backup. Next: the pick. **RULED 2026-10-06 10:41 (`09` §3 decision 142): A — the night backup skips an app whose saved version is not the one it runs.** Being built. | — | — | CC | | **R-698** | Backup & restore | P3 | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. **-- 2026-09-30 late (decision 53):** a box now keeps only an app's running and previous image; a restore to an older version re-pulls it — as every restore already did. The limit above is unchanged. | **OPEN — P3; owner: operator (a decision), CC measures** **Operator ruling 2026-10-05 18:23: kept OPEN as a known risk to a household's restore; owner the operator; not worked on in the burn-down.** | — | — | operator | | **R-91** | Backup & restore | P4 | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk **Checked from source 2026-10-05 (burn-down round 2):** Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word. | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | -| **R-99** | Backup & restore | P4 | Server-side prune **never removes** a phantom snapshot. Confirmed it does NOT count them toward `keep-last` (dry-run kept 2 real + the phantom) so there is **no retention/data-loss bug** — but one accumulates per aborted upload, forever | READY (S) **2026-10-05 (burn-down night): NEEDS THE OPERATOR.** Removing a phantom means deleting on a customer datastore (the row's own separate ruling), and the prune runs on ep0 (fenced). Next: the ruling. **RULED 2026-10-06 10:41 (`09` §3 decision 140): A — leftovers are deleted, by a runbook, when one is seen.** Being built. **2026-10-06: runbook written** (`runbooks/pbs-phantom-cleanup.md` + `runbooks/pbs-phantom-list.py`). **Read-only listing of ep0 today: NO phantom** — 9 snapshots in 5 namespaces, all ≥ 369,808,250 B, all verification `ok`, every directory has its manifest; nothing deleted (`audits/ten-answers-2026-10-06/r99-ep0-listing.txt`). Left: the agent's WARN names the runbook (agent change, then release). | — | Decide a cleanup path. Deletion on a **customer** datastore is a separate ruling — detection shipped, removal deliberately not automated | CC | | **R-164** | Backup & restore | P4 | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | | **R-213** | Backup & restore | P4 | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC | | **R-246** | Backup & restore | P4 | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor **Folded R-248 2026-10-03** (the same ruling: give `stale_at` a visible, evidence-bearing setter, or retire it). | — | — | operator | @@ -215,7 +212,6 @@ stopping line that lies. | **R-338** | Security & access | P3 | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it **Checked from source 2026-10-05 (burn-down round 2):** nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246). | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-616** | Security & access | P3 | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** **2026-10-05 (burn-down night): FIXED on controller `main`** (`28a5203`; the catalog clone stores no credentials; the token is supplied per fetch; R-615's repo comparison ignores credentials on both sides, so a token never re-clones (`TestR616_TokenSetSameRepoNoRecloneOriginClean`) and a credentialed origin is cleaned at the next pull. The operator's Gitea admin token rotation (the row's second half) is still owed). Ships with the next controller release; close after delivery. **2026-10-06: DELIVERED** in controller v0.298.0 (the clone stores no credentials). Left: the operator's Gitea admin token rotation. | — | — | CC + operator | | **R-717** | Security & access | P3 | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). | **OPEN — P3; owner: CC** | — | — | CC | -| **R-747** | Security & access | P3 | **[P3-LOW] A stranger can lock the household out of mealie with five wrong logins.** MEASURED 2026-09-30 on 9202 (the R-741 proof): after the install hold opened, a stranger's default-login tries were refused (401) and after five of them mealie answered 423 (locked) to every login — the generated, correct password included. mealie's own brute-force guard, on an app published on the internet; the setup gate and the install hold do not cover an app after its first setup. Not measured: how long the lock lasts. **Needs:** measure the lock's length; decide whether the page tells the household what to do. `audits/night-rulings-2026-09-30/C/C3-mealie-poll.txt` **-- 2026-10-01:** Measured at v3.28.0 (source + 9202): 5 wrong logins lock the ACCOUNT (not the IP) for `SECURITY_USER_LOCKOUT_TIME` hours (default 24); the lock is lifted by an hourly job; an admin can unlock others via `POST /api/admin/users/unlock`, but the household's only admin is the locked account. Both login names are public (`admin`, `changeme@example.com`). **Fixed** (`09` §3 decision 57, decided by CC unattended — operator may reverse): `SECURITY_USER_LOCKOUT_TIME=1`, catalog `a4597cd`; on 9202 the right password answered 423 for 120 min, then 200; a wrong one still 401 (`audits/rulings-2026-10-01/C/`). **Left:** a stranger can renew the lock every hour (per account, public name — option (d) in decision 57 would end that); the page does not tell the household why it is locked; an INSTALLED mealie takes the new setting only when its compose is rendered again (not measured which act does that). | **NARROWED — the lock is 1–2 h; renewal and the page remain; owner: CC** **RULED 2026-10-06 10:41 (`09` §3 decision 144): A — one sentence on mealie's page about the 1–2 hour lock.** Being built. | — | — | CC | | **R-763** | Security & access | P3 | **[P2-MEDIUM] On wger a stranger can make an account after the household's setup, and every anonymous visit to the dashboard creates a guest account.** MEASURED 2026-10-01 on 9202 (live template `82fff32`), found by checklist row 3.4: after the admin existed, a stranger with no dashboard session `POST /en/user/registration` → 302, signed in with it → 302, read the API → 200; two anonymous `GET /en/dashboard` raised the user count from 2 to 4 (wger's middleware `create_temporary_user`, `utils/middleware.py:69`). Settings read inside the app: `ALLOW_REGISTRATION True`, `ALLOW_GUEST_USERS True` (the image defaults; the template sets neither). `GET /en/user/demo-entries` as a stranger answered 500. wger is FIRST-ADMIN class 3 (a known default login, fixed by `after_install`), so it never got decision 47's sign-up lock — it was not in R-711's list. Every crawler visit adds a user row to the household's database. **Needs:** per decision 47, close it after the first admin: `ALLOW_REGISTRATION=False` and `ALLOW_GUEST_USERS=False` (env switches the settings read — measure that the admin can still add family members, row 3.7), proven on 9202 as a stranger. `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). | **READY — rank P2-MEDIUM; owner: CC (catalog)** **Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again.** **2026-10-06: PUSHED** to the live catalog (`c265b37`; wger stays hidden). | — | — | CC | | **R-775** | Security & access | P3 | **[P2-MEDIUM] Grimmory: a stranger's 5 wrong sign-ins lock EVERY visitor out of the web login for 15 minutes — so Grimmory was not published.** MEASURED 2026-10-01 on 9202 (drill catalog, v3.4.1, through traefik): after 5 wrong tries for `admin` every further sign-in answered 429 — the household's right password AND a different name — and stayed 429 for 10+ minutes of retries. Read in the jar: `AuthRateLimitService` — Caffeine `expireAfterWrite(ofMinutes(15))`, `MAX_ATTEMPTS 5`, keys `login:ip:` and `login:user:`; Spring `forward-headers-strategy: native` takes the address from X-Forwarded-For, and behind the tunnel every visitor is the tunnel container's address (R-753) — the wger shape (R-752), with no setting to change it. Everything else in the checklist passed (bench + box step v3.4.1 → v3.5.0, gate by its own probe, OPDS through traefik); two smaller findings for the publishing session: on a reinstall over the first install's kept books, a new upload was saved to the drive but not added to the library (`box/grimmory/reinstall-c1.txt`, not investigated); and the remove + restore round trip (2.5) cannot be shown on 9202 for a drive app — its backup lives on the scratch drive, which is not a registered drive (R-756). The template waits in `audits/new-apps-2026-10-01/wip/grimmory/`. **Needs (operator):** (A) publish with a sentence on the page that wrong guesses by others can lock the login for 15 minutes (MEASURED: during the lock an e-reader's OPDS feed still answered 200 with its own login, wrong 401 — `box/grimmory/opds-under-lock.txt`), or (B) wait until the box passes each visitor's real address (R-753). `audits/new-apps-2026-10-01/box/grimmory/throttle.txt` **-- 2026-10-01 (evening):** option B's precondition SHIPPED (controller v0.286.1, R-753): Grimmory's Tomcat RemoteIpValve walks from the right and counts `172.16.0.0/12` as a proxy (READ in source, Spring Boot 4.1.1 — not yet measured with Grimmory's own lock), so `login:ip:` becomes per visitor; `login:user:` still lets a stranger lock the public name `admin` 15 min. A third route was spiked and passed: Grimmory behind the permanent family gate with its e-reader paths excepted (R-780). Recommendation: publish behind the family gate if R-780 is built; otherwise B with a measured 3.6. **UPDATE 2026-10-02 — NARROWED, Grimmory PUBLISHED behind the family gate:** a stranger cannot reach Grimmory's web sign-in at all (6 tries through the simulated tunnel: the gate's 401, then the household signs in 200 — `audits/family-gate-2026-10-02/A/items.txt`), and the e-reader exceptions keep Grimmory's own login. The 2.5 round trip now WORKS on 9202 (`audits/family-gate-2026-10-02/B/box/life.txt`). **What is left:** a family member past the gate can still lock a NAME (the admin's) for 15 minutes with 5 wrong tries — hard-coded in Grimmory; and the reinstall-over-kept-books finding (`new-apps-2026-10-01/box/grimmory/reinstall-c1.txt`) is still not investigated. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P2→P3: Grimmory is now behind the family gate; only a family member can still lock a name.** | — | — | CC | | **R-782** | Security & access | P3 | **[P3-LOW] Two side observations of the R-753 sweep, inferred, not measured:** glance's seeded `glance.yml` has no `auth:` block (the dashboard is public to anyone with the address), and homepage's `/api/*` refuses a Host not in `HOMEPAGE_ALLOWED_HOSTS`, which the template does not set (widgets may 400). **Needs:** measure both on 9202; glance: decide whether a public link dashboard is intended (the setup gate does not cover it after setup). | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC | @@ -254,7 +250,6 @@ stopping line that lies. | **R-177** | Monitoring & notifications | P4 | **There is no operator-triggerable "run the fill check now" path.** `fill-watch` is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller | **READY (S) — NEW 2026-08-02** **2026-10-05 (burn-down night): NEEDS A DESIGN** — the same missing operator door as R-314 and R-279. | — | **Noticed while live-validating R-167 on 9201, not by a failure.** It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. **Partially mitigated already** — v0.191.2 makes every run log a positive observable (`checked N filesystem(s), M unreadable/skipped, K notification(s)`), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has `GetJobs` but no run-now, so this is a general affordance, not a fill-watch one — **scope it as "run a named scheduler job now", operator-gated.** **ID established free:** `grep -ro "R-177\b" documentation/ *.md` → 0 hits | CC | | **R-266** | Monitoring & notifications | P4 | **A failed root `statfs` still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one.** Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. `report/builder.go:93-95` copies `sysInfo.DiskTotalGB` / `DiskUsedGB` / `DiskPercent` into `r.Storage[0]` (`Mount: "/"`), and those are exactly the zeros a failed `statfs` leaves behind — the controller now KNOWS the measurement failed (`SystemInfo.DiskKnown`, controller v0.210.0) and the report still does not carry it. **Deliberately not fixed here, for a reason that is now structural rather than a preference:** adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (`scripts/wire_contract_gate.py` refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. **RANKED LOW, and the reason is that the consequence is bounded:** the hub bands host storage on `disk_percent`, so a failed read presents as 0% used — the *quiet* direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. **Fix shape when it is taken:** carry `disk_known` on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 | **READY** — owner Viktor | — | — | operator | | **R-285** | Monitoring & notifications | P4 | **A planned, supervised reinstall pages the operator as if the machine had died — there is no notion of expected downtime anywhere.** During the 2026-08-09 rehearsal the hub sent, all `status: sent` to the operator channel: `host_stale` 08:58 UTC, `node_stale` 09:00, **`host_down` 09:28 (error)**, **`node_down` 09:30 (error)**, `host_leaf_changed` 09:31, `host_recovered` 09:31, `node_recovered` 09:34, `offsite_delivery_stuck` 09:34 — eight operator mails for work that was deliberate, attended and announced. **This is the OPPOSITE gap from the one R-281 filed:** the alarms are not missing, they are indiscriminate. `host_stale` at 30 min and `host_down` at 60 min (`monitor/host_staleness.go:22-23`, `downAfter = 2 * threshold`) cannot distinguish a wiped-on-purpose box from a dead one, and `host_leaf_changed` firing on a reinstall is correct-but-expected. **Note the interaction with the mute used on 2026-08-09 evening:** blocking a customer silences everything, so today the only two settings are *page me for planned work* and *tell me nothing at all*. **What is owed is a middle:** a maintenance window, or an operator-set expected-downtime flag, that suppresses staleness and leaf-change while leaving genuine faults audible | **READY (M) — NEW 2026-08-09** | — | The evidence is the operator's mailbox plus `events`/`notification_log` for 2026-08-09 | CC | -| **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — RULED 2026-10-05 ~21:00 (`09` §3 decision 129): after a crash boot, the controller's app mails wait the same grace period as after a normal start. Being built (burn-down night).** **2026-10-05 (burn-down night): the ruling as worded ALREADY HOLDS — NEEDS THE OPERATOR again.** The 90 s boot grace runs after every controller start, crash boots included (`cmd/controller/main.go:286`, `:928`, `:2174`); the 2026-10-04 mails came 3.5 and 9 minutes after the boot, after the grace, so decision 129 would not have stopped them. Next: either a LONGER grace after a crash boot (how long — the incident needed > 9 min), or suppress only the household leg, or close the row. A longer grace needs the crash-boot fact over the local API. **RULED 2026-10-06 10:41 (`09` §3 decision 143): A — app mails wait about 15 minutes after a crash boot.** Being built. | — | Build decision 129 | CC | | **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** **NIGHT WATCH 2026-10-06 05:00 (burn-down night): the run WORKED AS DESIGNED but could not be seen to.** Hub log: `Deadline check: 4 customers, 0 backup missed … 1 skipped (down)` and NO per-box line. Read from a hub.db copy with -wal (a different channel): Tester 2's first host report was 2026-10-04 16:13:44Z, its last 2026-10-04 18:05:48Z — at 05:00 it was 34.8 h old, under the 48 h line, so `judgeDownCustomer` returned early („nothing expected yet") — silently. No alarm was owed, none fired. **Fixed without a row (hub, on `main`, ships with the next hub release):** each early return now logs why (`… is DOWN — not judged yet: first host report … ago`); `TestR872_DownButTooYoungIsLogged`, red-proved. **RE-DATED 2026-10-06 → 2026-10-07:** from 16:14Z today Tester 2 is old enough; if it is still off at 05:00 tomorrow the existing `… is DOWN — judged on the longer lines …` line and its events are the proof. `audits/night-burndown-2026-10-05/r872-watch.txt`. | R-871 | — | CC | | **R-886** | Monitoring & notifications | P3 | **DooPlex's Alertmanager cannot write its own state since the Longhorn restart of 2026-10-05 13:20Z** — every 15 min `Running maintenance failed … open /alertmanager/nflog.…: permission denied` (and the same for `silences`), 8 times by 14:21Z; the pod was recreated 13:20:48Z by that restart. Mail still goes out (`alertmanager_notifications_total{integration="email"}` 5 → 6, `failed_total` 0, 14:23Z), but a silence set now and the record of what was already sent do not survive the next pod restart — so a restart can re-send every active alarm or drop a silence. Likely collateral of the restart (volume ownership on re-attach), not measured. **Checked from source 2026-10-05 (burn-down round 2):** homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root `nobody` image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts silences now survive a restart -- an invariant with no test, which is exactly what this row says broke. Last structural change 58d1cd2 (2026-08-14, 'give alertmanager re | **OPEN** | — | Compare the volume's file owner with the pod's `securityContext` (`fsGroup`/`runAsUser`); fix in homelab-manifests; prove with a silence that survives a pod restart | operator | | **R-884** | Monitoring & notifications | P4 | **ArgoCD app `monitoring` shows `Deployment/prometheus` OutOfSync** (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. **Checked from source 2026-10-05 (burn-down round 2):** Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml `prom/prometheus:v3.14.0` -> `v3.15.0` (monitoring.yaml:419), and the `monitoring` Application has no `automated` syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an unsynced Renovate bump, i.e. a full sync = Prometheus 3.14 -> 3.15 upgrade (plus pod restart; R-211: no reloader). Not confirmed live. | **OPEN** | — | `argocd app diff monitoring` (or the CR's resource diff) to see what differs, then decide git or live | operator | @@ -303,10 +298,8 @@ stopping line that lies. | **R-392** | Process & tooling | P4 | **No architecture document covers the two-AI workflow.** `documentation/architecture/` holds eight documents and **all eight cover the product** — topology, host agent, control-plane authorization, hub, off-site connectivity, backup, controller modules, capability map. Nothing records how the Claude.ai / Claude Code split works, what each side owns, how skills and `.claude/rules/` are scoped, or why. **The absence was found by trying to fill the template field, not by a survey:** the task that added the five process skills (2026-08-25) had to name an owning architecture document and could not, and the template requires that be recorded rather than passed over. The exposure today is low — the split is stable and both sides work — but it lives entirely in the operator's head and in chat, which is precisely the shape of a commitment nothing enforces. | **OPEN — LOW** | — | Write one architecture document for the agent-tooling layer: which AI owns which artifact class (`TASK-*.md`, `RUNBOOK-*.md`, validation), how skills are scoped and installed, what belongs in a `CLAUDE.md` versus a skill versus a rules file, and the reasoning for each boundary. **The rules themselves already exist** in `skills/felhom-doc-authoring/SKILL.md`; what is missing is the map of who owns what. Do not restate the doc-authoring rules there — point at that skill. | CC | | **R-394** | Process & tooling | P4 | **`felhom-build-deploy/SKILL.md` is 179 lines, over the 150-line limit its own repo now enforces.** Found 2026-08-25 by `scripts/check_skills.py` on its first run — the over-length was discovered BY the new checker, on the day the limit was written down, which is the checker working as intended. **It is not edited and not trimmed here:** the task that introduced the limit explicitly scoped the four pre-existing skills out, and trimming a build-and-deploy skill without exercising its commands is how a wrong command ships to a live host. **It is a named single-entry exception in `GRANDFATHERED` in `scripts/check_skills.py`, printed as a WARN on every run**, so it cannot fade; a NEW skill over the limit is convicted normally, and growing the set requires editing that file in a commit with a row to name. **The rationale for the limit** — attention thins across the excess, so the lines that matter are not the ones that survive — is in `skills/felhom-doc-authoring/SKILL.md` §5. | **OPEN — LOW** | — | Trim `felhom-build-deploy/SKILL.md` under 150 lines in a session that can VERIFY the commands it keeps, then delete its `GRANDFATHERED` entry in the same commit. The likely trim is the per-artifact command blocks moving behind a pointer to the runbooks, keeping the gotchas inline — but that is a judgement for a session with a build to run, not a line-count exercise. | CC | | **R-421** | Process & tooling | P4 | **THE CLASS: an instrument that matches a LABEL rather than the fact it names — five instances, every one found by accident.** R-410 (a `mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was ABSENT), R-94 (a test comparing a constant to itself). **The gates are the machinery that enforces everything else in this project, and they were the one part nothing had ever checked.** The 2026-09-01 decoy sweep read all 29 scripts and fooled **16**. Ten were fixed the same day; four remain with rows (R-422..R-425); six could not be given a plausible decoy and are named. **The shapes, so the next one is cheap to recognise:** (1) name-for-fact — it matches a path or directory NAME while the fact lives inside the file; (2) substring-for-field — it matches a token anywhere in a body instead of in the field that carries it; (3) declaration-for-reachability — it checks a thing is declared, not that it RESOLVES; (4) constant-for-measurement — it compares a value against itself. **The single largest cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level), so every one was green and correct today and would have gone blind the moment anyone added a subdirectory. `decoy_coverage_gate.py` now refuses a new gate that ships without a decoy. | **OPEN — the class row; it stays open as the place the next instance is recorded** | — | — | CC | -| **R-502** | Process & tooling | P4 | **[P3-LOW] The bootstrap regression harness is run by NO gate and NO CI — and it had never exercised the pairing banner.** MEASURED 2026-09-14: `grep -rn bootstrap-modes scripts/*.py .gitea/workflows` → nothing; the harness is run by hand in `felhom-iso-assistant:trixie`. While adding the R-496 checks, its fake hub's `register` reply turned out to carry no `pairing_code`, so `print_pairing_banner` returned early in every run: the banner that a household reads first had no test at all until ISO v1.27.0. **Fix shape:** register the harness in `repo_gates.py` behind a container-available check that reports INCONCLUSIVE (never skip-as-pass) when docker is absent, with a decoy. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: test harness coverage; no household meets it directly.** **2026-10-05 (burn-down night): NEEDS THE OPERATOR.** The harness runs `docker run`; registered as a gate it would run Docker on DooPlex on every full gate run. Next: an operator word that a non-fast container gate may run there (INCONCLUSIVE without docker; never in --fast — CI has no docker). **RULED 2026-10-06 10:41 (`09` §3 decision 147): A — the ISO first-boot test runs on DooPlex in full runs only; „not checked" without Docker.** Being built. | — | — | CC | | **R-507** | Process & tooling | P4 | **[P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen.** MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): `qm sendkey 332 tab` did not move focus (both password copies landed in one field), `mouse_move 1237 772` + `mouse_button 1` did not move the cursor or press Next, while `alt-n` did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. **Fix shape:** measure QEMU `input-send-event` with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: release-test tooling; an operator click-through covers it.** | — | — | CC | | **R-579** | Process & tooling | P4 | **[P3-LOW] Five page shells loaded `style.css` with NO cache-buster, so a browser holding an older copy kept being served CSS that did not know about the newest UI.** FOUND 2026-09-18 by the operator's screenshot of the v0.254.0 language globe: it rendered as a bare, unstyled `
` — a stray triangle and two plain words outside the card — because `login.html`, `claim.html`, `recovery.html`, `launcher_shared.html` and `launcher_share_password.html` requested `/static/style.css` with no `?v=`, while `layout.html` has used `?v={{.Version}}` since v0.166.0. **`.Version` was also absent from three of those five data maps.** FIXED AND CLOSED in the same session (controller v0.255.0): the parameter on all five, and `Version` set once in `executeTemplateLang` so a new shell cannot miss it; `TestGlobeOnAnonymousShells` now refuses an absent or EMPTY `?v=`. **The general form, which is the part worth keeping: a template that loads a versioned asset WITHOUT its version is invisible to every test that reads markup — the markup is correct and the browser fetches the wrong file.** A gate over "every stylesheet/script link in a template carries `?v=`" would catch the class; not built, because 8 first-boot-wizard templates would fail it and R-554 deletes them. | **DEFERRED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the gate "every stylesheet/script link carries `?v=`" was deliberately not built until R-554 deletes the old wizard; nothing else owns it) — **CLOSED 2026-09-18 - controller v0.255.0** | — | — | CC | -| **R-624** | Process & tooling | P4 | **[P3-LOW] Three of the catalog's apps cannot be seeded by ANY headless route, and for two of them that is a deliberate security decision — so the upgrade harness has a permanent ceiling nobody has written down.** FOUND 2026-09-21 while widening R-462 from 3 apps to . **`vaultwarden` and `zipline` close self-registration ON PURPOSE** — vaultwarden by `SIGNUPS_ALLOWED=false` (R-512, *„a stranger who guesses vault. must not be able to register"*), zipline by answering `E1037: User registration is disabled` — and neither ships a CLI that could make an account instead. So there is no route to a first account without the admin secret, and **that is correct**: the harness must not be the reason a customer-facing app accepts strangers. **`gitea` is a different and fixable case:** the template sets no `INSTALL_LOCK`, so a fresh instance sits in its web-installer state and `gitea admin user create` refuses (`MustInstalled() [F] Unable to load config file for a installed Gitea instance`); POSTing the installer form first would work and was simply not written tonight. **Why this is a row rather than three notes:** `09` §3 decision 6 says the upgrade test goes to **all** apps, and R-462 is costed as if every app is reachable given enough fixture work. It is not. There is a class — *apps whose only account-creating route the catalog deliberately closes* — for which the honest maximum is `inconclusive` unless the harness is given the app's admin secret at deploy time, which is a decision nobody has taken. **Needs:** the class named in R-462's scope so the remaining count is honest; a decision on whether the harness may hold an app's admin secret (it already holds the ones IT generates — see R-619); and, separately and cheaply, a gitea installer-form fixture. Evidence: `audits/update-night-2026-09-21/apps/{vaultwarden,zipline,gitea}/verdict.json`. **— NIGHT 2026-09-23:** `gitea` stayed inconclusive on both venues (its installer); `vaultwarden`/`zipline` not attempted (closed sign-up, by design); `code-server`, `outline`, `rallly` have no front-door seed route (listed, not moved); `bentopdf`, `glance`, `crafty-controller`, `wger`, `wanderer` (meilisearch) and `uptime-kuma` have no fixture tonight (listed, not moved). **-- 2026-09-30: the ceiling was WRONG for three of the six named apps.** outline has a front-door first-run route (`POST /api/installation.create` — its self-hosted setup, refused once a team exists), rallly signs up through its own better-auth API with the e-mail code read from its own table in place of a mailbox, and zipline 4's first-run route is `POST /api/setup` (the fixture had tried `/api/auth/register` and `/api/auth/setup`, which are not it). Fixtures for outline and rallly are in `upgrade_fixtures_box.py` and both apps moved on both venues; zipline's fixture now tries `/api/setup` first (measured on the bench: a SUPERADMIN made, the login works). **What remains in the class:** vaultwarden (closed sign-up by design, no CLI), code-server (a browser IDE). `audits/pg-last-six-2026-09-30/`. **-- 2026-09-30 (evening): gitea's fixable case is DONE** — the fixture posts gitea's own first-run installer form (the form's own defaults + an admin; no CSRF on that form), measured on the bench and through the setup gate on 9202; gitea 1.27.0 → 1.27.3 published `d7ba60c`. What remains in the class: vaultwarden, code-server. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (catalog harness)** **Re-ranked 2026-10-03: P3→P4: an upgrade-test coverage limit; no household meets it.** **RULED 2026-10-06 10:41 (`09` §3 decision 146): A — the bench may hold an admin password to seed vaultwarden and zipline, bench only.** Being built. | — | — | CC | | **R-652** | Process & tooling | P4 | **[P3-LOW] The memory watch counted the kernel's file cache as the app's memory.** MEASURED 2026-09-23 night on the bench: nextcloud 34.0.4 read **100 %** of its 1 GiB and immich's PostgreSQL **100 %**, each with **0** kernel `oom_kill`s — `memory.peak` includes page cache, which the kernel drops before it kills anything. Under the watch as built (R-635 follow-up) both would be marked `memory_tight`, and the gate would demand a raised `mem_limit` — a customer-box capacity figure — for cache. **Done the same night (09 §3 decision 22, CC — operator may reverse):** the watch samples the app's own memory (`anon` of `memory.stat`) every 15 s; the mark and the ladder's `memory_peak_pct` read it where measured; the cgroup peak stays beside it (`memory_cgroup_peak_pct`). **Open:** romm's backfilled entry carries M1's 80.9 % cgroup peak (measured before the anon sample existed) — re-measure it on its next step; and decide whether an app whose anon is low but whose cgroup stays pinned at its limit (cache thrash) should be marked at all. Evidence: `audits/night-2026-09-23/apps/nextcloud/bench-1024M/`, `apps/immich/bench-noanon/`. | **READY — P3; owner: CC (catalog harness)** **Re-ranked 2026-10-03: P3→P4: the main fix is done; what is left is harness tuning, no household meets it.** | — | — | CC | | **R-693** | Process & tooling | P4 | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** **Re-ranked 2026-10-03: P3→P4: a harness judgement problem; no household meets it.** | — | — | CC | | **R-739** | Process & tooling | P4 | **[P3-LOW] The test bench cannot run wanderer at all: its web server calls the database at the PUBLIC name `https://.`, which the bench has no name or TLS for.** MEASURED 2026-09-30 (more-night-apps brief, Part B): on bench 9401 the catalog template (v0.20.0) came up `wanderer-db` healthy, `wanderer-search` healthy, `wanderer` **unhealthy** for 12 min, every page 500 („Error 0: Something went wrong"); `PUBLIC_POCKETBASE_URL=https://hike-db.gate.invalid`. So the harness can only ever answer `inconclusive` at FROM for wanderer, never about an update. Its step `v0.20.0 → v0.21.0` (web + db) and meilisearch `v1.36 → v1.54` stay untested; **the brief's question — does the search index survive meilisearch's move or is it rebuilt — is NOT measured.** **Needs:** a bench venue that gives the stack the DB name (an `extra_hosts` + plain-http override in the bench's render only, stated in the verdict), or wanderer proven on the box venue alone with the operator's word. `audits/more-night-apps-2026-09-30/B/wanderer-probe.txt` **-- 2026-09-30 late: the bench CAN run wanderer now** — `upgrade-test.py` `BENCH_ENV_OVERRIDES` points `PUBLIC_POCKETBASE_URL` at the database container on the bench only (the only address the web image reads, measured in v0.20.0 and v0.21.0), recorded in every verdict: web 200, sign-up (`PUT /api/v1/user`) 200, login 200. **The meilisearch question, answered:** v1.54.2 REFUSES v1.36's database („Your database version (1.36.0) is incompatible”); with `MEILI_UPGRADE_DB=true` it upgrades in place and all three indexes (actors, lists, trails) are there — so wanderer's step needs that switch in the template first. Not done: a fixture (creating a list answered PocketBase's „Failed to create record” — likely a rule for unverified users), whether a household's record survives in the index, the step itself. `audits/night-rulings-2026-09-30/` | **NARROWED — the bench runs it; the step needs a fixture and `MEILI_UPGRADE_DB`; owner: CC** **Re-ranked 2026-10-03: P3→P4: test bench coverage; the template switch and fixture are tooling work.** | — | — | CC | @@ -315,6 +308,7 @@ stopping line that lies. | **R-805** | Process & tooling | P4 | **[P3-LOW] The volume-persistence gate judges an empty NAMED volume (R-788) but not an empty BIND mount.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): Grimmory's `/app/data` bind held 0 files and read CLEAN (its state is in MariaDB); komga, paperless-ngx, radarr and sonarr also carry empty binds that are not judged. **Needs:** decide whether an empty bind after the exercise is UNDETERMINED like a named volume, with a decoy. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a test-gate rule; no household meets it.** | — | — | CC | | **R-806** | Process & tooling | P4 | **[P3-LOW] The gate's GET exercise speaks plain http to an HTTPS backend (crafty-controller :8443), and gramps-web did not answer on :5000.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): both reached data only through their fixture seed. **Needs:** read the `loadbalancer.server.scheme` label (upgrade_boxport already does); find why gramps-web's first GET got no answer. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a test-gate defect; no household meets it.** **NARROWED 2026-10-05 (app-catalog `4828dc7`): the harness GET now uses the traefik scheme label (https + -k for crafty-controller). LEFT: gramps-web :5000 not answering — a live look on a box.** | — | — | CC | | **R-807** | Process & tooling | P4 | **[P3-LOW] 13 apps stay UNDETERMINED because their seeds write only to the database, so their upload / media / cache / redis volumes stay empty.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): claper, crafty-controller, dawarich (redis), docmost, gramps-web, immich (ML cache), outline (+redis), sparkyfitness, tandoor, vikunja, wger, wishlist, zipline — plus plex (no seed route: a plex.tv claim token) and wanderer (no fixture; unhealthy under the gate). None is BROKEN: nothing was written outside a preserved folder. **Needs:** per app, a seed that uploads one file (or the volume named n/a with a reason: a redis cache, an ML model cache). | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: test coverage; nothing was found broken.** | — | — | CC | +| **R-891** | Process & tooling | P4 | **felhom.eu `CLAUDE.md` says `--fast` selects „all of them" — no longer true since the full-run-only `iso-bootstrap` gate (R-502, 2026-10-06).** An instruction file; the session may not edit it. | **OPEN — filed 2026-10-06** | — | Operator: change the line in felhom.eu `CLAUDE.md` „Gates — ONE entry point" to say `--fast` skips the full-run-only gates (today: `iso-bootstrap`) | operator |