From 30650cad6e384af35e343bc6b87126f8cfed94e8 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 5 Oct 2026 20:41:50 +0200 Subject: [PATCH] burn-down round 2: controller v0.297.0 rows closed (23), golden 0.297.0 evidence, delivery evidence, 23-row unchecked table, STATUS/CONTEXT/REPORT (292 -> 199; 1 opened, 94 closed) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- CONTEXT.md | 10 + REPORT-burndown2-2026-10-05.md | 115 ++++++ STATUS.md | 29 +- .../controller-red-proofs.txt | 6 + .../delivery/controller-delivery.txt | 4 + .../delivery/hub-deploy.txt | 11 + .../delivery/vouch-golden-floors.txt | 11 + .../unchecked-results.jsonl | 23 ++ documentation/backlog/CLOSED-ITEMS.md | 23 ++ documentation/backlog/OPEN-ITEMS.md | 49 +-- .../02-round-trip.txt | 5 + .../tests/golden-0.297.0-2026-10-05/README.md | 52 +++ .../tests/golden-0.297.0-2026-10-05/bake.log | 339 ++++++++++++++++++ 13 files changed, 640 insertions(+), 37 deletions(-) create mode 100644 REPORT-burndown2-2026-10-05.md create mode 100644 documentation/audits/burndown2-2026-10-05/delivery/controller-delivery.txt create mode 100644 documentation/audits/burndown2-2026-10-05/delivery/hub-deploy.txt create mode 100644 documentation/audits/burndown2-2026-10-05/delivery/vouch-golden-floors.txt create mode 100644 documentation/audits/burndown2-2026-10-05/unchecked-results.jsonl create mode 100644 documentation/tests/golden-0.297.0-2026-10-05/02-round-trip.txt create mode 100644 documentation/tests/golden-0.297.0-2026-10-05/README.md create mode 100644 documentation/tests/golden-0.297.0-2026-10-05/bake.log diff --git a/CONTEXT.md b/CONTEXT.md index 079a8c6e..662f5a5b 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -16,6 +16,16 @@ > and holds nothing of its own; this file does hold its own content, namely the standing rulings below. +> **2026-10-05 (late night) — burn-down round 2 (releases).** Register 292 → 199 (1 opened: R-888; 94 closed: 43 +> accepted by the operator 18:23, 51 fixed). Releases: hub v0.137.0 (`557629d`, deployed), agent v0.147.0 (tag, sha +> `642c4d19…`, bundle `326527d0…`, signed jobs to 3 boxes), controller v0.297.0 (`1453cfc` + `6f1ba1f`), golden 0.297.0 +> (`8cebc42e…`, vouched with agent 0.147.0 / min_agent 0.131.0; floors 0.297.0 for demo-hp, demo-felhom, tester-1), +> catalog `4828dc7`. R-124: recipe root namespace = `""` (+ runbook). R-887 mechanism from Gitea's log: a FetchTask the +> runner abandons after assignment → zombie stop after ~10 min; load = an outside crawler + the session's own 15-page CI +> polling (now one `runs?head_sha=` call per minute). New gates: `stands` (felhom.eu), `gofmt` (controller; NOT +> CHECKED out loud on the Go-less runner). R-469 not done: the permission check refused the catalog CLAUDE.md edit. +> Report: `REPORT-burndown2-2026-10-05.md`. + > **2026-10-05 (night) — the burn-down (no release; DooPlex/ep0 untouched).** Register 336 → 292 (1 opened — R-887 CI runner fault — 45 > closed): 24 fixed by later work + 2 duplicates (each re-checked; `audits/burndown-2026-10-05/partA-table.md` holds all > 317 P3/P4 verdicts), 19 small fixes with tests/red-proofs (catalog `29ac711`, agent `d833163`, controller `114ff27`, diff --git a/REPORT-burndown2-2026-10-05.md b/REPORT-burndown2-2026-10-05.md new file mode 100644 index 00000000..0b65daa8 --- /dev/null +++ b/REPORT-burndown2-2026-10-05.md @@ -0,0 +1,115 @@ +# REPORT — burn-down round 2: the operator's answer recorded, R-887 re-diagnosed, small rows fixed WITH releases — 2026-10-05 (late night) + +| Part | Result | +|---|---| +| **A** — rulings, then R-887 | **done** — rulings commit `301fe45` (count after: **249**); R-887 re-diagnosed from the logs (the restart idea refuted; the mechanism then SEEN in Gitea's own log), dated check 2026-10-12 | +| **B** — R-124, the small rows, the 23 unchecked | **done** — R-124 fixed (agent v0.147.0 + runbook); 50 more rows fixed and closed with tests and red-proofs; the 23 checked from source (1 duplicate closed, facts added to 8 rows, the rest left as they need a live box or a decision) | +| **B.4** — releases, delivered the normal way | **done** — hub v0.137.0 deployed; agent v0.147.0 released + signed jobs (binary and bundle) to demo-hp, demo-felhom, Tester 1; controller v0.297.0 + golden 0.297.0 baked, vouched, floors raised, all three boxes on 0.297.0; catalog pushed | +| **C** — numbers and record | **done** — STATUS shows 199 and asks nothing about the closed list | + +| Rows before | Rows after | Opened | Closed | +|---|---|---|---| +| **292** | **199** | **1** (R-888) | **94** (43 accepted by the operator + 51 fixed/merged) | + +Counted by `register_shape_gate.py`'s method. Target ≤ 220: met. + +## Baselines (re-verified at the start) + +felhom.eu `e8c56c440a` (hub v0.136.0) · controller `114ff2761a` (v0.296.0) · agent `d83316326e` (v0.146.1) · catalog +`29ac711d26` · golden 0.296.0 · register 292. The agent clone had a stray `scripts/__pycache__/` from round 1 — removed. + +## Part A — the rulings commit and R-887 + +- `301fe45`: 43 rows closed as „accepted by the operator, 2026-10-05", each with its one-line reason from the list; + R-124 and R-698 kept (R-698 owner → operator); R-831/R-870 carry the not-rotated rulings; R-887 records the screenshot + (one runner, ID 2, online). STATUS: the list and the rotate/runners requests removed. **Count after: 249.** CI run + 1363 success. +- **R-887, from the logs:** the runner's last restart was 13:24:42Z; the lost attempts started 15:05–15:46Z — **not a + restart**. Four lost attempts (not two): each without a runner `task` line, each failed at a :38-second mark 10–13 min + after assignment. Gitea's log for that hour had rotated. **Then it happened again at 17:15Z with the log intact:** + `slow POST …/RunnerService/FetchTask for 10.42.0.42, elapsed 3192ms` → `context canceled` → 17:28:39 + `clear_tasks.go … stopTasks() … task 1371` — the runner abandoned its fetch after Gitea assigned the task; Gitea's + zombie stop failed it. Load at that minute: an outside crawler on public commit pages, and this session's CI waiter + (15-page job listings at 13–31 s each). The waiter now makes ONE `runs?head_sha=` call a minute. A lost run re-runs + with `POST …/actions/runs//rerun` (used twice: controller run 1357 → success; catalog run 1368 → success). + **Dated check 2026-10-12** in DUE-CHECKS. The fix on DooPlex (runner fetch timeout, crawler) is the operator's. + +## Part B — fixes by repo + +**agent v0.147.0** (`f1b9b41`, CI 1365; tag `v0.147.0`; binary sha256 `642c4d19…`, bundle `326527d0…`, verified by +download; CHANGELOG `208fac8`, CI 1367): R-124, R-118, R-269, R-317 — red-proofs `audits/burndown2-2026-10-05/r124-red-proof.txt`, +`agent-red-proofs.txt`. **Delivery:** vouched (agent 0.147.0, golden 0.296.0 first), signed `agent_update` ×3, then +`agent_config_update` ×3 (felhom-op-1, ttl 45 m); hub System page: demo-hp, demo-felhom, Tester 1 — agent 0.147.0, +root files 0.147.0 (`delivery/`). Tester 2 offline — nothing sent. + +**hub v0.137.0** (`557629d`, CI 1369; manifest `81d04a6`; CI 1370): R-277, R-581, R-600, R-544, R-855, R-134, R-92, +R-292, R-599, R-725, R-728, R-208 (hub half) — red-proofs `felhom-eu-red-proofs.txt`. **Deployed:** ArgoCD Synced/Healthy +at `d75ad0f`, image `felhom-hub:0.137.0`, `felhom-hub 0.137.0 starting`, healthz 200; R-855's new line seen live +(„after 2 healthy ring-0 night(s)"). (The build ran while a helper was still appending to an audit text file outside +`hub/` — the image is the committed `hub/` tree; said here because the clean-tree gate is literal.) + +**felhom.eu gates/tools/docs** (same commits): R-819 (`stands` gate), R-857, R-555, R-364 (`hu_grep.py` + REUSE.md), +R-587, R-571, R-129 (demo-hp authenticates with DooPlex's own key — corrected everywhere it said „no key"), R-124 runbook. +New script tests pass under a BusyBox + bash + python3 + git PATH (the CI runner's tools): 19/19. + +**controller v0.297.0** (`1453cfc`; CI run 1371 **FAILED** — the new gofmt gate was INCONCLUSIVE on the Go-less runner; +fixed in `6f1ba1f`, CI 1372 success): R-591, R-568, R-567, R-363, R-547, R-10, R-552, R-251, R-104, R-619, R-362, R-675, +R-256, R-257, R-240, R-365, R-425, R-565, R-564, R-603, R-454, R-208, R-457 (swept, nothing left) + two twins found and +fixed on the way (the top-bar countdown at 0 days; nine more shared references in `deepCopyStack`). Red-proofs +`controller-red-proofs.txt` (two first attempts that did not convict are marked, with valid re-runs). **MinAgent 0.131.0.** +**Image** `felhom-controller:0.297.0`. **Golden 0.297.0** baked per RUNBOOK §4.0–4.1 (`documentation/tests/golden-0.297.0-2026-10-05/`: +all pass markers, round trip sha `8cebc42e…`, token leak 0 with a working control, teardown to `virgin`). **Vouched** +(agent 0.147.0, golden 0.297.0, min_agent 0.131.0) and **floors** 0.297.0 for demo-hp, demo-felhom, tester-1. +**Delivered:** demo-hp and demo-felhom `felhom-controller:0.297.0 … (healthy)`; Tester 1 reports Controller 0.297.0 +(„Controller frissítve: 0.296.0 → 0.297.0"). + +**catalog** (`4828dc7`; CI run 1368 lost by R-887, re-run success): R-593, R-760, R-594, R-605, R-781, R-806 (scheme half; +row narrowed), plus a stale runner test (expected 11 gates, 12 exist) and a test that never ran (outside its class) — +fixed, not filed. + +**Not done, and why:** R-469 and R-605's exit-code line in the catalog's `CLAUDE.md` — **the permission check refused +the instruction-file edit**; the operator is asked (rule 5). R-126 needs an operator choice. R-325 needs a same-step +felhom.eu gate change (left). R-377 (CONTEXT headings) not attempted. Installer rows (R-179, R-180, R-275, R-276, R-306, +R-130, R-310, R-881), R-136 (logs every operator out), R-502 (Docker in CI), R-798 (a live app definition) and the +larger controller rows (R-492, R-569, R-575, R-615, R-616, R-498, R-718) were left on purpose. + +**Opened:** R-888 — two report fields the hub never reads (a decision). **Seen, not a row:** Tester 1's crash guard reads +TRIPPED since 07:57Z — the morning's two deliberate test crashes; it re-arms by itself after 24 h (`runbooks/crash-guard.md`). + +## The 23 rows the first burn-down could not check + +Checked from source by a read-only agent (`audits/burndown2-2026-10-05/unchecked-results.jsonl`). Closed: R-350 (duplicate +of R-132, facts merged). Facts added to the open rows R-607, R-883, R-886, R-884, R-756, R-91, R-338, R-488. The three +„not worth it" ones are on STATUS for the operator. The rest need a live box reading (the settle command is in the table). + +| Row | Group | Evidence / how to settle (abridged) | +|---|---|---| +| R-76 | UNCHECKABLE-FROM-SOURCE | Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 FileBrowserImage = "gtstef/filebrowser:1.5.6-stable". The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --gre | +| R-91 | UNCHECKABLE-FROM-SOURCE | Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of delet | +| R-209a | UNCHECKABLE-FROM-SOURCE | Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test. — settle: uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h / | +| R-337 | NOT-WORTH-IT | The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 b, err : | +| R-375 | NOT-WORTH-IT | The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is pvesm status showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 if st == nil // st.Type == "pbs" // st.Avail <= 0 { return true, "" } (space preflight sk | +| R-488 | STILL-TRUE-SMALL | The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded interval := 5 * time.Second and time.Sleep(3 * time.Second) // initial settling time, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go: | +| R-504 | UNCHECKABLE-FROM-SOURCE | Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html. — settle: curl -sI https://iso.felhom.eu/ / head -1 | +| R-644 | UNCHECKABLE-FROM-SOURCE | Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs --deployment-password; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches | +| R-814 | UNCHECKABLE-FROM-SOURCE | Hetzner account state; nothing in source records a deletion. — settle: Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status | +| R-815 | UNCHECKABLE-FROM-SOURCE | PBS server-side state on ep0; no GC completion record in the docs (grep). — settle: ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 / grep -i garbage' | +| R-884 | UNCHECKABLE-FROM-SOURCE | Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml prom/prometheus:v3.14.0 -> v3.15.0 (monitoring.yaml:419), and the monitoring Application has no automated syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an un | +| R-132 | UNCHECKABLE-FROM-SOURCE | Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key operator_password_hash, setSetting :2200-2207 writes updated_at = datetime('now'). No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31 | +| R-298 | UNCHECKABLE-FROM-SOURCE | Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 if(d.role==='user-data'){ else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/ | +| R-338 | UNCHECKABLE-FROM-SOURCE | nodes.md:86-88 still claims demo-hp is on the R-50 island (local_api on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent con | +| R-350 | DUPLICATE | of R-132 — Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence. | +| R-542 | NOT-WORTH-IT | Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 initialize = append(initialize, c) // every unclaimed disk can be initialized; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkf | +| R-607 | STILL-TRUE-SMALL | Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY if len(newApps) > 0 // len(updated) > 0, and :257-258 says 'nincs változás' when both are empty. updated counts stack-dir copies only (copyTemplates, :447 updated = append(updated, appName) a | +| R-683 | UNCHECKABLE-FROM-SOURCE | Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power- | +| R-756 | UNCHECKABLE-FROM-SOURCE | Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH | +| R-862 | UNCHECKABLE-FROM-SOURCE | Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83). — settle: Ask the operator wheth | +| R-882 | UNCHECKABLE-FROM-SOURCE | Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it). — settle: sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p Act | +| R-883 | STILL-TRUE-SMALL | homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/rec | +| R-886 | STILL-TRUE-SMALL | homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root nobody image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts | + +## Teardown + +Machines: drill VM — build guest destroyed, token/scripts/log shredded, powered off, disk back on `virgin`. Boxes: only the +normal deliveries above. Hub: only the deploy, the vouch and the floors. Scratch secrets (hub password file, hub key file, +signed envelopes) are shredded at the end of the session. diff --git a/STATUS.md b/STATUS.md index 449e0c0c..a43e89ab 100644 --- a/STATUS.md +++ b/STATUS.md @@ -3,7 +3,34 @@ **Ready for the first real tester (Tester-2): yes. Tester 2 (a laptop, off at night) was offline again; nothing was sent to it.** -**Updated 2026-10-05 (late night, burn-down round 2 — in progress): 249 open rows after your answer (was 292).** +**Updated 2026-10-05 (late night, burn-down round 2): every box of ours healthy. The open-items list is at 199 (was 292 +at the start of this round, 336 this morning). Report: `REPORT-burndown2-2026-10-05.md`.** + +## Tonight, last (2026-10-05): the list at 199 + +**What happened:** +- **Your answer is recorded** (below): 43 rows closed as accepted. +- **51 more rows fixed and closed**, with one release per repository, delivered the normal way: hub **0.137.0** + (live), agent **0.147.0** (on demo-hp, demo-felhom and Tester 1, root files too), controller **0.297.0** (on all + three), new-install image **0.297.0** (baked, checked, approved), app catalog updated. +- **R-124 is fixed** (your ruling): the recovery recipe now writes the backup-server namespace the way the server reads it. +- **The CI fault (R-887) is understood:** when Gitea is busy, the CI runner's request for work can time out after Gitea + already gave it the job; the job is then never run and is failed 10–13 minutes later. Gitea was busy because of an + outside web crawler and because of MY CI checks, which asked too much — mine now ask once a minute in one small request. +- **Two of my own CI misses tonight:** a new check needed Go, which the CI machine does not have (fixed); a red run was + the fault above (re-run passed). + +**The numbers:** 292 before → **199 after**; 1 opened; 94 closed. + +**Needs you (none urgent; if you do nothing, each stays open as it is):** +1. **R-469** — a one-paragraph rewording in the app catalog's instruction file; my permission check refused editing + instruction files. Say "go" and it is done in a minute. +2. **R-126** — should a network share be offered as an export destination? Two options in the row; pick one. +3. **R-888** — two facts the boxes report that the hub never shows (installed-app list, retired drives). Needed or not? +4. **R-887** — CI: raise the runner's fetch timeout and/or slow the crawler on Gitea's public pages. If nothing: now and + then a CI run fails without running; it can be re-run. +5. Three more rows look „not worth doing" (R-337, R-375, R-542 — the check found each is by design or harmless). Close + them as accepted? If you say nothing they stay. ## Your answer to the burn-down list (2026-10-05 18:23), recorded diff --git a/documentation/audits/burndown2-2026-10-05/controller-red-proofs.txt b/documentation/audits/burndown2-2026-10-05/controller-red-proofs.txt index 213d3ad5..b104de08 100644 --- a/documentation/audits/burndown2-2026-10-05/controller-red-proofs.txt +++ b/documentation/audits/burndown2-2026-10-05/controller-red-proofs.txt @@ -497,3 +497,9 @@ rc=0 --- PASS: TestR365_BannerSaysDueAtZeroDays (0.12s) PASS ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.132s + +### CI fix (lead, 2026-10-05 ~18:20Z): the new gofmt gate was INCONCLUSIVE on the CI runner (no Go) — CI run 1371 FAILED +Fix: in CI (GITEA_ACTIONS/GITHUB_ACTIONS=true) with no gofmt reachable the gate prints "NOT CHECKED in CI" and exits 0; +elsewhere a missing gofmt stays INCONCLUSIVE. Decoys gofmt/ci-without-go and gofmt/dev-without-go (empty PATH). +Simulated CI (env -i, empty PATH, GITEA_ACTIONS=true): "gofmt gate NOT CHECKED in CI …" rc=0. +Red-proof (CI branch disabled): FAIL: gofmt/ci-without-go: want rc=0 and 'NOT CHECKED in CI', got rc=2. diff --git a/documentation/audits/burndown2-2026-10-05/delivery/controller-delivery.txt b/documentation/audits/burndown2-2026-10-05/delivery/controller-delivery.txt new file mode 100644 index 00000000..afba1cec --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/delivery/controller-delivery.txt @@ -0,0 +1,4 @@ +== controller delivery 2026-10-05T18:28:59Z +demo-hp gitea.dooplex.hu/admin/felhom-controller:0.297.0 Up 32 seconds (healthy) +demo-felhom gitea.dooplex.hu/admin/felhom-controller:0.297.0 Up 35 seconds (healthy) +tester-1 (hub customer page, version strings seen): 6 0.297.0 4 0.296.0 4 0.295.0 diff --git a/documentation/audits/burndown2-2026-10-05/delivery/hub-deploy.txt b/documentation/audits/burndown2-2026-10-05/delivery/hub-deploy.txt new file mode 100644 index 00000000..86407d7d --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/delivery/hub-deploy.txt @@ -0,0 +1,11 @@ +== hub deploy 2026-10-05T18:11:56Z +image=gitea.dooplex.hu/admin/felhom-hub:0.137.0 +sync=Synced health=Healthy op=Succeeded rev=d75ad0fdf3daf3f3b2a1690746d9a6a70ee4104c +2026/10/05 20:11:05 [INFO] felhom-hub 0.137.0 starting +healthz 200 +2026/10/05 20:11:06 [INFO] osupdates: the Docker engine set is approved only by the operator, after 2 healthy ring-0 night(s) +== System page 2026-10-05T18:12:05Z: root-files / agent cells +73: 0 armed unknown 0.142.0 → 0.147.0 (since 2026-10-05) +89: 0 armed 0.147.0 0.147.0 +105: 2 armed 0.147.0 0.147.0 +121: 2 TRIPPED 2026-10-05T07:57:17Z 0.147.0 0.147.0 diff --git a/documentation/audits/burndown2-2026-10-05/delivery/vouch-golden-floors.txt b/documentation/audits/burndown2-2026-10-05/delivery/vouch-golden-floors.txt new file mode 100644 index 00000000..14a36489 --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/delivery/vouch-golden-floors.txt @@ -0,0 +1,11 @@ +== vouch 2026-10-05T18:27:46Z: agent 0.147.0, golden 0.297.0, min_agent 0.131.0 +HTTP/1.1 303 See Other +Location: /configuration?flash=artifacts_set +== floors 2026-10-05T18:28:16Z: POST /customers//floor min_controller_version=0.297.0 min_agent=0.131.0 +demo-hp: Location: /customers/demo-hp?flash=floor_set +demo-felhom: Location: /customers/demo-felhom?flash=floor_set +tester-1: Location: /customers/tester-1?flash=floor_set +2026/10/05 20:28:16 [INFO] Artifact manifest set: agent=0.147.0 golden=0.297.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d" +2026/10/05 20:28:16 [INFO] Customer demo-hp controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0") +2026/10/05 20:28:17 [INFO] Customer demo-felhom controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0") +2026/10/05 20:28:17 [INFO] Customer tester-1 controller-version floor override set to "0.297.0" (declared MinAgent "0.131.0") diff --git a/documentation/audits/burndown2-2026-10-05/unchecked-results.jsonl b/documentation/audits/burndown2-2026-10-05/unchecked-results.jsonl new file mode 100644 index 00000000..5d2b3adb --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/unchecked-results.jsonl @@ -0,0 +1,23 @@ +{"id": "R-76", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Image changed since the 1.3.3 finding: felhom-controller@114ff27 controller/internal/infra/infra.go:27 `FileBrowserImage = \"gtstef/filebrowser:1.5.6-stable\"`. The comment infra.go:207-208 still asserts folders come out '2775 with the parent's setgid' -- the exact claim R-76 measured false on 1.3.3; no test pins it (git log --grep 'R-76|setgid' in controller: only 2026-06 commits). Whether 1.5.6 still drops setgid is runtime behaviour of the image.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- find /mnt/felhom-drives -path '*/userdata/*' -mindepth 3 -maxdepth 5 -type d ! -perm -2000 -printf '%m %u:%g %p\\n'\" | head (any folder a customer made in FileBrowser showing 755 without setgid = still true on 1.5.6)", "minutes_spent": 3} +{"id": "R-91", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word.", "dup_of": null, "unique_facts": "Stale doc to fix in the same commit as the delete: documentation/runbooks/offsite-endpoint.md:24 still says datastore `felhom-offsite` is at `/srv/pbs-felhom`; the real path since 2026-07-27 is `/mnt/pbs-datastore` (RUNBOOK-ep0-datastore-volume-2026-07-27.md:8). CONTEXT.md:3666 is the other line to change.", "small_fix": null, "not_worth": null, "settle_cmd": "ssh root@ep0 'du -sh /srv/pbs-felhom 2>&1; proxmox-backup-manager datastore list; df -h /'", "minutes_spent": 4} +{"id": "R-209a", "sev": "P4", "category": "Process & tooling", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Pure live state on DooPlex (whether a reboot has happened and the post-boot check passed). No source claim to test.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "uptime -s; cat /var/log/felhom-store-postboot-check.log; ls -d /var/lib/containerd.pre-move-2026-08-05; df -h /", "minutes_spent": 1} +{"id": "R-337", "sev": "P4", "category": "Monitoring & notifications", "group": "NOT-WORTH-IT", "evidence": "The row's first question ('establish the intended refresh path') is answered by source: GET /backup/status reads only the agent's in-memory store (felhom-agent@d833163 internal/localapi/server.go:1258 -> pickLatestBackup :1304-1318), and the ONLY writer is the job goroutine after the whole runner returns: server.go:885 `b, err := tier.Service.BackupWithSnapshotHook(...)` then :901 `s.store.RecordBackup(b)` (grep RecordBackup: no other caller). So it is NOT a collection cadence; the status appears when the agent's own job finishes (WaitTask + archive resolve, internal/backup/runner.go:205-232), and a backup the agent did not run itself never appears. The 4-min demo-hp skew is the gap between PBS writing the manifest and the job returning -- unmeasured.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: measure why demo-hp's job returned ~4 min after the manifest landed; cost: a constructed live repro on demo-hp with task-log timing; if never: during an incident the box's backup status can trail the PBS manifest by minutes after an out-of-schedule run, and a run started outside the agent never shows; pick: close with the source fact above written into the row (status = agent job end, not polling), reopen only if a lag is seen on a scheduled run.", "settle_cmd": null, "minutes_spent": 7} +{"id": "R-375", "sev": "P4", "category": "Backup & restore", "group": "NOT-WORTH-IT", "evidence": "The signal (audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:168-171) is `pvesm status` showing felhom-pbs Total/Used/Avail = 0. Nothing in the product consumes those numbers for a PBS target: felhom-agent@d833163 internal/backup/runner.go:265 `if st == nil || st.Type == \"pbs\" || st.Avail <= 0 { return true, \"\" }` (space preflight skips PBS and fails open on 0), and the same report says the hub's gauge reads the real 3.7/97.9 GB.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: confirm on ep0 that the namespace-scoped token lacks Datastore.Audit; cost: a read-only ep0 session; if never: the PVE UI on a box shows 0/0/0 for felhom-pbs, which no Felhom code reads (runner.go:265); pick: close as cosmetic with this pointer.", "settle_cmd": "(if ever wanted) ssh root@ep0 'proxmox-backup-manager acl list' ; on a box: pvesm status --storage felhom-pbs", "minutes_spent": 6} +{"id": "R-488", "sev": "P4", "category": "Process & tooling", "group": "STILL-TRUE-SMALL", "evidence": "The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded `interval := 5 * time.Second` and `time.Sleep(3 * time.Second) // initial settling time`, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go:90, restore_unit.go:471. No commit since 2026-09-13 touching internal/backup mentions R-488/test speed. Runtime (333 s) was NOT re-measured here (read-only).", "dup_of": null, "unique_facts": null, "small_fix": "controller: add Manager fields healthSettle/healthInterval (defaults 3s/5s, set in the constructor) used by waitForHealthy; a test helper (the existing Manager fixture constructor) sets them to 0/10ms. Test: TestWaitForHealthy_DefaultsAreProduction asserts a fresh Manager has 3s/5s (pins production), and measure `go test ./internal/backup` wall time before/after in the commit message (expect the 89 >=1 s tests to drop). Run with the R-650 docker-free seams.", "not_worth": null, "settle_cmd": null, "minutes_spent": 5} +{"id": "R-504", "sev": "P4", "category": "Install & onboarding", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Live HTTP behaviour of iso.felhom.eu; curl to hosts is outside this checker. Source side: documentation/runbooks/VOLUNTEER-first-hour.md:14 still says the root has no index (R-504); the download page exists at website/letoltes.html.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "If 404 is confirmed it is cosmetic (row's own re-rank); the operator could close it as accepted rather than add a Cloudflare rule.", "settle_cmd": "curl -sI https://iso.felhom.eu/ | head -1", "minutes_spent": 2} +{"id": "R-644", "sev": "P4", "category": "Apps & catalog", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Live scratch-box state. Source context: app-catalog-felhom.eu templates/gokapi/docker-compose.yml:26-29 seeds config.json with an EMPTY Password only when config.json is absent, then runs `--deployment-password`; a config.json that exists with an empty/plain password (e.g. a restored volume or an interrupted first boot) matches the crash text. Not proven.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- docker ps -a --filter name=gokapi --format '{{.Names}} {{.Status}}'; pct exec 9202 -- grep -s -c '\\\"Password\\\":\\\"\\\"' /opt/docker/stacks/gokapi/config/config.json\"", "minutes_spent": 3} +{"id": "R-814", "sev": "P4", "category": "Hub & operator", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Hetzner account state; nothing in source records a deletion.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Hetzner Storage Box API (read-only): GET https://api.hetzner.com/v1/storage_boxes/611421 with the operator's API token (stored out-of-band) -> 404 = deleted, else read .storage_box.status", "minutes_spent": 1} +{"id": "R-815", "sev": "P4", "category": "Backup & restore", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "PBS server-side state on ep0; no GC completion record in the docs (grep).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh root@ep0 'proxmox-backup-manager garbage-collection status felhom-offsite; proxmox-backup-manager task list --all --limit 20 | grep -i garbage'", "minutes_spent": 2} +{"id": "R-884", "sev": "P4", "category": "Monitoring & notifications", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml `prom/prometheus:v3.14.0` -> `v3.15.0` (monitoring.yaml:419), and the `monitoring` Application has no `automated` syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an unsynced Renovate bump, i.e. a full sync = Prometheus 3.14 -> 3.15 upgrade (plus pod restart; R-211: no reloader). Not confirmed live.", "dup_of": null, "unique_facts": "Cause candidate: Renovate 53c6e99 prometheus v3.15.0 merged 2026-10-03, never synced because monitoring has no auto-sync in git.", "small_fix": null, "not_worth": null, "settle_cmd": "sudo kubectl -n mon-system get deploy prometheus -o jsonpath='{.spec.template.spec.containers[0].image}' (v3.14.0 => the diff is the Renovate bump)", "minutes_spent": 5} +{"id": "R-132", "sev": "P3", "category": "Security & access", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Whether HUB_PW was rotated is not in source. The hub stores a UI-set password in hub_settings with an updated_at column: felhom.eu hub/internal/store/store.go:2182 key `operator_password_hash`, setSetting :2200-2207 writes `updated_at = datetime('now')`. No commit records a rotation (git log --grep rotate/HUB_PW since 2026-07-31).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "On a copy of the hub DB (hub.db + -wal + -shm): sqlite3 hub.db \"SELECT updated_at FROM hub_settings WHERE key='operator_password_hash'\" -- rotated only if later than 2026-09-18 (the last exposure, R-580 folded here); no row = still the ConfigMap seed", "minutes_spent": 4} +{"id": "R-298", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Template gate unchanged: felhom-controller@114ff27 controller/internal/web/templates/storage.html:364 `if(d.role==='user-data'){` else :368 protected, no actions. Dependency R-280 is CLOSED (CLOSED-ITEMS.md:600, v0.211.0). Whether the bug bites depends on the role the agent gives the drive: felhom-agent@d833163 internal/storage/role.go:172-186 -- a dir storage (felhom-backup) is user-data unless its backing device is on the system disk; disks.go:1236-1240 deliberately does not reclassify the backup-target drive. role.go unchanged since 2026-08-09. demo-hp was reprovisioned before 2026-09-21, so the 2026-08-10 topology may no longer hold.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp 'grep -A2 felhom-backup /etc/pve/storage.cfg; lsblk -no PKNAME $(findmnt -no SOURCE /) ; lsblk -no PKNAME $(findmnt -no SOURCE /mnt/nvme-1tb)' (same parent disk => role=system => the page still locks it => R-298 true; different => user-data => not reproducible on this box)", "minutes_spent": 8} +{"id": "R-338", "sev": "P3", "category": "Security & access", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"grep -E 'listen_addr|island_' /etc/felhom-agent/agent.json; pct config 9201 | grep ^net; ip -br link show master vmbr9\"", "minutes_spent": 6} +{"id": "R-350", "sev": "P3", "category": "Security & access", "group": "DUPLICATE", "evidence": "Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence.", "dup_of": "R-132", "unique_facts": "(1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked.", "small_fix": null, "not_worth": null, "settle_cmd": null, "minutes_spent": 3} +{"id": "R-542", "sev": "P3", "category": "Storage & devices", "group": "NOT-WORTH-IT", "evidence": "Still true in source, and by design: felhom-agent@d833163 internal/localapi/disks.go:438 `initialize = append(initialize, c) // every unclaimed disk can be initialized`; a disk mounted under /mnt/felhom-drives counts as UNCLAIMED on purpose (internal/storage/claim.go:88-103, R-220, so drives survive a guest rebuild), and the mkfs wrapper explicitly allows it: configs/felhom-mkfs-guarded.sh:59-60 'Mounts under /mnt/felhom-drives are our own drives (the agent detaches before a re-init) -> allowed'. Controller passes initialize through untouched (agent_disk_handlers.go:158-162). `already_mounted` is only set for controller-contributed stores (agentapi/client.go:447-451, omitempty) -- so 'null' is expected for agent candidates.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": "what: drop felhom-mounted disks from `initialize`; cost: reverses the R-220/re-init design (wrapper comment) and needs a decision on how a registered drive is re-initialized; if never: the raw endpoint lists a registered drive under initialize while the page (customer view) filters it -- only a session reading the raw endpoint is misled; pick: close as by-design, add one comment line at disks.go:438 saying a registered felhom drive is listed here deliberately.", "settle_cmd": null, "minutes_spent": 9} +{"id": "R-607", "sev": "P3", "category": "App updates", "group": "STILL-TRUE-SMALL", "evidence": "Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY `if len(newApps) > 0 || len(updated) > 0`, and :257-258 says 'nincs változás' when both are empty. `updated` counts stack-dir copies only (copyTemplates, :447 `updated = append(updated, appName)` after a hash mismatch); for a deployed+pinned app whose catalog moved, renderSource (:457-469 table) copies the STORED definition, so the hash matches and nothing is 'updated' although the git cache moved. CatalogImages is read from the catalog cache only inside ScanStacks (stacks/manager.go:666-672, assigned :690). So (a) the message measures the stack dir, not the catalog; (b) CatalogImages refreshes only on a ScanStacks, which this sync does not trigger. The 29 s nextcloud case (needed several rounds) is not explained by this.", "dup_of": null, "unique_facts": null, "small_fix": "controller sync.go: record the catalog git HEAD before and after gitCloneOrPull; if it moved, call s.rescanFn() even when newApps/updated are empty, and say 'Katalógus frissítve — az alkalmazások nem változtak' (and EN) instead of 'nincs változás'. Test (red first): a Syncer with a fake pull that moves HEAD and a frozen pinned app (renderSource returns the stored definition) asserts rescanFn was called once and the message is not the no-change one; and a no-move pull asserts rescanFn NOT called.", "not_worth": null, "settle_cmd": null, "minutes_spent": 9} +{"id": "R-683", "sev": "P3", "category": "App updates", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Evidence in repo: audits/night-2026-09-24/E/round-03-controller.log is the POST-cut log only (first lines 12:05:09Z: update.go:1335 'interrupted in verifying (started 2026-09-24T12:04:21Z)', :589 undo copies '.pre-update-20260924T120426Z'); the pre-cut log that would show a backing-up phase is lost (row says so). No later power-cut-mid-update drill: night-2026-10-04/MORNING-NOTE.md:64 'A1 power cut mid-update (demo-hp): NOT RUN'; DRILL-night-2026-09-25.md:102 was a cut during romm's verifying, not checked for the backup choice.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Not a read-only command: the next power-cut-during-update drill with the controller log saved at arm time, then grep \"phase backing-up\\|Tier-2\" in the saved pre-cut log", "minutes_spent": 7} +{"id": "R-756", "sev": "P3", "category": "Storage & devices", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 `if !m.DriveLive(hddPath)` -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 `return m.isMountPoint(hddPath)` -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH is compared to a registered storage path (api/router.go:574-575, web/handlers.go:3049, storage_handlers.go:528), i.e. a drive root. The message names `.../scratch_hdd/userdata/calibre-web`, so on 9202 HDD_PATH is a per-app subfolder, which can never be a mount point -> the 409 is certain for that value. Open: whether that HDD_PATH was written by the product (handlers.go:550 prefill from place.Drive) or by the test venue.", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "ssh demo-hp \"pct exec 9202 -- grep -H HDD_PATH /opt/docker/stacks/calibre-web/app.yaml /opt/docker/stacks/grimmory/app.yaml; pct exec 9202 -- findmnt -no TARGET,SOURCE /mnt/felhom-drives/scratch_hdd\" (HDD_PATH = per-app subfolder => product/venue wrote a non-root HDD_PATH; HDD_PATH = drive root and not mounted => scratch drive not registered)", "minutes_spent": 10} +{"id": "R-862", "sev": "P3", "category": "Box system & updates", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Waits on the operator's by-hand bootstrap on Tester 2; no commit records it (felhom.eu log since 2026-10-04). runbooks/config-bundle.md:77 'CC sends the bundle by the signed job and reads it back on the System page'; a box behind the vouched bundle for 7 days raises os_config_bundle_behind (:83).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "Ask the operator whether the three bootstrap commands ran; read-only proof: the hub's operator view of Tester 2 (OS/config-bundle sha vs the vouched sha), or whether os_config_bundle_behind has fired for it", "minutes_spent": 2} +{"id": "R-882", "sev": "P3", "category": "Hub & operator", "group": "UNCHECKABLE-FROM-SOURCE", "evidence": "Longhorn instance-manager runtime state on DooPlex; nothing in homelab-manifests addresses it (no commit since 87dfc29 names it).", "dup_of": null, "unique_facts": null, "small_fix": null, "not_worth": null, "settle_cmd": "sudo kubectl -n longhorn-system get pods -l longhorn.io/component=instance-manager -o custom-columns=NAME:.metadata.name,START:.status.startTime ; systemctl show k3s containerd iscsid -p ActiveEnterTimestamp (instance-manager older than a k3s/containerd/iscsid restart => the stale-PID state can recur)", "minutes_spent": 2} +{"id": "R-883", "sev": "P3", "category": "Hub & operator", "group": "STILL-TRUE-SMALL", "evidence": "homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/recipe-importer:latest (pull Always :27); adventurelog-system/adventurelog.yaml:100 and :256 adventurelog-backend/frontend:latest (pull Always :101,:257); jarrs-system/jarr-dev.yaml:311,:345,:630 gitea.dooplex.hu/admin/jarr:latest (pull Always). Zipline fixed (90f60e4, 4c8ec7a). The repo's own rule homelab-manifests/CLAUDE.md:120 'Image tags always pinned'. Helm values files (external-dns, pihole, plex, authentik, cnpg) not checked for tag fields beyond a `tag: latest|dev|empty` grep (no hits).", "dup_of": null, "unique_facts": null, "small_fix": "homelab-manifests only: for each line above, read the running digest/version (`sudo kubectl get pod -n -o jsonpath='{..imageID}'`), pin that exact tag (or @sha256 for the self-built gitea.dooplex.hu jarr/recipe-importer images, which have no version tags), add the version-checker match-regex annotation per CLAUDE.md:120, and set imagePullPolicy IfNotPresent. Test: `grep -rnE 'image:.*(:latest|:dev)\\s*$' --include=*.yaml .` returns nothing; after ArgoCD sync each pod's imageID equals the pre-change one (no upgrade). Operator-owned DooPlex change: needs the operator's word.", "not_worth": null, "settle_cmd": null, "minutes_spent": 6} +{"id": "R-886", "sev": "P3", "category": "Monitoring & notifications", "group": "STILL-TRUE-SMALL", "evidence": "homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root `nobody` image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts silences now survive a restart -- an invariant with no test, which is exactly what this row says broke. Last structural change 58d1cd2 (2026-08-14, 'give alertmanager real storage').", "dup_of": null, "unique_facts": null, "small_fix": "homelab-manifests: add pod `securityContext: {fsGroup: 65534, fsGroupChangePolicy: OnRootMismatch}` (65534 = nobody, the image's user -- confirm with `sudo kubectl -n mon-system exec deploy/alertmanager -- id`) to the alertmanager Deployment. Test (consequence, per CLAUDE.md): create a silence via amtool/API, delete the pod, after it returns the silence is still listed AND the log has no 'Running maintenance failed ... permission denied' within 15 min (positive observable: a 'maintenance done' line).", "not_worth": null, "settle_cmd": null, "minutes_spent": 5} diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index c32a0557..20de06c1 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -103,6 +103,29 @@ The full text of every row below: `git show e8c56c44:documentation/backlog/OPEN- | **R-269** | **A rotated-out per-guest local-API token still authorises, and the test that appears to pin the opposite passes only because of its lookup ORDER.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-agent v0.147.0 (delivered): token store reloads on growth before answering; `TestTokenStore_RotatedOutTokenRejectedFirst`; red-proof agent-red-proofs.txt | | **R-317** | **The agent decides whether to install dnsmasq by stat-ing a file the OTHER package owns.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-agent v0.147.0 (delivered): dnsmasq install probed by its service unit; `TestEnsureDnsmasq_*`; red-proof agent-red-proofs.txt | | **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** (P3) | CLOSED 2026-10-05 — DUPLICATE of R-132 (its unique facts moved there) | Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence. | +| **R-591** | **[P3-LOW] `Stack.Copy()` is a deep copy with one shallow field, and the field is new.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `deepCopyStack` copies every reference in Meta (i18n overlay and 11 more found by `TestDeepCopyStackMetaSharesNoReference`); `TestDeepCopyStackI18nIsNotShared` | +| **R-568** | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: disk-health rows sorted by durable id; `TestDiskHealthRows_OrderIsStable` | +| **R-567** | **[P3-LOW] The two drive wizard pages (/storage/init, /storage/attach) do not highlight the Tárhely menu group — the sidebar reads as if the household left the storage section.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: storage wizard pages open the Storage nav group; `TestStorageWizardPages_OpenTheStorageNavGroup`; two parity fixtures re-captured (nav only) | +| **R-363** | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: fill-watch also every 10 min (`sched.Every`), daily + start-up kept; `TestFillWatchRunsOnAnInterval` | +| **R-547** | **[P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: `disk_critical` is defined at ≥95 % used, but the fill-watch runs once a day.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: same change as R-363 (the interval watch); `TestFillWatchRunsOnAnInterval` | +| **R-10** | T-6E-1: DB-dump dir-fsync asymmetry (LOW, confirmed in 6E) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-15, size XS, roadmap state `idea`.** Moved verbatim; nothing added (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: dump directory fsynced after the rename; `TestDumpOneTo_SyncsTheDumpDirectoryAfterRename` | +| **R-552** | **[P3-LOW] An interrupted-restore notice for an app that is then REMOVED stays on the restore page for ever.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: remove clears the interrupted-restore notice (`ClearInterruptedRestore`, wired in the remove path); `TestR552_RemoveClearsTheInterruptedRestoreNotice` | +| **R-251** | **The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: off-site marker tag not listed as an app; `TestR251_MarkerTagIsNotAnApp` | +| **R-104** | **An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a lock surviving the self-heal is classed `locked` with its own cause line; `TestR104_SurvivingLockIsNamed` | +| **R-619** | **[P3-LOW] A `type: password` deploy field is MANDATORY however `required` reads, and the `deploy-fields` contract says the opposite — so any caller that trusts it is refused.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: password deploy fields served as required by the API (fresh metadata per call); `TestR619_PasswordFieldIsServedAsRequired` | +| **R-362** | **A data drive detached mid-restore is reported as „permission denied".** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a restore onto a detached drive names the drive; `TestR362_DetachedDriveIsNamed` | +| **R-675** | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: files-restore refusal names the second drive's whole copy (and is in both languages now); `TestR675_RefusalNamesTheWholeCopy` | +| **R-256** | **C2 — „A mentéskezelő nem elérhető." names no route at all.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `flash.offbox.mgr_unavailable/_unreachable` name a route (hu+en); `TestR256_R257_OffboxRefusalsNameARoute` | +| **R-257** | **C2 — „Az offsite tároló nincs elárvult állapotban." puts an English loanword and an internal state name in front of a Hungarian household customer, and names no route.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `flash.offbox.not_orphaned` reworded (hu+en); `TestR256_R257_OffboxRefusalsNameARoute` | +| **R-240** | **A backup that covered nothing calls itself „Sikeres".** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a zero-selection run no longer says „Sikeres"; `TestR240_ZeroSelectionRunDoesNotSaySuccess` | +| **R-365** | **An overdue abandonment countdown renders its past due-date in the future tense.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: a 0-day deletion countdown says it is due — page AND top bar (`TestR365_OverdueCountdownIsNotFutureTense`, `TestR365_BannerSaysDueAtZeroDays`) | +| **R-425** | **`offbox_rename_gate.py` scans a fixed three-entry `FILES` list.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: offbox rename gate finds files by pattern and judges the bundle; 3 decoys | +| **R-565** | **[P3-LOW] The English page test sees only ACCENTED Hungarian: an ASCII-only Hungarian word left in a template passes it on the English page.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: English-page test knows ASCII Hungarian words; a real `mp` leak became a key; `TestI18nEnglishPages` | +| **R-564** | **[P3-LOW] The retrieval-promise gate's Hungarian stems cannot see a SPLIT verb — „csak akkor állíthatók vissza", „hozod vissza" — so those Hungarian sentences were never scanned; the English translation exposed them.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: retrieval-promise gate knows split-verb Hungarian; 7 occurrences registered; 2 decoys | +| **R-603** | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: Go-side check for HTML-escapable values; `TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping` | +| **R-454** | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: `scripts/gofmt_gate.py` (NOT CHECKED out loud on the Go-less CI runner, INCONCLUSIVE elsewhere; 3 decoys); 12 files formatted | +| **R-208** | **Every Felhom Go build re-downloads its modules because `ARG VERSION` sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: controller and hub Dockerfiles: `ARG VERSION…` just above `go build`; `TestR208_DockerfileVersionArgsSitBelowModuleDownload`, felhom.eu `scripts/test_dockerfile_arg_order.py` (hub v0.137.0 deployed) | +| **R-457** | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** (P4) | CLOSED 2026-10-05 — CHECKED, NOTHING LEFT (burn-down round 2) | felhom-controller v0.297.0 (`1453cfc` + CI fix `6f1ba1f`, CI run 1372 success; image `felhom-controller:0.297.0`; golden 0.297.0 vouched; delivered to demo-hp, demo-felhom, tester-1 — `audits/burndown2-2026-10-05/delivery/controller-delivery.txt`); red-proofs `audits/burndown2-2026-10-05/controller-red-proofs.txt`: swept: none of the six candidate files has a date literal feeding an assertion against the real clock (identity/format/ordering checks only) — nothing to change; the faked-future-date CI idea is a separate, larger job | --- diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index d56a68f3..30008a6c 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -131,7 +131,7 @@ stopping line that lies. | **R-504** | Install & onboarding | P4 | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** **Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic.** | — | — | operator | | **R-881** | Install & onboarding | P4 | **The installer's uninstall does not remove `/usr/local/sbin/felhom-priv-apply`** (added to the bundle by agent v0.146.1, R-861), and its comment still says the guest hook is agent-installed at runtime (`scripts/felhom-host-install.sh` ~line 1679) — since v0.146.1 the hook is a bundle file. Found 2026-10-05 while checking the installer against the new bundle. | READY | — | Add the file to the uninstall list and fix the comment at the next installer tag | CC | -## Apps & catalog — 23 rows (P3 11, P4 12) +## Apps & catalog — 21 rows (P3 11, P4 10) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -147,9 +147,7 @@ stopping line that lies. | **R-762** | Apps & catalog | P3 | **[P2-MEDIUM] wger serves no CSS or JavaScript and no uploaded photo: every static file and every `/media/` file answers 404.** MEASURED 2026-10-01 on 9202 (drill catalog, the live template `82fff32`, wger 2.7), found by checklist rows 1.7 and 2.8: the login page links `/static/css/workout-manager.css`, `/static/bootstrap-compiled.css` — both 404 through traefik; the static root inside the container is empty (4 KB). A progress photo posted to `/api/v2/gallery/` answered 201 and the file is on the media volume, but `GET /media/gallery/…png` answers 404 signed in, without a session, and straight at the app inside the container. **Cause, read in the image:** the entrypoint runs `collectstatic` only when `DJANGO_DEBUG == "False"` and the template sets no `DJANGO_DEBUG`; and wger serves `/media/` only in development (`urls.py:393` „served like this during development only”) — upstream's production setup puts nginx in front for `/static` and `/media`. So the household gets an unstyled app and photos that never show. The same lines stand since the template was written (the 2026-09-29 template too). Not checked: whether any box runs wger (on 2026-09-30 none reported to the hub). **Needs:** `DJANGO_DEBUG=False` (collectstatic) and something that serves `/static` + `/media` (upstream's nginx sidecar, or the gunicorn switch of R-755 plus a static server), proven on the bench and on 9202 with a page that loads its CSS and a photo read back. Owner decides together with R-755 (same server question). `audits/new-app-checklist-2026-10-01/C/C8-signup-guest-media-static.txt`, `C/C4-seed-photo-size.txt` **-- 2026-10-01 (operator):** wger is `lifecycle: hidden` until this and its twin are fixed (catalog `55b8c8a`; read back on 9202: not on the app list, mealie control present). **Merged 2026-10-05 from R-755 (duplicate):** `templates/wger/docker-compose.yml` still sets no `WGER_USE_GUNICORN` — the gunicorn switch is the same server question. | **READY — rank P2-MEDIUM; owner: CC (catalog)** **Re-ranked 2026-10-03: P2→P3: wger is hidden from installs and no box runs it; needed only before it is offered again.** | — | — | CC | | **R-774** | Apps & catalog | P3 | **[P3-LOW] Two things the new apps' pages do not show yet: Karakeep's mail-ON path is unproven, and its official phone app reports crashes to its makers.** READ/MEASURED 2026-10-01: Karakeep's `smtp_mapping` (plaintext :2526, as Cal.com) was proven only with mail OFF (a fresh install boots) — 9202 has no hub, so the relay cannot be exercised there; the official mobile app ships Sentry crash reporting with a hard-coded DSN (`apps/mobile/app/_layout.tsx`, FIT.md). **Needs:** one password-reset mail from Karakeep on a hub-enabled box (demo-hp 9201, a throwaway install); and a sentence on the page about the phone app (copy, freeze). `audits/new-apps-2026-10-01/bench/karakeep-mail-off-boot.txt` | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC | | **R-76** | Apps & catalog | P4 | **FileBrowser-created folders break the setgid chain, and a drop-zone's mode is not stable** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-26, size S, roadmap state `idea (surfaced by the R-75 spike, 2026-07-26)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Two related findings from `audits/SPIKE-catalog-data-paths-2026-07-26.md` P3/P5, both **pre-existing** and deliberately left alone by that spike. **(a)** FileBrowser Quantum 1.3.3 creates files `0644` and folders `0755` and does **not** propagate the setgid bit — even though the entrypoint wrapper's `umask 002` really is in effect (`/proc/1/status` `Umask: 0002`). Group inheritance itself works (a file uploaded into a 2775 group-100 dir landed group 100, not the process gid 1000), so the convention's *group* half holds and only its *mode* half is lost. The consequence is proven with a control: inside a UI-created `0755` folder a gid-1000 process's file landed group **1000**, while the identical write into the 2775 parent landed group **100**. So **any folder a customer creates through FileBrowser breaks the shared-group chain one level down.** Latent today — every userdata-touching catalog app that declares an identity declares uid/gid **1000**, the same uid FileBrowser runs as, so owner permissions mask it; it bites the day a content app runs as a different non-root uid with gid 1000. The comment at `infra/infra.go:156` is right that the image ignores `-e UMASK` but does not say t | CC | -| **R-567** | Apps & catalog | P4 | **[P3-LOW] The two drive wizard pages (/storage/init, /storage/attach) do not highlight the Tárhely menu group — the sidebar reads as if the household left the storage section.** FOUND 2026-09-17 by slice 1 release C: the release C parity cases first used page name `storage` for the wizards; re-captured with the handler's real page name (`storage_handlers.go` `storageWizardPageHandler` passes the TEMPLATE name, `storage_init`/`storage_attach`, as `Page`) the fixtures lost `nav-group is-open` and the `active` link. Present since the wizards shipped; not caused by localisation. **Fix shape:** pass `Page` `storage` (or teach the layout's storage group both names), with a render assertion that /storage/init carries the open storage group. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic navigation polish.** | — | — | CC | | **R-577** | Apps & catalog | P4 | **[P3-LOW] A guest SHARE visitor has no way to pick a language, and the household's setting is the wrong default for them.** FOUND 2026-09-18 by localisation slice 2 release C (R-557, controller v0.254.0): every other page a person can reach now carries a language globe — the dashboard (the household's setting), and the sign-in and claim pages (the visitor's own cookie). The two guest share pages (`launcher_shared`, `launcher_share_password`) deliberately do NOT, and `TestGuestSharePagesHaveNoGlobe` pins that so it stays a decision rather than an oversight. **Why it is the operator's and not CC's:** a share visitor is a stranger the household sent a link to, and what language they are shown is a promise the SHARE FEATURE makes, not an implementation detail. The `felhom_lang` cookie already built would fit them exactly (display-only, their own browser, never the household's setting). **Fix shape, if the operator says yes:** add `{{template "lang_globe" .}}` to both shells with the anonymous form, and one render case per page per language. | **READY - rank P3-LOW; owner: operator (the decision), CC (the change)** **Re-ranked 2026-10-03: P3->P4: feature decision for the operator; Hungarian default works today.** | — | — | operator | -| **R-619** | Apps & catalog | P4 | **[P3-LOW] A `type: password` deploy field is MANDATORY however `required` reads, and the `deploy-fields` contract says the opposite — so any caller that trusts it is refused.** MEASURED 2026-09-21 on guest 9202 while widening the update drill. `GET /api/stacks/grafana/deploy-fields` serves `{"env_var":"GF_SECURITY_ADMIN_PASSWORD","type":"password","generate":"password:16","required":false}`; a deploy carrying only the two `required:true` fields is refused **400** „a(z) „Admin jelszó" mező kitöltése kötelező — használja a Generálás gombot…". **The BEHAVIOUR is right and is a decision, not a bug:** `deploy.go:305-312` refuses a `password` field with no caller value on purpose — *"We never silently auto-generate — the user needs to know their password"* — which is the opposite of the `secret` case one branch above, where a generated value the customer never sees is exactly correct. **The defect is the CONTRACT.** `.felhom.yml` declares `required: false`, the API serves that verbatim, and nothing on the wire distinguishes "optional because the box will generate it" (`secret`) from "optional in the template and mandatory in the code" (`password`). A person using the deploy page never meets this because the page renders a Generálás button; **anything that is not that page does**, which now includes this drill harness and would include `09` §6.2's unattended caller the day it deploys anything. **Fix shape (smallest that keeps the decision):** serve `required: true` for `type: password` in the deploy-fields response — one place, derived rather than stored, so templates need no edit — and a test asserting a `password` field always reaches the wire as required. Alternatively state it in the field's `description`, which is weaker because it is prose. Evidence: `audits/update-night-2026-09-21/apps/grafana/log.txt` (the refusal) and `batchA.log`. | **READY — rank P3-LOW; owner: CC (controller)** **Re-ranked 2026-10-03: P3→P4: only non-page callers (test harness) meet it; the deploy page works.** | — | — | CC | | **R-644** | Apps & catalog | P4 | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** **Re-ranked 2026-10-03: P3→P4: seen only on a scratch test guest; the row itself calls it not a customer fact.** | — | — | CC | | **R-707** | Apps & catalog | P4 | **[P2] 37 apps still start with a login a stranger can take (`09` §3 decision 45).** Audit of all 53 apps: `app-catalog-felhom.eu/FIRST-ADMIN.md` (class, fix route, status, measured or read). Open: **3 hard-coded defaults** — calibre-web (`admin / admin123`, measured working on demo-hp and 9202; its own `cps.py -s` route needs a generated password WITH a special character — our generator is letters+digits, a controller change), mealie (`changeme@example.com / MyPassword`), wger (`admin / adminadmin`); **34 open first-run screens** (the first visitor creates the admin: actualbudget, adventurelog, audiobookshelf, calcom, docmost, emby, ghost, gitea, gramps-web, home-assistant, homebox, immich, jellyfin, komga, n8n, navidrome, opengist, outline, papra, plant-it, radarr, rallly, recipe-importer, romm, seerr, sonarr, sparkyfitness, tandoor, termix, uptime-kuma, vikunja, wanderer, wishlist, zipline). **Stale notes:** romm's `default_creds` `admin / admin` answers 401 on demo-hp (like a wrong password) — the page now warns with a login that does not exist; zipline's looks stale too. **Measured on demo-hp 2026-09-28 (read-only):** bookstack's default still logs in on the INSTALLED app (the fix is for new installs; the page now warns). Each fix: route (a) env or (b) the app's own CLI/API via `after_install:`, proven on 9202 with the default failing and the generated password working; route (c) a page sentence. Several sessions (operator, 2026-09-28). **2026-09-29 (controller v0.280.0, catalog `d0e7e2e`):** every class-3 app fixed — mealie, wger, calibre-web by `after_install` (calibre-web with the new `password:24:special`), proven on 9202 fresh installs (`audits/login-gate-2026-09-29/D/`); the setup gate (decision 46, spike PASSED) built and live on immich, n8n, audiobookshelf (probes measured) and uptime-kuma (button) (`…/C/`); romm's and zipline's stale notes removed. **Left: 30 class-4 apps** — gate each (probe measured on 9202 where one exists — 11 upstream candidates listed in `…/B/B-VERDICT.md` §3; the button otherwise). **2026-09-29 afternoon (controller v0.281.0, catalog `6faf432`):** 28 more class-4 apps gated — 32 of 34 — each proven on 9202 (`audits/gate-rollout-2026-09-29/`B): stranger → gate page / 401, household reached the first-setup screen, the gate opened (9 by a measured probe, the rest by the press), the app answered after. seerr, outline, rallly: gated, their opening needs a media server / e-mail (not proven). **Left:** wanderer (R-714); plant-it is not installable. | **NARROWED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — seerr, outline and rallly are gated, but the gate OPENING is not proven (needs a media server / e-mail)) — **CLOSED — 2026-09-29 (the rest → R-714)** | — | — | CC | | **R-718** | Apps & catalog | P4 | **[P3-LOW] "Close sign-up now" restarts an app with its own switch, and the card does not say so.** MEASURED 2026-09-29 on demo-hp: pressing it on adventurelog recreated its backend (~30 s, one 500 on its login page). The window's card says the app restarts; the close card does not. **Fix direction:** the close card and the gate-open moment say "the app restarts once" where `after_setup.env` exists. **ALSO MEASURED 2026-09-29 (new-household drill):** the gate-open press on a fresh vikunja restarted it for its own switch — the front door answered 404 for ~2 s and nothing said so. | **OPEN — P3; owner: CC** **Re-ranked 2026-10-03: P3→P4: a short unannounced restart; copy polish only.** | — | — | CC | @@ -166,7 +164,7 @@ stopping line that lies. | **R-440** | App updates | P3 | **[P2-MEDIUM] 23 catalog image pins float, so an update is not reproducible.** `compose pull` on a moving tag fetches whatever upstream published that day. **MEASURED 2026-09-01 over `app-catalog-felhom.eu` @ `29edad9c5bf4`: 79 `image:` lines across 53 apps, 66 distinct; 23 of those lines carry a tag with no patch version.** `postgres:16-alpine` (8 apps), `redis:7-alpine` (6), `mariadb:11.6` (2), plus one each of `postgres:15-alpine`, `postgis/postgis:16-3.5-alpine`, `mariadb:11.4`, `mariadb:12.3`, `ghcr.io/claperco/claper:2.5`, `ghcr.io/thomiceli/opengist:1.13`, `wger/server:2.6`. **A 24th is arguable and is recorded rather than rounded away:** `ghcr.io/immich-app/postgres:16-vectorchord0.4.3-pgvectors0.2.0` pins both extensions exactly but leaves the PostgreSQL patch floating. A customer pressing Frissites can therefore swap their DATABASE ENGINE build with no catalog change and no record; two boxes updated on two days end up different. **Severity MEDIUM on its own; it becomes BLOCKING the moment a pre-update copy exists**, because "what did we upgrade from and to" must be recordable and today it is not — which is also why R-440 must be read next to the digest discipline in Rule 10 of the spike. **MEASURED LIVE 2026-09-01 — the floating pins have ALREADY moved, with a passing control.** Running digests on demo-hp compared against what the registry serves for the same tag today: **`mariadb:11.4` MOVED** (`sha256:4f1d8d20...` -> `sha256:611a2fcc...`) and **`mariadb:12.3` MOVED** (`sha256:a02fe89c...` -> `sha256:dd9b303a...`), while `postgres:16-alpine`, `redis:7-alpine`, `mariadb:11.6` and `opengist:1.13` were SAME — **and both fully-pinned CONTROLS (`rommapp/romm:5.0.0`, `privatebin/pdo:2.0.5`) were SAME.** So on a box with ZERO visible drift by tag, pressing Frissites today silently swaps the DATABASE ENGINE build under `romm` and `bookstack`, with no catalog change and no record. **Compounding fact found while reading:** the recovery unit records `ImagePins` but the manifest comment says *"image NOT stored - re-pulled on restore"*, so a RESTORE of a floating-pinned app also re-pulls whatever is current — the same non-reproducibility on the recovery path. **HALF OF THE ANSWER SHIPPED 2026-09-02 (controller v0.233.0, slice 1): `app.yaml.installed_images` now records, per compose SERVICE, the reference AND the repo digest each container was actually created from — so "what did we upgrade FROM" is answerable on any box that has taken one lifecycle action since the upgrade.** What is still missing is the other half: comparing that digest against what the registry serves for the same tag TODAY, which needs a network call the render path deliberately does not make (see R-446). **The row therefore stays OPEN and its rank is unchanged** — recording a digest does not make a floating pin reproducible; it makes the drift measurable after the fact. `audits/SPIKE-app-update-2026-09-01.md` **— NIGHT 2026-09-23:** every image moved tonight (and the 21 backfilled moves) carries its resolved digest in the catalog's test record; the floating pins themselves still float until the box pulls by digest (`09` §6.4 part 6). **-- MEASURED 2026-09-30 on demo-hp 9201: a fresh install or a guarded Update renders `name:tag@sha256` from the ladder (`digest.go` `RenderWithLadderDigests`) — 9 of 20 services there run such a definition (adventurelog, docmost, paperless-ngx). The other 11 run TAG-ONLY definitions: bookstack, kimai, opengist, privatebin, romm were installed before v0.269.0 and not updated since, and `CarryDigests` (`digest.go:141`) keeps only a digest the running definition already has; bentopdf and calibre-web have no ladder at all. So a re-pull of those (a restore) takes whatever the tag serves that day. `audits/pg-last-six-2026-09-30/E/E2-digests-demo-hp.txt`.** **-- 2026-09-30 (evening):** calibre-web now has a ladder entry (`53a4a1d`), so a new install or a guarded Update renders its tested digest; bentopdf still has none. `audits/more-night-apps-2026-09-30/` **Merged 2026-10-05 from R-446 (duplicate):** the household-visible consequence — for these apps the „Naprakész" badge can never turn „behind" (`felhom-controller controller/internal/stacks/updateorder.go:96` returns false with no catalog digests / no test date); R-740 (same-tag re-test) is closed, so floating-tag drift for laddered apps is handled. | **NARROWED 2026-09-30 — floating only for (1) apps with no ladder entry and (2) apps installed before v0.269.0 until their next guarded Update; owner: CC.** **Re-ranked 2026-10-03: P2→P3: narrowed to apps with no ladder entry and pre-v0.269 installs; each guarded update fixes the latter.** | — | — | CC | | **R-462** | App updates | P3 | **[P2-MEDIUM] Widen the upgrade harness beyond three apps — and the cost is dominated by FIXTURES, not by machine time.** The R-449 harness works and is proven by a red negative control (`audits/SPIKE-upgrade-test-2026-09-06.md` §1). **Costed with this run's REAL numbers rather than an estimate:** a successful edge takes **6.4 s – 305.1 s, median 71.8 s**; a FAILING edge takes **556 s**, roughly **8×**, because a negative is only honest if it waits out the full settle window; 3 apps / 11 images cost **5.07 GB**, so 53 apps naively extrapolate to **~90 GB** and, at the median, about an hour of harness time for one edge each. **THAT EXTRAPOLATION UNDERSTATES THE REAL COST BY AN ORDER OF MAGNITUDE, and that is the point of this row.** Two of the three apps needed a bespoke non-browser seed route; one needed two attempts and a discarded approach; one (bookstack) can only ever be half-proven (R-460). **Fixture time scales with apps and does not amortise.** **The decision this row is really asking for is scope, not schedule:** all 53, or only the apps a customer would lose data from, or only apps whose catalog transition is a MAJOR. **Recommended shape, NOT a design — the operator picks:** start with the apps that carry a database, because §3 measured that the abort question only ever bites there. Owner: **VIKTOR rules on scope, CC implements.** `audits/SPIKE-upgrade-test-2026-09-06.md` §5 **ROW CORRECTED 2026-09-21: the scope is NOT open and this row said it was.** It read *"VIKTOR rules on scope"*; the operator ruled on 2026-09-13 (`09` §3 decision 6) that the upgrade test goes to **ALL** apps through the nightly rotation, explicitly *not* "database apps first". What is open is the WORK, not the scope. An ORDER inside that ruling — the 15 database services first, because that is where a wrong answer costs data rather than uptime — is costed as a drill brief in `09` §6.4: legs A–E ≈ **21–34 CC-hours** plus ~25–30 GB on a scratch host, with legs C (a PostgreSQL `pg_upgrade` rehearsal) and E (one automatic night on a throwaway) the two that unblock a decision. **-- UPDATE NIGHT 2026-09-21:** **The count moved from 3 apps to 21 EDGES ACROSS 19 APPS.** The update night walked real within-a-major upstream edges on scratch guest 9202 through the product's own guarded Update, each app seeded and read back through its OWN front door with a negative control on every readback: **14 proven, 3 failed, 4 inconclusive**. Proven: actualbudget, audiobookshelf, bookstack, docmost, grafana, home-assistant, mealie, n8n, navidrome, papra, privatebin, romm, vikunja, and nextcloud's MariaDB engine major. **Ten of the fourteen printed a verbatim migration line**, so the database really was rewritten and the data still read back. Box-side fixtures for 20 apps now exist at `audits/update-night-2026-09-21/fixtures.py`, and four (actualbudget, navidrome, audiobookshelf, vikunja) are ported into `app-catalog-felhom.eu/scripts/upgrade_fixtures.py` with seven new EDGES (U1-U7) so the same edges can be run on the harness venue **with their ABORT step**, which the box deliberately does not offer. **OWED, stated so it is not mistaken for done:** the U1-U7 harness RUNS (the code is in, the runs are not), and fixtures for the four inconclusive apps, of which two (vaultwarden, zipline) cannot be seeded at all while the catalog rightly closes their sign-up (see R-624). **-- 2026-09-23, the RomM lesson is IN THE HARNESS:** `upgrade-test.py` v2 runs every edge that read back under light load for `--soak` seconds (default 600) and reads the kernel's `oom_kill` counter host-side; a kill or restart turns `proven` into `failed`, a peak over 80 % of the limit adds `memory_tight`. **Red-proof:** romm 5.0.0 → 5.3.0 on the template AS PROMOTED (512M, four workers) — seeded, migrated, read back, and then **OOM-killed at +76 s** under four light callers, verdict `failed` (Docker's own OOMKilled read true here; restarts stayed 0, which is why the walk never saw it). Ten minutes is ample for this failure; demo-hp's first kill came at two hours only because nothing was loading it. Evidence `audits/update-rulings-2026-09-23/harness/`. **— NIGHT 2026-09-23:** the bench and the box now share ONE fixture set — the box walk's fixtures run on the bench through `upgrade_boxport.py` (ported verbatim into the catalog), plus a new wishlist fixture and fixes for opengist 1.15, komga (`/api/v2/users/me`) and nextcloud (wait for `occ status`). 14 apps / 15 edges tried; 12 proven on both venues and published with their test records (`audits/DRILL-night-2026-09-23.md` Part C). **-- 2026-09-30 (day brief): three new front-door fixtures (sparkyfitness, rallly, outline; catalog `e6f3ec2`) and five apps moved on both venues — the three PostgreSQL conversions (R-463) plus bookstack 26.05.5 → 26.09.1 and kimai apache-2.57.0 → 2.67.0 (upstream dropped the `apache-` prefix; the plain tag's digest equals `apache`'s, measured). The catalog currency and the next-apps ranking: `audits/catalog-currency-2026-09-30.md` — 25 of 53 apps behind upstream inside a major, 19 across one (morning); after the session **31 apps can update themselves at night (+ nextcloud with a whole copy)**, 32 carry a proven step, 21 have no ladder.** **-- 2026-09-30 (evening, more-night-apps brief): ten more steps published on both venues** — first ladders for calibre-web, gitea, wger, crafty-controller, uptime-kuma, zipline (new front-door fixtures; gitea's installer form), within-major steps for emby, ghost, home-assistant, outline, rallly. **Night-updatable by the audit's method (a): 30 + 2 conditional at the session's start → 35 + 3 conditional** (calibre-web conditional: `files_may_change`). zipline got a two-step ladder 4.6.1 → 4.7.0 → 4.8.0 (the direct jump fails, R-742). Stopped: wanderer (the bench cannot run it, R-739). Fixture/harness defects fixed on the way: R-735, gitea's installer, the HTTPS backend on the bench, per-file names for `files_may_change`. `audits/more-night-apps-2026-09-30/` | **READY — rank P2-MEDIUM; owner: CC (scope already ruled, `09` §3 decision 6)** **Re-ranked 2026-10-03: P2→P3: 35+ apps already update by proven steps; an app without a ladder simply does not move, which is safe.** | — | — | CC | | **R-469** | App updates | P3 | **[P3-LOW] REMOVE THE ENGINE-MAJOR RULE when Slice 4 (R-448) ships — a tracked act, not a lapse.** Since 2026-09-13 `app-catalog-felhom.eu` `CLAUDE.md` rules that *until the Update button takes a verified backup as its precondition, no template may move a database-engine image across a major version* (four MariaDB, eleven PostgreSQL services), and `scripts/check-engine-major.py` (fourth row of `catalog_gates.py`, run by `.githooks/pre-push` with the push range) refuses one, naming the rule and this expiry. **Why the rule:** every `mariadb:` sidecar now carries `MARIADB_AUTO_UPGRADE=1` (R-459), so a MariaDB major move CONVERTS the customer's datadir on the next Update; PostgreSQL converts nothing and refuses to start (R-463). Either way a customer-data event with no backup in front of it. **Honest limit, not re-filed:** the gate needs a parent commit and CI fetches at `--depth 1` — the R-452 gap — so on a shallow clone the runner skips it out loud and only the hook bites. **When R-448 ships:** delete the CLAUDE.md rule, the gate's row and the gate, in one commit that cites this row; then close this. **2026-09-13 — UNBLOCKED, NOT LIFTED.** R-448 shipped in controller v0.237.0/v0.238.0 (slice 4): an update now refuses without a restorable, proven Tier-2 copy, backs up first when it is stale, takes a safety dump, and holds an app that does not come up — the precondition this rule was waiting for. **The rule stays in force until someone deliberately removes it**, which is a separate act (and is worth weighing against R-475: an app with no Tier-2 copy cannot be updated at all, so the guard does not yet cover every app a major engine move would touch). **HALF-LIFTED 2026-09-21, catalog `5ff36d098cbc`.** Slice 4 shipped 2026-09-13, so the rule's own expiry condition is met — **for MariaDB**: the four `mariadb:` sidecars have both halves they need, a verified backup in front of the Update (any tier since v0.239.0) and `MARIADB_AUTO_UPGRADE=1` whose conversion the harness WATCHED run on E3/E3b with the seeded data read back after. **PostgreSQL and MySQL stay refused** — postgres performs no `pg_upgrade` and REFUSES to start on an older major's datadir across eleven templates (R-463); a backup is a route BACK, not a conversion. The refusal text now cites R-463 instead of the shipped R-448. **R-450's second half is enforced in its place:** a MariaDB major must be the ONLY image move in its template in that commit (the bookstack `0b73e5e` shape — two migrations behind one edge). The gate now PRINTS what it allowed, by name — a lifted rule that goes quiet is a lifted rule nobody can audit. Two new decoy cases; two red-proofs, each seen to fail; 40 cases green. **What remains of this row:** the PostgreSQL half, which is R-463's to clear — see `09` §3b **Q5**. **-- NARROWED 2026-09-25 (evening), catalog `6a4a5f0`, `09` §3 decision 35:** the PostgreSQL half now passes ONE app at a time — only a template whose ladder entry for the step is proven on BOTH venues and carries `engine_conversion` (the box converts it, controller v0.273.0), as the only image move in its commit. Every other PostgreSQL app stays refused; the postgis family is judged now (it was not). CLAUDE.md rule text updated the same commit. Decoys + red-proof: `audits/night-2026-09-26/C/`. | **NARROWED** — **PARTLY CLOSED 2026-09-21 — MariaDB lifted; the PostgreSQL half stands until R-463** | — | — | CC | -| **R-607** | App updates | P3 | **[P3-LOW] A forced catalog sync answers „nincs változás" while the cache DOES change, and `catalog_images` stays stale until a separate rescan — so the update badge can be wrong for a window nobody bounds.** MEASURED 2026-09-21 on scratch guest 9202 (controller v0.260.0) while proving R-524. A real catalog move was pushed, `POST /api/sync` was invoked, and it answered **„nincs változás"** — yet the box's own cache file `/catalog-cache/templates/uptime-kuma/docker-compose.yml` **had moved to the new tag**. `Stack.CatalogImages` as served by the API stayed at the OLD value until a separate `POST /api/stacks/rescan`. **WHY IT MATTERS AND WHY IT IS NOT COSMETIC:** `CatalogImages` is the single input `stacks.CatalogOrder` compares against (v0.260.0), so for that window the badge answers from a stale catalog — it can read „Naprakész" on an app that IS behind, which is the exact failure §5.6 of `09` was written to prevent, arriving by a different door. **It also cost a measurement:** the session that found it nearly recorded a `tag-ok` badge as proof of the R-524 ahead arm when the badge was in fact stale; the honest reading came only after the rescan. **An instrument that can report an old value as a current one is not a measurement.** **TWO SEPARATE QUESTIONS, and the row does not conflate them:** (a) why the sync REPORTS no change when the working tree moved — a wrong sentence, possibly a comparison against the wrong ref; (b) whether `CatalogImages` is refreshed by the sync at all or only by `ScanStacks` on its own timer — if the latter, the staleness window is the scan interval and is bounded but unstated. **Neither was isolated** — this row records the observation, not a diagnosis. **Also observed in the same run, NOT diagnosed and folded in here rather than filed twice:** a removed app leaves `applied-compose.yml` behind in its stack directory. Stated as observed; it was not established whether that is intended. **Fix shape:** first reproduce with a loop that pushes a tag, syncs, and reads `catalog_images` on a timer, so the window is a NUMBER before anything is changed. Evidence: `audits/update-arc-2026-09-21/06-sync-after-bump.txt` and `16-r524-sync-box-ahead.txt`. **— UPDATE NIGHT 2026-09-21:** **Seen again 2026-09-21 (update night), a dozen times in one session, and for the first time with a USER-VISIBLE consequence rather than a measurement one.** On `mealie` the bump was pushed, `POST /api/sync` AND `POST /api/stacks/rescan` were both run, and the badge still read the up-to-date one (HU „Naprakesz", EN "Up to date") — the catalog had not reached `catalog_images` yet. The guarded Update was then pressed and **reported „Frissitve" after 2.1 seconds having moved nothing at all**: pinned, installed, the live compose line and `docker inspect` all still read `v3.20.1`. That is honest given a stale cache — the pin is written from *the catalog's current definition*, which was still the old one — but what the household sees is a button that says it updated them and did not. **A NUMBER, at last, which is what this row asks for:** the night's harness was changed to poll `catalog_images` until the pushed reference appears and to report how long that took; those figures are each edge's `badge_catchup_seconds`, and here they are: **4.4 s, 4.4 s, 4.5 s, 4.5 s — and 29.0 s.** The four fast ones are one sync+rescan round; the 29-second one (`nextcloud`, an engine-sidecar bump) needed **additional** sync+rescan rounds before `catalog_images` carried the pushed reference. **So the window is not a fixed scan interval — it varies by roughly 7x between edges on the same box in the same hour**, which is why a caller (or a household) cannot know when the badge is safe to read. Before tonight this row had no number at all; it now has five, and they disagree with each other, which is itself the most useful thing about them. Every drill-catalog bump of the night was followed by `POST /api/sync` answering „Sablonok naprakészek — nincs változás" while the box's cache HAD moved, with `catalog_images` staying stale until a separate `POST /api/stacks/rescan`. The night's harness therefore rescans unconditionally after every sync, which is a workaround and not a fix. **The window was still never measured as a NUMBER** — that is what the row asks for and what remains owed. **— NIGHT 2026-09-23, measured twice more:** for an INSTALLED app the stack directory's compose is the applied one, so waiting on it for a drill commit took ~7 min (chaos rounds 2 and 4); the badge's own `catalog_images` caught up in seconds. And twice a fresh box walk deployed a STALE template (the uptime-kuma "after" run, the opengist re-walk) because the deploy ran before the sync reached the file — the walks now wait for the box's template to show FROM. | **READY — rank P3-LOW; owner: CC (controller)** | — | — | CC | +| **R-607** | App updates | P3 | **[P3-LOW] A forced catalog sync answers „nincs változás" while the cache DOES change, and `catalog_images` stays stale until a separate rescan — so the update badge can be wrong for a window nobody bounds.** MEASURED 2026-09-21 on scratch guest 9202 (controller v0.260.0) while proving R-524. A real catalog move was pushed, `POST /api/sync` was invoked, and it answered **„nincs változás"** — yet the box's own cache file `/catalog-cache/templates/uptime-kuma/docker-compose.yml` **had moved to the new tag**. `Stack.CatalogImages` as served by the API stayed at the OLD value until a separate `POST /api/stacks/rescan`. **WHY IT MATTERS AND WHY IT IS NOT COSMETIC:** `CatalogImages` is the single input `stacks.CatalogOrder` compares against (v0.260.0), so for that window the badge answers from a stale catalog — it can read „Naprakész" on an app that IS behind, which is the exact failure §5.6 of `09` was written to prevent, arriving by a different door. **It also cost a measurement:** the session that found it nearly recorded a `tag-ok` badge as proof of the R-524 ahead arm when the badge was in fact stale; the honest reading came only after the rescan. **An instrument that can report an old value as a current one is not a measurement.** **TWO SEPARATE QUESTIONS, and the row does not conflate them:** (a) why the sync REPORTS no change when the working tree moved — a wrong sentence, possibly a comparison against the wrong ref; (b) whether `CatalogImages` is refreshed by the sync at all or only by `ScanStacks` on its own timer — if the latter, the staleness window is the scan interval and is bounded but unstated. **Neither was isolated** — this row records the observation, not a diagnosis. **Also observed in the same run, NOT diagnosed and folded in here rather than filed twice:** a removed app leaves `applied-compose.yml` behind in its stack directory. Stated as observed; it was not established whether that is intended. **Fix shape:** first reproduce with a loop that pushes a tag, syncs, and reads `catalog_images` on a timer, so the window is a NUMBER before anything is changed. Evidence: `audits/update-arc-2026-09-21/06-sync-after-bump.txt` and `16-r524-sync-box-ahead.txt`. **— UPDATE NIGHT 2026-09-21:** **Seen again 2026-09-21 (update night), a dozen times in one session, and for the first time with a USER-VISIBLE consequence rather than a measurement one.** On `mealie` the bump was pushed, `POST /api/sync` AND `POST /api/stacks/rescan` were both run, and the badge still read the up-to-date one (HU „Naprakesz", EN "Up to date") — the catalog had not reached `catalog_images` yet. The guarded Update was then pressed and **reported „Frissitve" after 2.1 seconds having moved nothing at all**: pinned, installed, the live compose line and `docker inspect` all still read `v3.20.1`. That is honest given a stale cache — the pin is written from *the catalog's current definition*, which was still the old one — but what the household sees is a button that says it updated them and did not. **A NUMBER, at last, which is what this row asks for:** the night's harness was changed to poll `catalog_images` until the pushed reference appears and to report how long that took; those figures are each edge's `badge_catchup_seconds`, and here they are: **4.4 s, 4.4 s, 4.5 s, 4.5 s — and 29.0 s.** The four fast ones are one sync+rescan round; the 29-second one (`nextcloud`, an engine-sidecar bump) needed **additional** sync+rescan rounds before `catalog_images` carried the pushed reference. **So the window is not a fixed scan interval — it varies by roughly 7x between edges on the same box in the same hour**, which is why a caller (or a household) cannot know when the badge is safe to read. Before tonight this row had no number at all; it now has five, and they disagree with each other, which is itself the most useful thing about them. Every drill-catalog bump of the night was followed by `POST /api/sync` answering „Sablonok naprakészek — nincs változás" while the box's cache HAD moved, with `catalog_images` staying stale until a separate `POST /api/stacks/rescan`. The night's harness therefore rescans unconditionally after every sync, which is a workaround and not a fix. **The window was still never measured as a NUMBER** — that is what the row asks for and what remains owed. **— NIGHT 2026-09-23, measured twice more:** for an INSTALLED app the stack directory's compose is the applied one, so waiting on it for a drill commit took ~7 min (chaos rounds 2 and 4); the badge's own `catalog_images` caught up in seconds. And twice a fresh box walk deployed a STALE template (the uptime-kuma "after" run, the opengist re-walk) because the deploy ran before the sync reached the file — the walks now wait for the box's template to show FROM. **Checked from source 2026-10-05 (burn-down round 2):** Diagnosed from source (both questions the row asks). felhom-controller@114ff27 controller/internal/sync/sync.go:236-241 rescans ONLY `if len(newApps) > 0 // len(updated) > 0`, and :257-258 says 'nincs változás' when both are empty. `updated` counts stack-dir copies only (copyTemplates, :447 `updated = append(updated, appName)` after a hash mismatch); for a deployed+pinned app whose catalog moved, renderSource (:457-469 table) copies the STORED definition, so the hash matches and nothing is 'upda | **READY — rank P3-LOW; owner: CC (controller)** | — | — | CC | | **R-615** | App updates | P3 | **[P3-LOW] Pointing a box at a different app catalog by `git.repo_url` alone is INERT — the box keeps fetching from the repository it first cloned.** FOUND 2026-09-21 by reading `sync.go` **before** running it, which is the only reason the update night's drill catalog worked at all. `Syncer.gitCloneOrPull` (`controller/internal/sync/sync.go:274-306`) clones **only when `/catalog-cache/.git` is absent**; on every later cycle it runs `git fetch --depth 1 origin ` + `git reset --hard origin/` **against the remote stored in the clone**, which `buildRepoURL` wrote at clone time. Changing `git.repo_url` in `controller.yaml` and restarting therefore changes **nothing**: the sync keeps pulling the old catalog and reports success. Measured: after the repoint, `git -C /catalog-cache remote -v` still read `app-catalog-felhom.eu`; the box only followed the drill repo once the cache directory was removed. **Why it matters beyond a drill:** this is the one knob that would move a box to a different or a staged catalog — for a migration, a per-customer catalog, or a rollback of the catalog itself — and it silently does not work. **Nothing is wrong with the CACHING**, which is right; what is missing is that a changed `repo_url` must invalidate the clone. **Fix shape:** on start, compare `git.repo_url` with the clone's `origin` and re-clone when they differ (or `git remote set-url` + a full fetch); log which happened. A test that changes `repo_url` under an existing cache and asserts the next sync reads the NEW repo — it fails today. Evidence: `audits/update-night-2026-09-21/04-9202-config-pre.txt`, `05-9202-follows-drill.txt`. **-- 2026-10-01 (night, new apps):** the other direction measured on 9202: after `repoint restore` a DRILL-only template (grimmory) stayed OFFERED on the app list, because its `/opt/docker/stacks/grimmory` folder (written by the drill sync) remained; removing the folder by hand ended it (`audits/new-apps-2026-10-01/B/B9-restore-live.txt`). Older drill folders (`chaoscrash`, `chaosoom`, `chaosoomb`) sit there too. The drill teardown should list and remove stack folders the live catalog does not have. | **READY — rank P3-LOW; owner: CC (controller)** | — | — | CC | | **R-683** | App updates | P3 | **[P3-LOW] Watch: after a power cut during an update's health check, the hold named an HOUR-OLD second-drive copy, not the one the update's own backup should have just made.** 2026-09-24 chaos round 3 (nextcloud, `backup_max_age: 1m`): no `backing-up` phase was seen and the hold named Tier 2 at 13:04 for an update pressed at 14:04; the pre-cut controller log was lost with the container (the runner now saves it at arm time — R-320). Round 11, the same action without a power cut, named a fresh 14:34 copy and logged the Tier-2 copy. The sentence was TRUE (it named the copy it offered); the question is why the update did not back up first. Not reproduced; watch the next power-cut drill. `audits/night-2026-09-24/E/round-03*.json`, `E/round-11-controller-pre.log` | **OPEN — P3; owner: CC (watch)** | — | — | CC | | **R-785** | App updates | P3 | **[P3-LOW] SparkyFitness is pinned 11 releases and a major behind upstream (v0.17.3; upstream v1.7.3, v1.6.0 dated 2026-07-24).** READ 2026-10-01 (`audits/visitors-2026-10-01/C/bench/C1-previous-tag.txt`). **Needs:** an update walk 0.17 → 1.x through the ladder (bench + box), after R-784 is decided. | **OPEN — rank P3-LOW; owner: CC (after R-784)** | — | — | CC | @@ -174,7 +172,7 @@ stopping line that lies. | **R-621** | App updates | P4 | **[P2-MEDIUM] A held update DESTROYS the evidence of why it failed: `failAndHold` runs `compose down`, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it.** MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: `adventurelog v0.12.1 → v0.13.0`. The new backend applied **nine Django migrations successfully** and then never listened; the update held after the full 5-minute health wait. **`Manager.failAndHold` (`stacks/update.go:723`) calls `updateCompose(dir, env, "down")`**, which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded `Logs result for adventurelog: 0 bytes returned (empty)` and `docker ps -a` held nothing at all. **What survives is the WHAT and not the WHY:** the controller line `update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy)` and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. **This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit** — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. **Why it matters beyond a drill:** the hold sentence sends the household to a restore, and after the restore the only remaining question is *should I press Update again?* — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. **Fix shape:** capture `compose logs --no-color --tail N` into the stack directory (beside `applied-compose.yml`, which already travels with the stack) IMMEDIATELY before the `down`, and surface it on the app page's hold panel or at least through the existing `/api/stacks//logs` fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt`, `state-after-hold.txt`. **FIXED in v0.262.0.** `failAndHold` now writes each service's log into `/hold-logs//compose-logs.txt` (`compose logs --no-color --tail 400`) **before** the `down` that destroys them, best-effort by design: a hold must never fail because its evidence could not be written. This is what `adventurelog` cost — nine migrations ran, the app never bound its port, and the only log that could have said why was gone before anyone looked. | **VERIFY** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the fix shape also said to show the captured hold log on the app page's hold panel; the row does not say that was done) — **CLOSED 2026-09-22 — v0.262.0: the hold keeps the app's own log before stopping it** | — | — | CC | | **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator | -## Backup & restore — 47 rows (P2 8, P3 19, P4 20) +## Backup & restore — 37 rows (P2 8, P3 13, P4 16) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -189,33 +187,23 @@ stopping line that lies. | **R-49** | Backup & restore | P3 | **[P2] The offsite capture set is ~90% cache and duplication — 1.1 GB of a 1.2 GB immich "photo backup".** Measured 2026-07-19: `immich_ml_cache.tar` **823 660 032 B (~60%)** — re-downloadable ML model weights; `immich_postgres_data.tar` **308 251 136 B (~23%)** — a raw tar of the postgres data dir that DUPLICATES the logical `.sql` dump captured beside it; `upload/backups/` **18 MB** — immich's own nightly dump, a backup inside the backup, growing daily; plus the stranded pre-v3 `dccc13fe…` tree (~36 MB) no DB has ever referenced. Actual irreplaceable content: **72 MB of originals**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S–M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: wasted space and transfer, no data risk; needs a capture-set ruling, not a sale blocker.** | — | **Evidence: `audits/DIAG-immich-restore-round2-2026-07-19.md` §4 (full byte breakdown).** This is the customer's offsite quota and transfer cost, and it lands on the Hetzner sub-account they are billed for. **Recorded, deliberately not changed** — a capture-set exclusion is a data-loss-shaped decision and gets its own ruling, not a drive-by edit. Candidates in priority order: (a) `immich_ml_cache` — pure cache, strongest case; (b) the `postgres_data` volume tar where a logical dump of the same DB is already captured (the dump is what the restore path actually replays); (c) `upload/backups/`. Likely generalises past immich into a template-classification rule about cache volumes and self-backup directories, so it should be specified against the catalog, not one app | CC | | **R-127** | Backup & restore | P3 | **The catalog's `data_key: true` flag is UNRELIABLE — at least four data-encrypting keys the catalog itself labels as encryption keys are unflagged; and the O4 restore path can regenerate a DB password that then does not match the restored data directory** | **READY (S/M)** | — | **Found by D5's Part 0, and it is why D5's boundary is `type: secret` rather than `data_key`.** Two separable legs. **(a) The misclassification.** Only 5 fields across 4 apps set `data_key: true` (`adventurelog/SECRET_KEY`, `homebox/HBOX_AUTH_API_KEY_PEPPER`, `papra/AUTH_SECRET`, `sparkyfitness/{API_ENCRYPTION_KEY,BETTER_AUTH_SECRET}`), yet `n8n/N8N_ENCRYPTION_KEY` („Titkosítási kulcs"), `wanderer/POCKETBASE_ENCRYPTION_KEY` („Adatbázis titkosítási kulcs"), `calcom/CALENDSO_ENCRYPTION_KEY` and `bookstack/APP_KEY` are unflagged — the catalog's own Hungarian labels contradict the flag. **D5 makes this non-urgent but not harmless:** everything `type: secret` now travels, so the keys DO reach the drive; what stays wrong is the **fail-closed gate**, which only refuses for `data_key` names — so if one of these is missing from both sources the restore proceeds onto data it cannot decrypt instead of refusing. Fix = flag them (`app-catalog-felhom.eu`, a catalog-only change) + a gate/test that the flag set and the label set agree. **(b) The regenerated-DB-password trap.** `internal/backup/restore_unit.go` O4 generates a replacement for any missing non-data-key secret. Proven on `postgres:16-alpine`: with PGDATA restored from the volume tar, `POSTGRES_PASSWORD` is ignored (initdb skipped), so the app fails over the compose network while the dump replay still succeeds through the container's local **trust** socket — success reported, data unreachable. v0.188.0 corrected the WARN's false claim that *"stored data is unaffected"* and scoped it, but did **not** add a guard: D5 shrinks this to the rare case (the secret was empty at capture AND absent from the guest). Real fix = either treat a DB password as fail-closed like a data key, or `ALTER USER` to the regenerated value after the volume restore. 18 DB/root-password fields are in scope; MariaDB fails loudly instead (Access denied), which is the safer half | CC | | **R-231** | Backup & restore | P3 | **`/opt/backup/scripts/` on DooPlex is unversioned host state** — found 2026-08-06 while adding the auto-memory store to the backup set. No repository tracks the scripts that protect the recovery chain, so the edit made that day (`CLAUDE_MEMORY_DIR` in `backup-config.sh`, multi-path restic call in `backup-data.sh`) exists only on the box. This is the same class the part-2 session was closing, found inside the fix for it; the change is transcribed in `felhom.eu/workspace/README.md` so it is at least *recorded*. **Two related facts, both understating current safety:** the backup destination (`/mnt/5_hdd/backup`) is on the **same physical disk** as the workspace it protects, and the DooPlex backup set has **no off-site leg** (`sync-hetzner-backups.sh` is jarrs.eu and pulls *from* Hetzner *to* DooPlex). Bringing a root-owned production backup script under version control, and deciding what installs it, is its own scoped change. | **READY** — owner Viktor | — | — | operator | -| **R-240** | Backup & restore | P3 | **A backup that covered nothing calls itself „Sikeres".** On a configured box with no app selected for off-site backup, a run reports status `ok` with the warning „Sikeres — nincs mentésre jelölt alkalmazás" — *successful* immediately beside *nothing is selected*. Measured as T4 on the final walk, 2026-08-07; flagged once before (2026-08-06) and deliberately not touched then, because the task that noticed it forbade changing that path. **It is the same rhetorical shape the project has spent a fortnight removing** — R-203's *a warning beside a success is read as a success*, R-234's *„✓ Rendben" over an app that was skipped*, R-225's *unknown rendered as zero* — one notch weaker each time, and this is the weakest and last of them. The state itself is honest and must stay `ok`: an unconfigured box reporting `incomplete` forever is its own defect, pinned by a test. **The defect is the word „Sikeres", not the verdict.** Wording such as „Nincs mentésre jelölt alkalmazás — ez a futás semmit nem mentett" says the same thing without congratulating the customer on it. | **READY** — owner Viktor | — | — | operator | -| **R-251** | Backup & restore | P3 | **The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice.** Measured on the fifth walk, 2026-08-07, on the screen the customer reaches after entering R. The snapshot carries tags `felhom-offbox,calibre-web`; the listing renders **two rows** — `calibre-web · 2026-08-07 14:57 · 12.8 MB` and `felhom-offbox · 2026-08-07 14:57 · 12.8 MB`. `felhom-offbox` is the tier's own marker tag, not an application. **The screen's whole job is to let the customer check that what is in the store is what they expect** (*"Nézd át, hogy tényleg azt találod-e itt, amire számítasz"*), and it shows them a stranger's name beside their own data and a total that is double the truth. **Cosmetic, not a data defect** — the restore page correctly offers only `calibre-web`. **Fix:** filter the marker tag out of the listing, or key the rows on the app tag. | **READY** — owner Viktor | — | — | operator | -| **R-257** | Backup & restore | P3 | **C2 — „Az offsite tároló nincs elárvult állapotban." puts an English loanword and an internal state name in front of a Hungarian household customer, and names no route.** `web/offbox_handlers.go:270` (the Go error it mirrors is `backup/offbox.go:343`). „Offsite" is untranslated; „elárvult állapot" is the codebase's own `OffboxOrphaned()` predicate surfacing verbatim. A customer who pressed a button and got this cannot tell whether something failed, whether they did something wrong, or what to do instead. **This is a refusal that is CORRECT and fail-closed and still a dead end** — the same shape R-241 recorded for `--recover-offsite-install`. **Fix shape, not a decision:** say what the customer tried to do, why it does not apply right now, and where to look — or, since this is a state they cannot reach deliberately, do not offer the action at all | **READY** — owner Viktor | — | — | operator | | **R-314** | Backup & restore | P3 | **`StopAbandon` has no web route — a customer who telephones is served by a command line.** `--abandon-stop` exists on the controller binary (`cmd/controller/main.go:86`) and refuses rather than silently no-opping, which is right. But there is no handler: the only in-product way to cancel a countdown is the customer finding their recovery code (Scenario G cancels it at the moment the code proves they still have it). **An operator who is telephoned instead has to reach a shell on the customer's machine.** Used this session on the operator's ruling, container stopped first so the running controller could not overwrite `settings.json` from memory — a sequencing subtlety that is itself an argument for a route | **READY (S) — NEW 2026-08-12, RANK 3** | R-241, R-307 | An operator-authenticated POST that calls the same `StopAbandon`, so the telephone path and the code path converge | CC | -| **R-362** | Backup & restore | P3 | **A data drive detached mid-restore is reported as „permission denied".** Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (`IsDisconnected`, used by both backup legs) and the restore path never consults it. **A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist.** Creditable in the same test: the agent re-bound the drive 5 s later, unaided. | **OPEN — MEDIUM** | — | Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. | CC | | **R-401** | Backup & restore | P3 | **Revisit the off-site integrity depth when a real store is LARGE — the default rests on ONE measurement, on ONE 134 MB store.** Controller v0.228.0 (R-399) made `--read-data-subset=100%` the default for every box. The whole justification is a single data point: `demo-hp`, 2026-08-30, 140 829 678 B / 2 651 blobs / 67 snapshots, structure 35.0 s vs 100% **39.2 s** — four seconds. Re-proven live 2026-08-31 at 38.7 s. **It does not extrapolate, and the reason is structural: the structure check's cost tracks the INDEX, a read-data run's tracks the DATA.** A 50 GB store is ~370x the data and this curve says nothing about it. **Nothing was invented from that one point** — no rotation schedule, no size threshold, no bandwidth budget — because four production designs in this project were specced against unvalidated mechanisms and all four were wrong. **`readDataSubsetRe` already accepts `n/m`**, so a rotating schedule (`1/7` on a different seventh each week) needs no parser work when the time comes; the missing input is a measurement on a large store, not code. **THE TRIGGER IS AN EVENT, NOT A DATE:** the slow-check WARN from v0.228.0 firing on any box (`integritySlowNoticeThreshold`, 5 min) — that line names the duration, the depth and this row. **WHAT HAPPENS IF NOBODY ACTS:** every box re-reads its entire store every week, however large it grows, and the first person to notice is a customer whose upload is saturated. **Whoever acts must also revisit `integrityCheckTimeout` (30 min)**, which is now the number a large store meets first. | **OPEN — WATCHING** | — | When the WARN fires: measure the curve on that store, then choose between a rotation (`n/m`), a size-conditional default, or leaving it. Do NOT choose from this row's numbers — they are the small-store case. | CC | | **R-409** | Backup & restore | P3 | **Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it.** MEASURED on demo-hp 2026-08-31 against kimai's restored unit: `manifest.json`'s `checksums` object carries sha256 for `.felhom.yml` (2 235 B), `app.yaml` (488 B) and `docker-compose.yml` (2 195 B) — **4 918 bytes of a 213 231 242-byte unit**. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. **And nothing else supplies one:** restic 0.14.0's `restore --verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and **no content hash**. **So "the restore produced correct files" is currently unanswerable by any automated means.** **What is NOT claimed here:** `restic check --read-data-subset=100%` already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | **OPEN — MEDIUM** | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's `checksums` to cover `db_dumps` and `volume_dumps` — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: `audits/SPIKE-restic-restore-test-2026-08-31.md` §Q2, §Q3. | CC | | **R-433** | Backup & restore | P3 | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01.** The one question that can move this is drafted and ready to send: **Question 1** of `documentation/runbooks/provider-questions-2026-09-01.md` — *can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?* **Neither answer leaves this row where it is:** "yes" makes per-file recovery real but permanently operator-only, and closes this; "no, full restore only" means the snapshots do not bound a single customer's exposure at all, because using them costs every other customer on the box their newer snapshots — **and R-95 becomes urgent.** Nothing here can progress without it, and it is not CC's to send (§11-D). Tracked by a dated check in the DUE-CHECKS block. **⚠ A TRAP ON THE WAY TO THIS ANSWER, found by the operator 2026-09-01 and armed against in the questions file: HETZNER HAS TWO SIMILARLY-NAMED PRODUCTS WHOSE DOCS SAY OPPOSITE THINGS.** **Storage SHARE** (a managed Nextcloud, NOT us) documents *"Currently, we only support restores for the full backup ZFS snapshot to a specific point in time"* (`docs.hetzner.com/storage/storage-share/faq/backup-snapshot/`). **Storage BOX** (ours) documents the opposite — *"You can download individual files or entire directories as usual"* (`docs.hetzner.com/storage/storage-box/snapshots/`). **A web search for the obvious phrasing surfaces the SHARE page, and it reads like a definitive NO.** Tell them apart by the giveaways: the Share page talks about *Nextcloud's data cache*, a *database dump* and the *konsoleH* interface, and never mentions Storage Box. **Why this is filed and not left as trivia: a support agent could answer Question 1 from the wrong page, and that answer would push this row to the top of the register for no reason.** Question 1 now names the product, quotes the Storage Box line, and says explicitly that we know what the Share FAQ says. **If an answer cites full-snapshot-only, check WHICH PRODUCT it is about before acting on it.** **Nothing in this repository ever leaned on the Share claim** — verified by grep at the time; the only vendor line we cite is the Storage Box one. **RE-DATED 2026-09-15 → 2026-09-22 (due-checks gate):** the operator mailbox read through the Gmail connector (`(from:hetzner OR subject:hetzner OR "Storage Box") after:2026/09/01`) holds no Hetzner reply — one match, our own `offsite_snapshots_dropped` alarm. Whether the two tickets were ever sent is not visible to CC; if they were not, sending them is the operator's act. **-- ANSWERED. Hetzner replied on ticket #2026090103040671; the operator supplied the thread on 2026-09-22 after this session searched for it and WRONGLY reported no answer (see the instrument note below and R-628).** **Q1 - can the MAIN account read individual files and directories out of a specific snapshot, without restoring the whole box?** Hetzner: *"With the main account it should be possible to download files and directories from the snapshot as it is stated in the documentation"*, citing `docs.hetzner.com/storage/storage-box/snapshots#access-to-snapshots`; and separately *"A restore of a snapshot will revert the Storage Box completely to the state it was in when the snapshot was taken."* **So the Storage BOX documentation governs, not the Storage Share FAQ** - which is exactly the trap `provider-questions-2026-09-01.md` warned the reader about, and the answer came back on the right side of it. **File-level snapshot access is a MAIN-ACCOUNT capability; the sub-account cannot do it**, which matches R-433's own measured finding (777,600 names swept from a sub-account, zero hits). **HEDGE, kept because it is in the reply: "should be possible" is not "is", and nobody has yet read a file out of a snapshot from the main account on `u629488`. That is now the measurement this row needs, and it is cheap.** **Q2 - is `--append-only` enforced by Hetzner, or taken from what the client sends?** Hetzner: *"You can use --append-only and it would look like this: `command="rclone serve restic --stdio --append-only path/to/repo" `"*, with `fluix.one/blog/hetzner-restic-append-only/`. **So it is enforced by US, by pinning a FORCED COMMAND to the SSH key in the Storage Box's `authorized_keys`** - not by Hetzner globally, and not by anything the client asks for. A key without that prefix still gets a deleting server. **This is the answer R-95 has been blocked on** and it says append-only IS expressible on a Storage Box, which R-95 records as impossible over SFTP - see R-95 and R-436. **NOT YET MEASURED, and that distinction is the whole of what this row is worth: this is a support answer, not a proof.** What is owed is a real forced-command key on `u629488`, a restic `forget --prune` through it that is REFUSED, and a backup through it that still succeeds. | **BLOCKED** — **BLOCKED-ON-PROVIDER — Question 1 of `runbooks/provider-questions-2026-09-01.md`** | — | — | CC + operator | | **R-540** | Backup & restore | P3 | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** | — | — | CC | | **R-545** | Backup & restore | P3 | **[P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository.** FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: `POST /backup/offbox/config` configures a target and can disable it (`enabled` unchecked), but nothing removes it. `POST /backup/offbox/reset` refuses unless `OffboxOrphaned()` is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in `data/offbox/`. **Why P3 and not higher:** a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. **Fix shape:** a „Távoli cél törlése" action beside the config form that clears the target and shreds `data/offbox/`, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | **READY — rank P3-LOW; owner: CC** | — | — | CC | | **R-548** | Backup & restore | P3 | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. **-- 2026-09-30, the class on a real box: demo-hp's whole-guest LOCAL tier was refused by the space preflight (R-685) 10 times since 2026-09-27** („local has 12.1 GiB free; the last archive of guest 9201 was 8.9 GiB, so a new one needs about 12.1 GiB") — its newest local archive was 2026-09-29 04:42, 30 h old at the day's read (the weekly PBS tier, last 2026-09-24, is its whole copy meanwhile). The refusal is correct and named; what is missing is that nothing gives the box the room back (the host's `local` holds the 8.9 GiB archive of the only guest plus templates on a 39 GB root). `audits/pg-last-six-2026-09-30/C/`. | **READY — rank P3-LOW; owner: CC** | — | — | CC | -| **R-552** | Backup & restore | P3 | **[P3-LOW] An interrupted-restore notice for an app that is then REMOVED stays on the restore page for ever.** FOUND 2026-09-17 by CC reviewing its own controller v0.246.0 (R-550) during live validation. The per-app notice (`Manager.opInterrupted`, persisted in `restore-status.json`) is cleared in exactly one place — `BeginRestoreOp` for that app (`internal/backup/opstatus.go`) — and `removeStack` (`internal/api/router.go`) never touches the restore record. So a household that answers „A visszaállítás megszakadt … indítsd el újra" by REMOVING the app instead of restoring it keeps a „Megszakadt visszaállítás" card about an app that no longer exists. Measured shape, not hypothetical: on 9201 the notice cleared only when homebox was restored again (08:54:50Z, card count 0) — the teardown deliberately took that path before removing it. **Fix shape:** `removeStack` clears the app's notice (a `ClearInterruptedRestore(stack)` beside the existing update-hold clear, R-491's precedent), with a wiring test. Not fixed in v0.246.0: found after the release was built; one release per repo per session. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-645** | Backup & restore | P3 | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** | — | — | CC | -| **R-675** | Backup & restore | P3 | **[P3-LOW] The unit-only restore's refusal for a file app still points to „Fájlok visszaállítása" instead of the second drive's whole restore.** `missingFileLegsRefusal` predates decision 26 (v0.269.0); when a whole copy exists on the second drive the sentence should name it. | **READY — P3; owner: CC (controller)** | — | — | CC | | **R-698** | Backup & restore | P3 | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. **-- 2026-09-30 late (decision 53):** a box now keeps only an app's running and previous image; a restore to an older version re-pulls it — as every restore already did. The limit above is unchanged. | **OPEN — P3; owner: operator (a decision), CC measures** **Operator ruling 2026-10-05 18:23: kept OPEN as a known risk to a household's restore; owner the operator; not worked on in the burn-down.** | — | — | operator | | **R-729** | Backup & restore | P3 | **[P3-LOW] An off-site target, once saved on the page, cannot be removed through the product.** MEASURED 2026-09-30 on 9202: `/backup/offbox/config` refuses an empty address and no route clears the target; the session removed its throwaway target from `settings.json` by hand, with the controller stopped (harness teardown on a scratch guest). A household that tries its own NAS and gives up keeps a disabled target forever. **Fix direction:** a „Távoli mentési cél törlése" press that clears the target (never the repository). | **READY — rank P3-LOW; owner: CC (controller)** | — | — | CC | -| **R-10** | Backup & restore | P4 | T-6E-1: DB-dump dir-fsync asymmetry (LOW, confirmed in 6E) **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-15, size XS, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **T-6E-1, confirmed in CAMPAIGN-6E.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | One-line hardening; batch with the next controller task | CC | -| **R-91** | Backup & restore | P4 | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | +| **R-91** | Backup & restore | P4 | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk **Checked from source 2026-10-05 (burn-down round 2):** Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word. | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | | **R-99** | Backup & restore | P4 | Server-side prune **never removes** a phantom snapshot. Confirmed it does NOT count them toward `keep-last` (dry-run kept 2 real + the phantom) so there is **no retention/data-loss bug** — but one accumulates per aborted upload, forever | READY (S) | — | Decide a cleanup path. Deletion on a **customer** datastore is a separate ruling — detection shipped, removal deliberately not automated | CC | -| **R-104** | Backup & restore | P4 | **An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach.** `resticStep` has `unlock --remove-all` (`internal/backup/offbox.go:634-648`) but `ensureOffboxRepo`'s probe fails first, `classifyResticProbe` (`:77-93`) has no lock case → `"other"` → fail-fast; `ClassifyOffsiteFailure` likewise, so the operator is told *„A távoli mentés ismeretlen okból nem sikerült"* for a precisely-known, self-healable condition **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size S, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY STALE, checked against live source 2026-08-22 — migrated as written per the task rule, with the staleness named rather than edited away.** The self-heal this row calls unreachable was built: `resticStep` escalates to `unlock --remove-all` and retries once (`internal/backup/offbox.go:763-768`), and `unlockStale` runs before every off-site run and restore (`:1274`, `offbox_restore.go:261`). Its premise that the probe fails first is also doubtful: the probe is `restic cat config`, a read that takes no lock. **What REMAINS true:** `ClassifyOffsiteFailure` (`offbox.go:179-193`) still has no lock case, so if a lock ever did survive both layers the customer would still be told an unknown reason. Re-rank on that basis, not on the original text.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Was C9-F3.** Reachable by any interruption — container restart, OOM, network drop, host reboot mid-backup. The tier stays dead until a human runs `restic unlock --remove-all`. Flips: the offsite row in map §C; `07` §8 row 15 | CC | | **R-164** | Backup & restore | P4 | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | | **R-213** | Backup & restore | P4 | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC | | **R-246** | Backup & restore | P4 | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor **Folded R-248 2026-10-03** (the same ruling: give `stale_at` a visible, evidence-bearing setter, or retire it). | — | — | operator | -| **R-256** | Backup & restore | P4 | **C2 — „A mentéskezelő nem elérhető." names no route at all.** `web/offbox_handlers.go:47` and `:181` flash this to the customer on the off-site backup surface. It states an internal component's unavailability in the operator's vocabulary („mentéskezelő" = the backup Manager object), gives no reason the customer can act on, and names no next step — not "try again in a few minutes", not "contact support", not a page to go to. **Contrast, in the same subsystem and shipped the same week:** R-252's fix reads *„Meghajtók", „Meglévő meghajtó csatolása". Utána gyere vissza ide.* — a route. **Severity is low and stated so it is not over-ranked:** the condition is a nil backup manager, which on a healthy box does not occur; this is about the copy, not a broken path. Found by the C2 sample (19 refusals on the recovery/restore/offbox surface; **~202 of the repo's 221 refusal strings were NOT examined**) | **READY** — owner Viktor | — | — | operator | | **R-279** | Backup & restore | P4 | **There is no operator-triggerable off-site backup.** The only route to `POST /backup/offbox/run` is the customer's own dashboard session; `signed_jobs` carries opaque operator-SIGNED blobs and the hub holds no signing key. This cost the rehearsal a stop: preparing the run needed one off-site push and there was no operator path to it. Sibling of **R-177** (no operator-triggerable fill check) | **READY (XS) — NEW 2026-08-09** | — | Same shape as R-177; solve both together | CC | | **R-336** | Backup & restore | P4 | **The offsite DR endpoint is polled about once per second, and that is what turned a slow leak into an outage.** ep0's PBS proxy served **~85,000 requests/day** — a flat **3,538/hour**, every hour, from two boxes: `74,445 GET /api2/json/admin/datastore` (`libwww-perl`, i.e. PVE's `pvestatd`) and `73,171 GET /admin/datastore/felhom-offsite/status` (`proxmox-backup-client`). Two pollers asking substantially the same question at the same rate. On 2026-08-18 this walked a connection leak in the proxy to its 1024-fd soft limit in **14 days**, wedging the offsite tier for 9½ hours (`audits/INCIDENT-ep0-pbs-fd-exhaustion-2026-08-18.md`). The `LimitNOFILE=65536` drop-in applied that morning raises the ceiling **but does not fix the leak** — it converts a fortnightly outage into a multi-year one, which is mitigation, not a fix. A DR endpoint that is written to weekly does not need to be asked about every second | **READY (M) — NEW 2026-08-18** | — | **CORRECTED 2026-08-18 (evening) — the easy lever named here does not exist.** This cell used to read *"PVE storage status is the prime suspect, and its interval is tunable"*. **The first half is right and the second half is false.** `pvestatd` stats EVERY configured storage on each 10-second cycle, and Proxmox staff have stated the interval is not designed to be configurable — so there is no knob to turn down. The only lever PVE actually offers is disabling the storage entry (`pvesm set --disable 1`) around the backup window, and that is **substantially more than a tuning knob**: it collides with `felhom-agent/internal/pbsdr/manager.go`'s health model, where an inactive-but-existing entry drives the consume-the-one-time-secret recovery path. So the fix is a design question (does the hub still need a 15-minute fill reading at all, given R-339 now reports reachability separately?), not a config edit. **Doc-only correction — no agent code was changed.** The remaining step is unchanged: cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing — the positive observable, per standing rule 3. **Baseline measured 2026-08-18, and the FIRST measurement published was WRONG.** The initial "~85/day, matching the ~73/day implied by the failure" came from a single 17-minute window whose delta was **one descriptor** — a sample of one cannot carry a daily rate, and the agreement that made it feel solid was coincidence. **Re-measured over two independent windows the same morning: 183/day (31 min) and 200/day (5.6 h)** — ~2.6x the published figure, putting the runway to the 65536 ceiling at **~357 days, not the ~2 years first claimed**. **And the named mechanism is the minority one:** across that window `CLOSE-WAIT` held flat at 1 while `ESTAB` grew 45→49 — *all* the growth was established connections, and at the wedge the split was 1011 ESTAB / 543 CLOSE-WAIT. **The fix must target connections the proxy never reaps, not just `CLOSE-WAIT` sockets.** The PBS 4.2.5-1 upgrade (2026-08-18) did NOT change the slope and was never expected to — see R-341 **SPIKE 2026-08-20 — THE PREMISE OF THIS ROW DOES NOT SURVIVE MEASUREMENT, and that is a change in what the row IS, not new evidence on it.** `audits/SPIKE-ep0-established-connections-2026-08-20.md`. **The leak is OURS, and the poll rate is not what feeds it.** Every one of the 388 leaked descriptors is an ESTABLISHED connection held open by **`felhom-agent`** on the boxes — 194 on each, `ss -tnp` naming a single PID per box, and **zero** held by `pvestatd` or `proxmox-backup-client`. Confirmed independently from ep0's access log over the same 46.18 h window: `libwww-perl` (pvestatd) **81,192 requests -> 0 descriptors**, `proxmox-backup-client` **80,061 requests -> 0 descriptors**, `Go-http-client/1.1` (the agent) **387 `/snapshots` calls -> 388 sockets — one per call, within one**. So **162,404 requests, 99.5% of the traffic, produce 0% of the leak.** **Mechanism, named from source:** `felhom-agent/internal/pbs/client.go:56-60` builds `&http.Transport{TLSClientConfig: tlsCfg}` — a composite literal, so `IdleConnTimeout` is the zero value = **no limit** (`http.DefaultTransport` sets 90 s; a literal does not inherit it) — and `cmd/felhom-agent/main.go:1486` (`pbsTargetsFromPVE`) builds **a fresh client every cycle**, as its own doc comment states. Each cycle therefore strands one idle keep-alive connection in a transport nothing ever closes; `CloseIdleConnections`/`IdleConnTimeout`/`MaxIdleConns` appear **nowhere** in the agent repo. Cadences reconcile without fitting: 900 s hub poll (184.7 cycles) + 6 h `DefaultVerifyCadence` (7.7 cycles) = 192.4 predicted vs **194 observed per box**. **CONSEQUENCE — RE-RANK.** The remaining step recorded above ("cut the poll rate, then confirm the fd count stops climbing") **would have produced a null result and read as a failed fix.** Cutting the Proxmox poll rate removes ~99.5% of ep0's request load and **zero** descriptors. The poll rate is still wrong on its own terms — 85,000 requests/day to a weekly-write DR endpoint — but it is now a **scaling/cost item, not the leak fix**, and the leak fix is **R-344**. **Q3 (is the leak proportional to the request rate?) is PREDICTED not-proportional and NOT YET MEASURED** — Phase C is held at STOP 1 with its prediction pre-registered in `evidence-ep0-established-connections-2026-08-20/phaseC-prediction.txt`. Do not record a proportionality verdict here until that window has run. **RE-SCOPED 2026-08-20 — THIS ROW IS NO LONGER A LEAK FIX, AND ITS RECORDED NEXT-STEP WOULD HAVE "FIXED" NOTHING WHILE LOOKING LIKE A FAILED FIX.** That near-miss is the reason the spike-first rule exists and it is kept here deliberately. The old next-step read: *cut the poll rate by whatever means survives that question, then confirm the fd count between restarts stops climbing.* Had it been executed, the fd count would have kept climbing at the same ~200/day, the poll reduction would have been recorded as ineffective, and the real defect — **ours, in `felhom-agent`, R-344** — would have been further from being found, not closer. **Measured 2026-08-20:** `pvestatd` (`libwww-perl`) and `proxmox-backup-client` made **162,404 requests** in a 46 h window and leaked **zero** descriptors; the agent made 811 and leaked **388**. The fix (agent 0.130.0) took ep0 from 388 accumulated descriptors to its **baseline of 17**, with the poll rate completely unchanged — 85,000/day before and after. **WHAT THIS ROW ACTUALLY IS NOW — a SCALING concern, still worth fixing on its own merits:** ~85,000 requests/day to a DR endpoint that is WRITTEN TO WEEKLY, from two boxes. That is ~42,500/box/day, so **at fifty customers it is ~2.1 million requests/day — about 25 requests/second, constantly, against a CX33**. The design question is unchanged and is still the hard part: does the hub still need a 15-minute fill reading at all, given R-339 reports reachability separately? And the lever remains awkward — `pvestatd` stats every configured storage on each 10-second cycle with no tunable interval, so the only PVE-side lever is disabling the storage entry, which collides with `felhom-agent/internal/pbsdr/manager.go`'s health model. **NEW ACCEPTANCE CRITERION, since the old one is void:** the fd count is NOT the observable for this row any more — that belongs to R-344 and is already satisfied. Measure the REQUEST RATE at ep0's access log, and state the projected rate at the target customer count. | CC | -| **R-365** | Backup & restore | P4 | **An overdue abandonment countdown renders its past due-date in the future tense.** With the terminal step due and the daily sweep not yet run, the card reads „A kérésed szerint a korábbi távoli mentéseidet **2026-08-20** napján véglegesen töröljük" — on 2026-08-21. The window is up to ~29 h in production (due moment → next 05:10 sweep). | **OPEN — LOW** | — | Say "due, will run at the next daily sweep" once the date has passed. | CC | | **R-375** | Backup & restore | P4 | **A PBS datastore signal was noted and explicitly not filed.** `audits/REPORT-ep0-pbs-upgrade-2026-08-18.md:171`: *"Likely the namespace-scoped token lacking datastore-level audit. Not filed; noted here."* Recorded so the note has a number and stops depending on someone re-reading that report. **Age when filed: 4 days.** | **OPEN — LOW** | — | Confirm the cause on ep0 the next time it is touched; it is a read-only check. | CC | | **R-526** | Backup & restore | P4 | **[P3-LOW] A host delete cannot release only the customer's ep0 PBS token: the endpoint's one removal op destroys every backup group too.** MEASURED 2026-09-15 from source: `tenantsync.Deprovision` „DESTROYS the customer's PBS namespace, all its backup groups, and its token". The task asked for „PBS token elengedése" on host delete; building it needs a new token-only op in the ep0 tenantsync script — a new operation on a protected box. Not built. R-511's adopt path makes the kept token usable instead. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (new ep0 op yes/no), CC (build)** **Re-ranked 2026-10-03: P3->P4: operator teardown op on a protected box; adopt path covers the need.** | — | — | operator | | **R-541** | Backup & restore | P4 | **[P3-LOW] There is no path to move a customer between off-site boxes, or from shared to dedicated.** Read from source 2026-09-16: provisioning is idempotent-reuse keyed on the customer (`shared already provisioned for tester-1 (subaccount 311327)`), and a dedicated deprovision destroys the repository — so "move this customer" has no safe route today. It becomes reachable the moment R-540's second pool box exists, or when a customer outgrows the shared model. **Needs:** a move that copies the repository, re-keys, and only then releases the old sub-account — a new mechanism nobody has measured. | **READY — rank P3-LOW; owner: CC (hub) — design first** **Re-ranked 2026-10-03: P3->P4: needs a second pool box or an outgrown customer first; operator-only and far off (0.3% full).** | — | — | CC | @@ -226,7 +214,7 @@ stopping line that lies. | **R-832** | Backup & restore | P4 | **ep0's copy in a place outside both Hetzner and the operator's home (roadmap).** Today DooPlex (the operator's home) holds it (decision 71). A Hetzner Storage Box would share a provider with ep0 and with every household's file backups, and cannot run PBS, so the copy could not be verified or restored from directly. | **DEFERRED — later, if the product grows** | — | — | operator | | **R-878** | Backup & restore | P4 | **A catch-up (R-871) runs the database-dump leg in the DAY, and that leg stops an app with a volume for its copy — the household may notice the stop, and a large volume makes it longer.** MEASURED 2026-10-05 on demo-felhom: the catch-up at 08:25:02 stopped opengist, copied 182.5 KB, started it again — about 1 s, then a few seconds of `health: starting`; the night does exactly the same, unseen. Nothing measured for a large volume. Fix direction (if it matters): skip the volume copy of a running app in a DAYTIME catch-up and leave it to the next night, or warn. `audits/catchup-2026-10-05/partA/live-demo-felhom.txt` | **READY — owner: CC** | — | measure a large volume first | CC | -## Storage & devices — 9 rows (P3 5, P4 4) +## Storage & devices — 8 rows (P3 5, P4 3) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -234,11 +222,10 @@ stopping line that lies. | **R-330** | Storage & devices | P3 | **Disk health Phase 2 — the three SMART attributes the wire does not carry.** The failing drive's most telling counter was **187 `Reported_Uncorrect`**, sitting at normalized **1** against threshold **0** with a raw count of **1001** — one point from failing and structurally unable to get there. Also wanted: **199 `UDMA_CRC_Error_Count`** (cabling) and **188 `Command_Timeout`**. None are on the agent→controller wire today, so v0.215.0's ladder could not use them. Phase 2 also persists periodic SMART **samples** (the right home is `metrics.MetricsStore`, NOT the Phase-1 state file, which is one record per disk and must stay that way). **This is a declared WIRE change, so under the G-1 gate the hub must model the new fields in the SAME session** — that is precisely why it was kept out of Phase 1, where it would have turned a one-word severity fix into a three-repo change | **READY (M) — NEW 2026-08-14** | R-328 (closed) | Add 187/199/188 to the agent's `SmartSummary` + hub model in one session; then persist samples | CC | | **R-332** | Storage & devices | P3 | **The new Hiba-from-counters path has never fired on real hardware.** v0.215.0's whole point is a verdict the product could not previously reach, and it is proven only against the committed fixture's values in unit tests (12 scenario groups, 11 of 12 red-proofs failing as required). The live validation on demo-hp proved the **negative** — three healthy disks still read Rendben across the deploy, no false alert — and the **severity wire** end to end, but no live disk has actually reached Hiba. **This is the honest gap and it must not be closed by pointing at the fixture tests**: the drive that produced the fixture is in DooPlex, which is Tier 2 and never a drill target, and the demo boxes are all-flash and healthy | **WATCHING — NEW 2026-08-14, NARROWED same day.** One item originally in this gap is now PROVEN LIVE: the **persisted state surviving a controller restart**. The v0.215.0→v0.216.0 redeploy destroyed and rebuilt the container, and the new one read back a `changed_at` written by the PREVIOUS version (`2026-08-14T07:23:14.640216851Z`, still intact at 09:31:35Z) instead of re-baselining — Scenario L on real hardware, not just the production-path unit test. **What remains unproven is the verdict itself, plus the stronger restart half: an already-ALERTED disk not re-alerting** | a real degrading disk, or an injection harness | **Closing condition:** a live disk reaching Hiba from counters, OR a deliberate injection through the REAL pipeline (agent `/disks` → controller check → hub event), not a hand-set verdict | CC | | **R-542** | Storage & devices | P3 | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** | — | — | CC | -| **R-756** | Storage & devices | P3 | **[P3-LOW] On 9202, "remove with drive data" refuses calibre-web with 409 „…/scratch_hdd/userdata/calibre-web tárhely jelenleg nem elérhető", while the controller container lists that folder.** MEASURED twice on 2026-10-01 (`audits/lockouts-2026-10-01/B/B1…`, `audits/calibre-name-and-prune-2026-10-01/A/A1…`): `POST /api/stacks/calibre-web/remove` with `remove_hdd_data` → 409; `docker exec felhom-controller ls -ld /mnt/felhom-drives/scratch_hdd/userdata/calibre-web` → the directory (dated 2026-09-22). The walk then removed the app keeping the data (R-442's fail-closed answer). Either the drive is not a registered drive on this scratch box (a test-venue artefact) or the resolver reads another path than the one it names. Not measured which. **-- 2026-10-01 (night, new apps):** the same 409 for Grimmory (`remove_hdd_data`), and a drive app's per-app backup is NOT offered for restore after a remove — the controller logged `drive … not mounted — skipping ensure (held by drive gate)`; the unit lives on that drive. So on 9202 checklist row 2.5 cannot be shown for any drive app until its scratch drive is a registered drive (`audits/new-apps-2026-10-01/box/grimmory/restore-why.txt`). | **OPEN — rank P3-LOW; owner: CC** | — | — | CC | +| **R-756** | Storage & devices | P3 | **[P3-LOW] On 9202, "remove with drive data" refuses calibre-web with 409 „…/scratch_hdd/userdata/calibre-web tárhely jelenleg nem elérhető", while the controller container lists that folder.** MEASURED twice on 2026-10-01 (`audits/lockouts-2026-10-01/B/B1…`, `audits/calibre-name-and-prune-2026-10-01/A/A1…`): `POST /api/stacks/calibre-web/remove` with `remove_hdd_data` → 409; `docker exec felhom-controller ls -ld /mnt/felhom-drives/scratch_hdd/userdata/calibre-web` → the directory (dated 2026-09-22). The walk then removed the app keeping the data (R-442's fail-closed answer). Either the drive is not a registered drive on this scratch box (a test-venue artefact) or the resolver reads another path than the one it names. Not measured which. **-- 2026-10-01 (night, new apps):** the same 409 for Grimmory (`remove_hdd_data`), and a drive app's per-app backup is NOT offered for restore after a remove — the controller logged `drive … not mounted — skipping ensure (held by drive gate)`; the unit lives on that drive. So on 9202 checklist row 2.5 cannot be shown for any drive app until its scratch drive is a registered drive (`audits/new-apps-2026-10-01/box/grimmory/restore-why.txt`). **Checked from source 2026-10-05 (burn-down round 2):** Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH is compared to a registered storage path (api/router.go:574-575, web/handlers.go:3049, storage_handlers.go:528), i.e. a drive root. The message names .../scratch_hdd/use | **OPEN — rank P3-LOW; owner: CC** | — | — | CC | | **R-25** | Storage & devices | P4 | **Device-node TOCTOU hardening (drive init).** Graduate the controller v0.141.0 Observation: the `format → resolveEnrollUUID(path) → AssignDisk(uuid)` sequence has a narrow /dev-re-enumeration window (agent-guarded on the destructive format via anti-retarget durable-id; benign fs-UUID mount). Bind resolve+assign to the format's durable-id so the mount can't target a moved node. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | From the v0.141.0 F6 commit's security-review finding (`felhom-controller` REPORT). Low real risk (single-operator, agent-guarded), but cheap to close | CC | | **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC | | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | -| **R-568** | Storage & devices | P4 | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: `/dashboard` fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (`audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt`). `diskHealthRows` (`disk_health.go` L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. **Fix shape:** sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic.** | — | — | CC | ## Security & access — 24 rows (P2 2, P3 20, P4 2) @@ -255,7 +242,7 @@ stopping line that lies. | **R-270** | Security & access | P3 | **R-268's stated rotation recipe is incomplete: the controller never re-reads `bootstrap.json`'s `local_api`, so a rotation leaves the agent channel dead across restarts.** `bootstrap.ensureLocalAPI` returns early when `cfg.LocalAPI.Endpoint != ""` — by design it FILLS an absent block and never refreshes a present one — so the token the controller uses lives in its own `controller.yaml`, not in the mount. Proved live 2026-08-09: two controller restarts after a correct `bootstrap.json` rotation, still `HTTP 401`; the channel came up only once `local_api.token` was written into `controller.yaml`. The neighbouring `DetectEndpointDrift` compares the ENDPOINT and deliberately does not compare the token (*"a token mismatch is a different failure"*), so this shape is knowingly unmodelled. Parent question — which file is authoritative — is **R-78** | **READY (S) — NEW 2026-08-09** | — | Either teach the drift detector the token, or make the rotation path write both files. Correct the R-268 row's recipe either way | CC | | **R-275** | Security & access | P3 | **`--uninstall` leaves five 0600 `agent.json.*` credential backups, and the reinstall hands them to the new service account.** `/etc/felhom-agent/` survives with `agent.json.{campaign8-before,campaign9-before,campaign9-prev,pre-e-target-move,pre-prunegate.bak}`, each carrying a 64-char `hub.api_key` and a 59-char `proxmox.token`. The teardown claims to remove *"config (+ its .bak backups)"* and `scripts/CHANGELOG` F1 records *"uninstall now purges the agent config's `.bak*` siblings (one held a live hub api_key)"* — **that fix does not match the filenames in use, and it misses `agent.json.pre-prunegate.bak`, a file that literally ends in `.bak`.** **Exposure assessed, not assumed:** these are SUPERSEDED — the orphaned key hashes to `a5d2222a…`, the hub's current demo-hp key to `8c59d1b6…`, and the Proxmox token was deleted by the same uninstall. **But the reinstall recreates `felhom-agent` at uid 999, the same uid the deleted account had**, so three of the backups become the new account's files — verified readable as `felhom-agent`. A fresh install's service account inherits read access to the prior install's credentials; superseded today, live if the backups were recent (R-179's precedent). **Also left, undeclared:** `/etc/felhom/{.bootstrap-done,appliance-pairing-code}`, `felhom-bootstrap.service` + `/usr/local/sbin/felhom-bootstrap.sh`, the `vmbr9` stanza in `/etc/network/interfaces`, and `/etc/sudoers.d/felhom-agent.bak-pre-e2a` (21 KB — **INERT: sudo skips dotted filenames, verified with `sudo -l -U felhom-agent`; `visudo -c -f` parsing it OK is NOT evidence sudo loads it**) | **READY (S) — NEW 2026-08-09** | — | Purge by directory, not by glob; and do not let a new service account reuse a uid that owns old secrets | CC | | **R-276** | Security & access | P3 | **RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing.** After `--uninstall` on demo-hp, `wg-quick@wg-felhom` is **enabled and active**, `/etc/wireguard/wg-felhom.conf` present, handshake to `167.233.158.164:443` **52 s old**, counters 5.86 GiB in / 2.48 GiB sent. It appears in **neither** the WIPED nor the KEPT list, though the BYO install disclosure names it prominently on the way in (*"an OUTBOUND WireGuard tunnel to the Felhom hub"*). A host told to leave Felhom retains a live network path into Felhom infrastructure, its hub-side peer registration intact, and nobody is told | **READY (S) — NEW 2026-08-09** | — | Tear the tunnel down and deregister the peer, or list it under KEPT with the reason and the removal command | CC | -| **R-338** | Security & access | P3 | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | +| **R-338** | Security & access | P3 | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it **Checked from source 2026-10-05 (burn-down round 2):** nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246). | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-616** | Security & access | P3 | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** | — | — | CC + operator | | **R-717** | Security & access | P3 | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). | **OPEN — P3; owner: CC** | — | — | CC | | **R-747** | Security & access | P3 | **[P3-LOW] A stranger can lock the household out of mealie with five wrong logins.** MEASURED 2026-09-30 on 9202 (the R-741 proof): after the install hold opened, a stranger's default-login tries were refused (401) and after five of them mealie answered 423 (locked) to every login — the generated, correct password included. mealie's own brute-force guard, on an app published on the internet; the setup gate and the install hold do not cover an app after its first setup. Not measured: how long the lock lasts. **Needs:** measure the lock's length; decide whether the page tells the household what to do. `audits/night-rulings-2026-09-30/C/C3-mealie-poll.txt` **-- 2026-10-01:** Measured at v3.28.0 (source + 9202): 5 wrong logins lock the ACCOUNT (not the IP) for `SECURITY_USER_LOCKOUT_TIME` hours (default 24); the lock is lifted by an hourly job; an admin can unlock others via `POST /api/admin/users/unlock`, but the household's only admin is the locked account. Both login names are public (`admin`, `changeme@example.com`). **Fixed** (`09` §3 decision 57, decided by CC unattended — operator may reverse): `SECURITY_USER_LOCKOUT_TIME=1`, catalog `a4597cd`; on 9202 the right password answered 423 for 120 min, then 200; a wrong one still 401 (`audits/rulings-2026-10-01/C/`). **Left:** a stranger can renew the lock every hour (per account, public name — option (d) in decision 57 would end that); the page does not tell the household why it is locked; an INSTALLED mealie takes the new setting only when its compose is rendered again (not measured which act does that). | **NARROWED — the lock is 1–2 h; renewal and the page remain; owner: CC** | — | — | CC | @@ -286,7 +273,7 @@ stopping line that lies. | **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC | | **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC | -## Monitoring & notifications — 23 rows (P2 3, P3 14, P4 6) +## Monitoring & notifications — 21 rows (P2 3, P3 12, P4 6) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -297,12 +284,10 @@ stopping line that lies. | **R-271** | Monitoring & notifications | P3 | **The `agent_channel_unauthorized` alarm can never be closed, because its own prescribed remedy is what silences the recovery.** `channelhealth.Checker.Check`'s UP branch notifies only when `prev != "" && prev != "up"`; a controller restart resets `state` to `""`, so an unseeded→up transition is silent by construction. The alert text says *"token stale/rotated (**re-bootstrap**)"* — i.e. restart the controller — so **following the instruction guarantees no recovery event.** Observed live 2026-08-09: two `agent_channel_unauthorized` errors on the hub (one `sent`, one `suppressed` by the 1 h operator cooldown) and **nothing afterwards**, though the channel came up 3 minutes later and stayed up. The down side is deliberately asymmetric (F2: a born-down channel alerts on cycle 1); the up side never got the matching treatment. Customer dashboard is fine — `SetDashboard` reflects current state every cycle. It is the OPERATOR's trail that ends on "down" | **READY (S) — NEW 2026-08-09** | — | Notify on unseeded→up when the previous *persisted* state was down, or seed from the hub's last event | CC | | **R-333** | Monitoring & notifications | P3 | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) | | **R-340** | Monitoring & notifications | P3 | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc//fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC | -| **R-363** | Monitoring & notifications | P3 | **The fill watcher runs once a day, so a filesystem that fills at 03:31 goes unannounced for ~24 h while the backup is already refusing apps.** `sched.Daily("fill-watch", "03:30", …)` (`cmd/controller/main.go:1092`) plus one startup check. Proven 2026-08-21 23:17: the 69 GB filesystem carrying the Docker data-root, the system namespace and ALL 40-class app data was filled to 99% / 1.2 GiB free; the backup reserve refused `kimai` per app and the hub received `recovery_unit_capture_failed` (error) naming the filesystem, **and the fill watcher said nothing at all**. The package comment says it "warns the CUSTOMER that a filesystem is filling, BEFORE anything fails"; at a daily cadence it frequently cannot. | **OPEN — MEDIUM** | — | The reserve already computes the same numbers every run. Let the watcher share that reading rather than owning a separate daily one. | CC | | **R-388** | Monitoring & notifications | P3 | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor | | **R-435** | Monitoring & notifications | P3 | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** | — | — | CC | | **R-521** | Monitoring & notifications | P3 | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** | — | — | CC + operator | | **R-522** | Monitoring & notifications | P3 | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** | — | — | CC | -| **R-547** | Monitoring & notifications | P3 | **[P3-LOW] A disk that fills and empties between sweeps is never mentioned to anyone: `disk_critical` is defined at ≥95 % used, but the fill-watch runs once a day.** MEASURED 2026-09-17 (chaos night) on a fresh box (`tester-1-022354`, controller 0.245.0): the customer guest’s root filesystem was held at **96 % for ten minutes** (29 G used, 1.5 G free) and **no alarm of any kind fired** — checked twice, once by the round’s own runner and once independently after the fill was released. Cause, established from the ladder BEFORE the round rather than after: `fillwatch` runs **daily at 03:30** plus once ~90 s after a controller start, so a ten-minute window contains no check unless a restart lands inside it. The timing here was almost comic — the controller restarted at 21:28 after the previous round’s power cut, so its one opportunistic check ran about **twenty seconds before** the disk filled. Meanwhile all twelve apps kept serving and the background household loop logged 12 operations with 0 failures, so nothing else would have hinted at it either. **This is the ladder working as designed, not a missed alarm** — the fill-watch is a daily sweep, not a monitor. It is filed because the honest answer to „would the household be told their disk is full?” is **no, unless the controller happens to restart while it is full**, and that answer is not written down anywhere. **Fix shape (one of):** sample the fill more often than daily (a cheap `statfs` on the 5-minute health pass would do it); or say plainly in `08-alarm-ladder.md` that a transient full disk is out of scope. Evidence: `audits/evidence-chaos-night-2026-09-17/round-3.txt`. | **READY — rank P3-LOW; owner: CC** | — | — | CC | | **R-585** | Monitoring & notifications | P3 | **[P3-LOW] Six event producers still send Hungarian only, so an English household can see one Hungarian line in some mails.** FOUND 2026-09-18 by localisation slice 3 Part B (R-558, controller v0.256.0), which converted 19 of the customer-facing producers and left these: `backup_failed`, `db_dump_failed`, `backup_integrity_ok`, `backup_integrity_failed`, `offbox_enlarge_blocked` and `local_api_endpoint_drift`. **Why they were left:** each receives its sentence already FINISHED from another package, so the key and its arguments no longer exist by the time the notifier sees it — converting them means changing their callers, not the notifier. The 15 operator-tier types are deliberately excluded and are NOT part of this row: the operator reads Hungarian. **Why it matters more than it looks: `offbox_enlarge_blocked` has no `customerMessages` entry on the hub**, so its raw sentence IS the household's mail rather than an extra line under a translated headline — for that one type an English household gets a wholly Hungarian mail, not a mostly-English one. **Fix shape:** push the key and its arguments down from each caller (the shape slice 2 release B already used for errors, `util.MsgError`), then add each to the `convertedProducers` table in `internal/notify/message_customer_test.go`, which is the list both language tests walk. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-724** | Monitoring & notifications | P3 | **[P3-LOW] The status pages disagree with each other in small ways a household notices.** MEASURED 2026-09-29 on a fresh box: Beállítások reads „Mentés ütemezés 02:30 / 03:00" while Biztonsági mentés says 02:30 / 03:30 / 04:15 / 04:30–08:30; Beállítások shows a raw `2026-09-29T19:22:30Z`; its „Helyi cím (LAN)" and „Átjáró" read „nem állapítható meg" on the page the household is asked to read out for remote help; the backups overview still says „Következő mentés — 0 órája" (the age of the LAST run, under the word "next" — noted 2026-09-14, never filed); the dashboard shows the backup at 19:33 where the backup pages say 21:33 (R-500, still reproducing). **FIXED 2026-09-30 (controller v0.283.0):** Beállítások shows the real schedule and a local time; the dashboard's last backup is local time (R-500); „az előző N órája készült" replaces „0 órája". **NARROWED — remaining:** „Helyi cím (LAN)" / „Átjáró" read through the file-sharing container, so a box without it says „nem állapítható meg" — not a text fix (the read needs another path). | **NARROWED — the LAN/gateway read only; owner: CC** | — | — | CC | | **R-177** | Monitoring & notifications | P4 | **There is no operator-triggerable "run the fill check now" path.** `fill-watch` is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller | **READY (S) — NEW 2026-08-02** | — | **Noticed while live-validating R-167 on 9201, not by a failure.** It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. **Partially mitigated already** — v0.191.2 makes every run log a positive observable (`checked N filesystem(s), M unreadable/skipped, K notification(s)`), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has `GetJobs` but no run-now, so this is a general affordance, not a fill-watch one — **scope it as "run a named scheduler job now", operator-gated.** **ID established free:** `grep -ro "R-177\b" documentation/ *.md` → 0 hits | CC | @@ -311,8 +296,8 @@ stopping line that lies. | **R-337** | Monitoring & notifications | P4 | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC | | **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — operator decision** | — | — | operator | | **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** | R-871 | — | CC | -| **R-886** | Monitoring & notifications | P3 | **DooPlex's Alertmanager cannot write its own state since the Longhorn restart of 2026-10-05 13:20Z** — every 15 min `Running maintenance failed … open /alertmanager/nflog.…: permission denied` (and the same for `silences`), 8 times by 14:21Z; the pod was recreated 13:20:48Z by that restart. Mail still goes out (`alertmanager_notifications_total{integration="email"}` 5 → 6, `failed_total` 0, 14:23Z), but a silence set now and the record of what was already sent do not survive the next pod restart — so a restart can re-send every active alarm or drop a silence. Likely collateral of the restart (volume ownership on re-attach), not measured. | **OPEN** | — | Compare the volume's file owner with the pod's `securityContext` (`fsGroup`/`runAsUser`); fix in homelab-manifests; prove with a silence that survives a pod restart | operator | -| **R-884** | Monitoring & notifications | P4 | **ArgoCD app `monitoring` shows `Deployment/prometheus` OutOfSync** (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. | **OPEN** | — | `argocd app diff monitoring` (or the CR's resource diff) to see what differs, then decide git or live | operator | +| **R-886** | Monitoring & notifications | P3 | **DooPlex's Alertmanager cannot write its own state since the Longhorn restart of 2026-10-05 13:20Z** — every 15 min `Running maintenance failed … open /alertmanager/nflog.…: permission denied` (and the same for `silences`), 8 times by 14:21Z; the pod was recreated 13:20:48Z by that restart. Mail still goes out (`alertmanager_notifications_total{integration="email"}` 5 → 6, `failed_total` 0, 14:23Z), but a silence set now and the record of what was already sent do not survive the next pod restart — so a restart can re-send every active alarm or drop a silence. Likely collateral of the restart (volume ownership on re-attach), not measured. **Checked from source 2026-10-05 (burn-down round 2):** homelab-manifests@87dfc29 mon-system/alertmanager.yaml:137-247: the Deployment has NO securityContext / fsGroup / runAsUser at all (grep), runs prom/alertmanager:v0.34.1 (:199, non-root `nobody` image) with --storage.path=/alertmanager on the Longhorn PVC alertmanager-data (:202, :212-213, :245-247). The comment :239-244 asserts silences now survive a restart -- an invariant with no test, which is exactly what this row says broke. Last structural change 58d1cd2 (2026-08-14, 'give alertmanager re | **OPEN** | — | Compare the volume's file owner with the pod's `securityContext` (`fsGroup`/`runAsUser`); fix in homelab-manifests; prove with a silence that survives a pod restart | operator | +| **R-884** | Monitoring & notifications | P4 | **ArgoCD app `monitoring` shows `Deployment/prometheus` OutOfSync** (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. **Checked from source 2026-10-05 (burn-down round 2):** Strong lead from source: homelab-manifests@87dfc29 commit 53c6e99 (Renovate, 2026-10-03) changed ONLY mon-system/monitoring.yaml `prom/prometheus:v3.14.0` -> `v3.15.0` (monitoring.yaml:419), and the `monitoring` Application has no `automated` syncPolicy in git (argocd-apps/homelab.yaml:602-605). So the drift is most likely an unsynced Renovate bump, i.e. a full sync = Prometheus 3.14 -> 3.15 upgrade (plus pod restart; R-211: no reloader). Not confirmed live. | **OPEN** | — | `argocd app diff monitoring` (or the CR's resource diff) to see what differs, then decide git or live | operator | ## Hub & operator — 13 rows (P2 1, P3 5, P4 7) @@ -323,7 +308,7 @@ stopping line that lies. | **R-31** | Hub & operator | P3 | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-only; a known workaround (click once, wait, verify) exists.** | — | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | CC | | **R-244** | Hub & operator | P3 | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | — | — | operator | | **R-882** | Hub & operator | P3 | **Longhorn on DooPlex could not grow a volume online: its `instance-manager` (116 days up) called a host process that no longer existed** — `nsenter: cannot open /host/proc/196610/ns/mnt` on every expansion retry, and an offline growth was blocked by the expansion's own attachment ticket (found 2026-10-05 growing `hub-data` to 2 Gi). A restart of the instance-manager (operator-approved) fixed it: 77/77 volumes back `attached/healthy` in 110 s. **Why the cached PID went stale was not established** (likely a containerd/k3s or iscsid restart after the instance-manager started), so it will recur after the next such restart and stay invisible until a volume needs to grow. `audits/hub-db-offsite-2026-10-05/partA/step1-*.txt` | **OPEN** | — | Find which host process the PID was and whether Longhorn 1.10.x re-resolves it; until then, before growing any volume, check the instance-manager's age against the last k3s/containerd restart | operator | -| **R-883** | Hub & operator | P3 | **8 DooPlex workloads run an image by a moving tag (`:latest` or none), so any pod restart is a silent upgrade.** Measured 2026-10-05: the Longhorn restart restarted zipline on `ghcr.io/diced/zipline:latest` (pull Always), which pulled 4.8.0; 4.8.0 refused its database (`cannot safely migrate from prisma to drizzle: expected migration 20260508022000 … was not applied`) and crash-looped. Fixed for zipline by pinning `4.7.0` (homelab-manifests `90f60e4`, `4c8ec7a`; 4.7.0 applied the four missing migrations; a dump from before is kept out-of-band). The other 7 were counted, not named or changed. | **OPEN** | — | List the 7 (`kubectl get deploy,sts -A` images without a fixed tag), pin each to the running version, and let Renovate move them | operator | +| **R-883** | Hub & operator | P3 | **8 DooPlex workloads run an image by a moving tag (`:latest` or none), so any pod restart is a silent upgrade.** Measured 2026-10-05: the Longhorn restart restarted zipline on `ghcr.io/diced/zipline:latest` (pull Always), which pulled 4.8.0; 4.8.0 refused its database (`cannot safely migrate from prisma to drizzle: expected migration 20260508022000 … was not applied`) and crash-looped. Fixed for zipline by pinning `4.7.0` (homelab-manifests `90f60e4`, `4c8ec7a`; 4.7.0 applied the four missing migrations; a dump from before is kept out-of-band). The other 7 were counted, not named or changed. **Checked from source 2026-10-05 (burn-down round 2):** homelab-manifests@87dfc29 still has moving tags (grep image lines without a numeric tag): admin-system/toolbox.yaml:12 nicolaka/netshoot:latest (a bare Pod); calibre-system/cwa.yaml:826 calibre-web-automated:dev; outline-system/outline.yaml:270 minio/minio:latest; tandoor-system/recipe-importer.yaml:26 gitea.dooplex.hu/admin/recipe-importer:latest (pull Always :27); adventurelog-system/adventurelog.yaml:100 and :256 adventurelog-backend/frontend:latest (pull Always :101,:257); jarrs-system/jarr- | **OPEN** | — | List the 7 (`kubectl get deploy,sts -A` images without a fixed tag), pin each to the running version, and let Renovate move them | operator | | **R-264** | Hub & operator | P4 | **Twenty-one facts the boxes report that the hub can now decode nowhere, each allowlisted with a reason rather than silently skipped — and for these the reason is "no consumer today, and one is arguably owed".** Split out of R-260 on 2026-08-08 so that closing the CLASS (gated) and fixing its sharpest instance (`operator_key_configured`) could not be mistaken for having decided what the hub should do with the rest. **The list, grouped by what a consumer would be for.** **(a) Guest-network health — `guest_net` and its seven children** (`checked_at`, `has_route`, `dhclient_alive`, `heal_succeeded`, `heals_last_hour`, `last_heal_at`, `damped`). The R-54 watchdog reports per-guest network state and self-heal counts every cycle and the hub — the component that emails the operator — models none of it. There is a live incident in this project's own record where a killed `dhclient` took a tunnel down for 1 h 15 m (`audits/INCIDENT-guest-dhclient-killed-2026-07-20.md`); a recurring-heal signal is exactly what would have surfaced it. **This is the strongest candidate of the twenty-one.** **(b) `selfupdate_pending` + `selfupdate_pending_version`** — an agent that has flipped its binary and never committed reports pending on every heartbeat so that "the operator sees WHY the version isn't advancing", and no operator can see it. **(c) `mgmt_plane.healed_recently`** — bounded: the hub DOES alarm on the `privsep_healed_at` timestamp beside it, so the recurring-clobber signal is not lost, only this flag. **(d) `restore_tests.mount_parity` + `mount_inventory`** — R-262's subject; the verdict is not lost (a mismatch fails before `Pass` is set) but the hub cannot tell a full-fidelity pass from a boot-only one. **(e) `pbs_dr.applied_at`.** **(f) Controller-side: `config_hash`, `reporting_disabled`, `stacks`, `storage.migrated_to`, `backup.last_db_dump`, `backup.last_integrity_check`** — the last two are backup-integrity timestamps, which is the "presence is not success" neighbourhood. **For each the question is the same and is NOT answered here: is it wanted? If the hub should act on it, model it and name what consults it. If it should not, the honest end is that the emitter stops sending it** — a fact emitted forever and consumed nowhere is a future false green waiting for someone to write a check against it. **Deliberately not decided in the G-1 session**, whose scope was the gate plus the operator-access instance; unilaterally removing emitters would also break the byte-identical cross-repo host-report golden and is a coordinated two-repo change. **⚠ DISPOSITIONS RECORDED 2026-08-13 — and the first thing to say is that they were NOT in this register before today.** The rulings were made on 2026-08-12; this row still read **READY — owner Viktor** and the gate's twenty entries still all said *"arguably owed"*, so a session told to *"re-read the dispositions from the register"* would have found none. They are written down now, which is the point of writing them down. **THE COUNT WAS ALSO WRONG:** this row says *twenty-one*; the gate's allowlist held **twenty**, measured. Twenty is the number the dispositions below account for, exactly. **(1) BUILD A READER — four groups, fourteen facts.** (a) guest-network health, (b) the staged-update-pending pair, (d) the restore-test depth pair, (f-part) the two backup-integrity timestamps. **(2) NO READER WANTED — five facts, now recorded as `not consumed, DELIBERATELY` with the ruling and its date** in `scripts/wire_contract_gate.py`, each with its own reason rather than a bare refusal: `mgmt_plane.healed_recently` (the hub already alarms on the timestamp beside it), `pbs_dr.applied_at` (`pbs_dr.state` is the verdict; the timestamp alone is the attempt-read-as-result trap), `config_hash` (the hub authors the config and knows its own generation), `stacks` (the app view is built from the purpose-built `app_telemetry` wire), `storage.migrated_to` (box-local bookkeeping with no hub-side intent to reconcile against). **The emitters are deliberately left alone** — removing one is a coordinated two-repo change and breaks the host-report golden; the honest end here is a recorded decision, not a deletion. The gate grew a THIRD entry kind to carry them, because an undecided fact and a decided one must not read alike. **(3) `reporting_disabled`, decided on its own merits: RECLASSIFIED `redundant`** — `health.status = "disabled"` travels in the same minimal report, is decoded into `reports.health_status`, and IS rendered. **The decision surfaced a real defect the flag would not have fixed: the staleness checker is age-only, so a deliberately-silent box still alarms → R-321.** **PROGRESS, 2026-08-13: the first reader is BUILT — guest-network health, R-319.** Its eight allowlist entries are **removed** (an allowlisted tag is skipped, so leaving them would have meant the new reader's fields were never checked); the gate's checked-tag count rose **182 → 190** and skipped fell **88 → 80**, which is the positive control that the wiring is real. **WHERE THE TWENTY NOW STAND: 8 read · 5 deliberately unread · 1 redundant · 6 still owed a reader** (`selfupdate_pending`, `selfupdate_pending_version`, `restore_tests.mount_parity`, `restore_tests.mount_inventory`, `backup.last_db_dump`, `backup.last_integrity_check`) — counts measured from the allowlist, not estimated. **Only ONE reader was built on purpose:** four at once is a design session pretending to be an implementation, and this one now tells us what the other three cost | **OPEN — 6 of 20 still owed a reader; dispositions recorded 2026-08-13, first reader shipped (R-319)** — owner Viktor | — | — | operator | | **R-402** | Hub & operator | P4 | **The off-site integrity verdict and its depth are published to the hub and NO hub surface reads either.** `offsite.last_integrity_ok` has been on the wire since controller v0.227.0 and `offsite.last_integrity_depth` since v0.228.0; both are allowlisted in `scripts/wire_contract_gate.py` **with their reason**, which is why the gate is green rather than silent. **The order is deliberate and is the opposite of the one that produced R-331:** publish the value first, build the display when someone decides what the screen should say. R-331 removed a hub Backup card that rendered `Integrity Unknown` for every customer forever from fields nothing wrote. **The depth is not decoration:** "checked, OK" means two different things at structure depth and at 100%, so a card showing the verdict without the depth shows the same words for a check that re-read every byte and one that only read the index. **WHAT HAPPENS IF NOBODY ACTS:** the operator can only answer "was this customer's off-site store verified, and how deeply?" by reading that box's own log. | **OPEN — SMALL, needs a HUB decision first** | — | Decide what the hub screen should say, then model both fields hub-side and delete the two allowlist entries together. `offsite.last_integrity_check` is already decodable and is not allowlisted. | Viktor decides, CC builds | | **R-451** | Hub & operator | P4 | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 **BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies.** The controller's payload carries no image (`internal/report/types.go` L98–103) and the hub's `Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts** — but **the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in `09` §6.2–6.3; the payload question is `09` §3b **Q7**. **-- RULED 2026-09-23 (`09` §3 decision 18):** the report carries, per compose service (database included), the installed reference, the catalog reference and the badge state. **Built later, when the fleet grows** — Q7's recommendation, confirmed. | **DEFERRED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the fleet sweep (report field, hub denormalisation, fleet page) is ruled and not built — deferred until the fleet grows (`09` §6.3)) — **RULED — build deferred until the fleet grows; owner: CC** | — | — | CC | @@ -343,16 +328,14 @@ stopping line that lies. | **R-89** | Business & legal | P4 | Retention as a per-customer **commercial** policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 43 rows (P3 4, P4 39) +## Process & tooling — 35 rows (P3 3, P4 32) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-565** | Process & tooling | P3 | **[P3-LOW] The English page test sees only ACCENTED Hungarian: an ASCII-only Hungarian word left in a template passes it on the English page.** FOUND 2026-09-17 by slice 1 release C (R-556, controller v0.250.0): after the extractor and the tests were green, a by-eye review of the English renders found six Hungarian fragments still in JavaScript strings — „, majd a(z)” and „FIGYELEM:” in the storage decommission dialog, „jelenlegi:” on the drive-init list, the uptime units „mp” and „p” and the count word „ db” on the debug page. All six were converted by hand; **no test failed on any of them**, because `TestI18nEnglishPages` looks for Hungarian letters and the extractor's ASCII word list (`i18n_extract.py` `ASCII_HU`) is used by neither test nor gate. Release B's review had found more of the same kind (Konfig, Megtartva, helyi, pl., Befejezve, automatikus, jelenleg:, kedd/szerda/szombat, szint). **Fix shape:** run the ASCII word list over the English renders in `TestI18nEnglishPages` (after the data mask), with a negative control on an English sentence and a decoy planting „mp” in an English value; extend the list with the words releases B and C found. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-578** | Process & tooling | P3 | **[P3-LOW] A helper that takes the settings lock must never be called from inside a settings callback — there is no gate, only one test in one package.** FOUND 2026-09-18 the hard way, by localisation slice 2 release C introducing exactly that: `UpdateOffboxStatus` holds the settings WRITE lock while it runs its callback, `boxLang()` reads the language through the READ lock, and `sync.RWMutex` is not reentrant — so the off-site run's final status write DEADLOCKED, **holding the settings lock**, which would wedge everything else on that box that touches `settings.json`. The only symptom was `go test ./internal/backup/` going from 8 minutes to a 25-minute timeout. Fixed by hoisting the language resolution; `TestNoteHelpersAreNotCalledUnderTheSettingsLock` (internal/backup) now names the file and line in a second. **What is still open:** that test covers `internal/backup` only, and it knows only the `note`/`noteErr`/`boxLang` helpers. Any other settings-reading helper, in any other package, can make the same mistake with nothing to catch it but a hang. **Fix shape:** promote it to a gate over every package, keyed on "a call to a method that reads settings, inside a literal passed to a `settings.Update*` function"; or give `Settings` a re-entrant read path and remove the class. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** | — | — | CC | | **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** **MECHANISM SEEN 2026-10-05 17:15–17:28Z, with Gitea's own log (round 2):** catalog run 1368 (`4828dc7`) — 17:15:14 the job is marked started; 17:15:16 `router: slow POST /api/actions/runner.v1.RunnerService/FetchTask for 10.42.0.42 (the runner), elapsed 3192ms`, then `UpdateRepoRunsNumbers … context canceled` and `GetActionWorkflow: EOF` — **the runner abandoned its fetch after Gitea had assigned the task**; the runner log has no line for task 1371; 17:28:39 `actions/clear_tasks.go:174 stopTasks() [W] Cannot transfer logs of task 1371` — Gitea's zombie-task stop. **The load at that minute:** an outside crawler (216.73.216.78) walking commit pages and `archive/*.tar.gz`, and THIS session's CI waiter, whose 15-page job listings took 13–31 s each. An API re-run passed in 7 s. **Done in-session:** the waiter now asks `GET …/actions/runs?head_sha=` once a minute (1 s). **Not done (DooPlex, the operator's):** the runner's fetch timeout and Gitea's exposure to the crawler. | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. | — | Operator: decide whether to raise the act-runner fetch timeout and/or rate-limit the public Gitea pages the crawler walks; meanwhile re-run a lost job via `POST /repos/admin//actions/runs//rerun`. Keep the 2026-10-12 check | operator | | **R-206** | Process & tooling | P4 | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC | -| **R-208** | Process & tooling | P4 | **Every Felhom Go build re-downloads its modules because `ARG VERSION` sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB** | **READY (S) — NEW 2026-08-05, ROOT CAUSE PROVEN** | — | **Measured, not inferred.** All **208** retained `go mod download` records carried **`Usage count: 1`** — not one was ever reused in ~a month of builds. The mechanism was isolated by four controlled builds: an unchanged tree rebuilt with the **same** `--build-arg VERSION` → `RUN go mod download` **CACHED**; the **same** tree with a **new** `VERSION` → **executed**. `COPY go.mod ./` stays CACHED either way, which is the tell: a `COPY`'s key is content-based, while a `RUN`'s key includes the stage **environment**, and `ARG VERSION`/`ARG GIT_COMMIT` are declared *before* the download in `felhom-controller/controller/Dockerfile`. Since every real build passes a fresh version, the layer is invalidated **every single time**. **`felhom.eu/hub/Dockerfile` has the identical defect** (`ARG VERSION`/`ARG BUILD_TIME` above `COPY go.mod go.sum*` → `RUN go mod download`) — and because both Dockerfiles produce byte-identical `buildx du` description strings, the 208 records are a COMBINED count and must not be attributed to one project. **Fix shape (one line each, not applied here):** move the `ARG VERSION`/`ARG GIT_COMMIT`/`ARG BUILD_TIME` declarations down to just above the final `go build`. **Worth more than the cap and the move combined** — the cap bounds the symptom, this removes the source. `build.sh`'s `rm -rf` + `cp -a` and its host-side `go mod tidy` were **ruled out by fingerprinting**: the tree is byte-identical across runs and `tidy` is a no-op | CC | | **R-209a** | Process & tooling | P4 | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC | | **R-230** | Process & tooling | P4 | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | — | — | operator | | **R-288** | Process & tooling | P4 | **The capability map is too long to be read, and that is why it stops being true.** `architecture/00-capability-map.md` is **134 642 bytes / 19 456 words across 99 table rows in only 159 lines** — because the rows ARE the length. Measured, longest first: the unaided-recovery-journey row is **3 024 words**, the offsite-password-recovery row **1 087**, the unattended-restore-proof row **971**, the app/guest-network-failure row **904**. That single longest row is a novella of nested corrections, each appended rather than resolved. Its own verification stamp reads **2026-07-16 against evidence corpus @ felhom.eu tip `4b18cc5`** (line 23) — three weeks stale, which is the measurable consequence: nobody re-reads a row they cannot finish. **This is the project's memory, so restructuring it is surgery and wants daylight** — filed, deliberately not attempted in the 2026-08-09 session. **What the shape should probably be:** one line of status per capability plus a dated evidence pointer, with the argument moved to the audit it came from | **READY (M) — NEW 2026-08-09** | — | Do not fold this into another session; it needs its own. **SECOND CONCRETE COST, 2026-08-10 — and it is a different failure mode from the first.** An *execution record* — 33 package deletions on the operator's rule — was undiscoverable for two days because it lives **inside the row about the Configuration page being slow**. Two sessions searched for it: one reported "no register row records a package prune", the other exhausted the Gitea logs, the activity feed and the schema before concluding it might be unestablishable. It was in `OPEN-ITEMS.md` the whole time. **The first cost (2026-08-09) was two records that looked contradictory and were not; this one is a record that could not be found at all.** Illegibility now has two measured costs and they are different in kind: prose rows make claims ambiguous, and rows-about-other-things make facts unfindable. **The rule this earns is in CONTEXT.md: a record that lives inside a row about something else has not been recorded** | Viktor | @@ -365,21 +348,15 @@ stopping line that lies. | **R-394** | Process & tooling | P4 | **`felhom-build-deploy/SKILL.md` is 179 lines, over the 150-line limit its own repo now enforces.** Found 2026-08-25 by `scripts/check_skills.py` on its first run — the over-length was discovered BY the new checker, on the day the limit was written down, which is the checker working as intended. **It is not edited and not trimmed here:** the task that introduced the limit explicitly scoped the four pre-existing skills out, and trimming a build-and-deploy skill without exercising its commands is how a wrong command ships to a live host. **It is a named single-entry exception in `GRANDFATHERED` in `scripts/check_skills.py`, printed as a WARN on every run**, so it cannot fade; a NEW skill over the limit is convicted normally, and growing the set requires editing that file in a commit with a row to name. **The rationale for the limit** — attention thins across the excess, so the lines that matter are not the ones that survive — is in `skills/felhom-doc-authoring/SKILL.md` §5. | **OPEN — LOW** | — | Trim `felhom-build-deploy/SKILL.md` under 150 lines in a session that can VERIFY the commands it keeps, then delete its `GRANDFATHERED` entry in the same commit. The likely trim is the per-artifact command blocks moving behind a pointer to the runbooks, keeping the gotchas inline — but that is a judgement for a session with a build to run, not a line-count exercise. | CC | | **R-421** | Process & tooling | P4 | **THE CLASS: an instrument that matches a LABEL rather than the fact it names — five instances, every one found by accident.** R-410 (a `mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was ABSENT), R-94 (a test comparing a constant to itself). **The gates are the machinery that enforces everything else in this project, and they were the one part nothing had ever checked.** The 2026-09-01 decoy sweep read all 29 scripts and fooled **16**. Ten were fixed the same day; four remain with rows (R-422..R-425); six could not be given a plausible decoy and are named. **The shapes, so the next one is cheap to recognise:** (1) name-for-fact — it matches a path or directory NAME while the fact lives inside the file; (2) substring-for-field — it matches a token anywhere in a body instead of in the field that carries it; (3) declaration-for-reachability — it checks a thing is declared, not that it RESOLVES; (4) constant-for-measurement — it compares a value against itself. **The single largest cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level), so every one was green and correct today and would have gone blind the moment anyone added a subdirectory. `decoy_coverage_gate.py` now refuses a new gate that ships without a decoy. | **OPEN — the class row; it stays open as the place the next instance is recorded** | — | — | CC | | **R-422** | Process & tooling | P4 | **`reuse_refs_check.py` only checks citations whose extension is one of `go py html css yml yaml sh`.** A cited `.md` path that does not exist is invisible — MEASURED 2026-09-01: `documentation/architecture/99-does-not-exist.md` added to `REUSE.md` passed, while the `.go` control was correctly convicted. REUSE.md and the CLAUDE.md files cite `.md` paths routinely, so this is the common case, not an exotic one. Fix: widen `PATH_RE`, then walk the false positives it produces across all four repos — that pass is the work, not the regex. The decoy is kept in `scripts/test_gate_decoys.py` asserting TODAY's behaviour, so the day this is fixed the test fails and is updated deliberately. | **OPEN** | — | — | CC | -| **R-425** | Process & tooling | P4 | **`offbox_rename_gate.py` scans a fixed three-entry `FILES` list.** MEASURED 2026-09-01: `NAS-mentés` in a new `backups_offbox_extra.html` passed. The scope was correct when written and silently narrows every time the feature grows a file. Fix: scan the offbox feature's files by pattern, or assert the FILES list against a discovered set so a new file fails until it is classified. | **OPEN** | — | — | CC | | **R-426** | Process & tooling | P4 | **The decoy-coverage exemption list — 20 registered gates that ship WITHOUT a decoy test, each named.** `scripts/decoy_coverage_gate.py`'s `EXEMPT` map is debt, and this row owns it so it lives in the register and not only in a Python literal. **Four kinds:** (a) genuinely covered in the 2026-09-01 sweep but not yet moved into a suite — `hub-copy`, `instructions`, `docker-v`, `image-pins`; (b) blocked by an open hole and therefore un-assertable as rejecting — `site` (R-423), `one-register` (R-424), `offbox-rename` (R-425); (c) shared scripts whose decoy lives in `felhom.eu` and is counted there — `reuse-refs`, `instructions`, `observations` in the controller and agent runners; (d) **no plausible decoy constructed yet** — `hostinstall`, `wire-contract`, `due-checks`, `published`, `image-resolvable`, `volume-persistence`. Group (d) is the honest unknown: six gates whose soundness is UNTESTED, not established. **The list is green today and shrinks; a NEW gate with no decoy fails immediately.** | **OPEN — 20 names; group (d) is six untested gates** | — | — | CC | -| **R-454** | Process & tooling | P4 | **[P3-LOW] Five `internal/web` test files have been `gofmt`-unclean for an unknown length of time, and nothing notices.** MEASURED 2026-09-02: `gofmt -l controller/internal/web/` reports `backups_split_test.go`, `claim_code_naming_test.go`, `disk_health_test.go`, `r400_debug_routes_test.go`, `recovery_test.go` — at the **baseline** commit `960d29b0612c`, i.e. not introduced by v0.233.0 (both files added that day are clean). **`go vet` does not check formatting and `controller_gates.py` has no formatting gate**, so the only thing that would ever surface this is someone running `gofmt -l` by hand, which is how it was found. **Not reformatted in the same session, deliberately** — the minimal-changes rule, and a five-file whitespace commit inside a feature release makes that release's diff unreadable. **Small, and the cost of NOT having the instrument is the row:** the count can only grow, and every future `gofmt -l` run produces noise that hides a real one. Fix is two lines: a `gofmt -l` gate in `controller_gates.py` plus one formatting commit, in that order (the gate first, so the commit is provably complete). Owner: **CC.** | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: code-formatting tooling only; no household meets it.** | — | — | CC | -| **R-457** | Process & tooling | P4 | **[P3-LOW] A test that hardcodes a date AND asserts an age derived from it is green on the day it is written and red the next morning — one instance PROVEN, six candidate files named.** MEASURED 2026-09-03: `TestGroupD_BadgeRendersOnBothSurfaces` (shipped the previous day in v0.233.0) pinned a fixture `catalog_since: "2026-07-18"` and asserted the rendered string `"Frissítés elérhető — 46 napja"`. **The pure badge tests inject a clock; the RENDER test does not and cannot** — it goes through the production templates, which call the funcmap entry `updateBadge`, which reads `time.Now()`. The suite was green on 2026-09-02 and **FAILED on 2026-09-03** with *"the behind badge is missing"* on both surfaces, because the true answer had become 47. **Fixed by DERIVING the fixture** — `catalog_since` is computed as *today minus 46 days*, so the test asserts the real number through the real clock and cannot rot. **THE CLASS, which is why this is a row and not just a fix:** a clock-reading test that also carries a date LITERAL is a bomb with a fuse of unknown length, and the suite being green is not evidence it is defused — it is evidence the fuse has not burned down yet. **NAMED AS UNCHECKED CANDIDATES, NOT ACCUSED** — six other test files contain both a `20xx-xx-xx` literal and `time.Now()`: `internal/backup/offbox_test.go`, `internal/web/handler_export_upload_test.go`, `internal/web/r103_tier2_action_test.go`, `internal/web/dashboard_backup_card_test.go`, `internal/web/async_restore_test.go`, `internal/stacks/installed_test.go`. Mixing the two is not itself a defect — it is one only where a literal feeds an assertion evaluated against the real clock — so each needs reading, which is a sweep and not this session. **The instrument that would end the class:** run the suite once under a faked future date in CI and see what turns red. Owner: **CC.** `felhom-controller` v0.234.0 CHANGELOG | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: test-suite hygiene; no household meets it.** | — | — | CC | -| **R-488** | Process & tooling | P4 | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: developer speed only.** | — | — | CC | +| **R-488** | Process & tooling | P4 | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. **Checked from source 2026-10-05 (burn-down round 2):** The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded `interval := 5 * time.Second` and `time.Sleep(3 * time.Second) // initial settling time`, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go:90, restore_unit.go:471. No commit since 2026-09-13 touching internal/backup mentions R-488/test speed. Runtime (333 s) was NOT re-measured here (read-only). | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: developer speed only.** | — | — | CC | | **R-492** | Process & tooling | P4 | **[P3-LOW] `cfg.Paths.HDDPath` is empty on every box and still has readers; delete it.** R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, `systemInfo`, the same fallback. The global now carries no information on any box and its deletion was deferred twice. **Fix shape:** remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: dead-code cleanup; every reader already falls back correctly.** | — | — | CC | | **R-502** | Process & tooling | P4 | **[P3-LOW] The bootstrap regression harness is run by NO gate and NO CI — and it had never exercised the pairing banner.** MEASURED 2026-09-14: `grep -rn bootstrap-modes scripts/*.py .gitea/workflows` → nothing; the harness is run by hand in `felhom-iso-assistant:trixie`. While adding the R-496 checks, its fake hub's `register` reply turned out to carry no `pairing_code`, so `print_pairing_banner` returned early in every run: the banner that a household reads first had no test at all until ISO v1.27.0. **Fix shape:** register the harness in `repo_gates.py` behind a container-available check that reports INCONCLUSIVE (never skip-as-pass) when docker is absent, with a decoy. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: test harness coverage; no household meets it directly.** | — | — | CC | | **R-507** | Process & tooling | P4 | **[P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen.** MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): `qm sendkey 332 tab` did not move focus (both password copies landed in one field), `mouse_move 1237 772` + `mouse_button 1` did not move the cursor or press Next, while `alt-n` did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. **Fix shape:** measure QEMU `input-send-event` with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: release-test tooling; an operator click-through covers it.** | — | — | CC | -| **R-564** | Process & tooling | P4 | **[P3-LOW] The retrieval-promise gate's Hungarian stems cannot see a SPLIT verb — „csak akkor állíthatók vissza", „hozod vissza" — so those Hungarian sentences were never scanned; the English translation exposed them.** FOUND 2026-09-17 by slice 1 release B (R-556): after the gate learnt English (`EN_PATTERNS`), seven English retrieval phrases on `backups_remote`, `backups_escrow`, `backups_restore` and `backups_restore_wizard` had NO Hungarian registration, because their Hungarian carries the verb particle after the verb („A távoli mentések csak akkor állíthatók vissza …", „a távoli mentések CSAK ezzel a kóddal állíthatók vissza", „csak a hiányzó fájlokat hozod vissza"). The stems (`visszaállíthat`, `visszaszerezhet`, `visszahozhat`, `visszanyit`) match only the joined form. The seven were registered in English with reasons (none is a false promise: two are preconditions, five describe the action on the same page). **Fix shape:** add split-form patterns to the Hungarian scan (`állíthatók? vissza`, `(hoz|szerez|nyit)\w* vissza`), register the Hungarian occurrences found, decoy with a planted split-verb promise. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: copy gate precision; the seven found sentences were reviewed and are true.** | — | — | CC | | **R-569** | Process & tooling | P4 | **[P3-LOW] Four more API handlers pick their status code by matching ENGLISH words in an error — the same shape as R-553, one language over.** FOUND 2026-09-17 while fixing R-553: `controller/internal/api/router.go` matches `"protected"`, `"not found"`, `"not deployed"`, `"still running"`, `"not orphaned"` in `err.Error()` at the stop/start, remove and orphan-cleanup handlers (three separate blocks). These strings are internal English, so localisation does not move them — the risk is a reworded internal error, not a translation, which is why this is P3 and was NOT folded into R-553's release. **Fix shape:** the same `util.KindErrorf` sentinels in `internal/stacks` (`ErrProtectedStack`, `ErrStackNotFound`, `ErrStillRunning`, …), a `statusFor` helper per handler family, and one table test per family passing a reworded message. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: internal robustness; only a future rewording would break it.** | — | — | CC | | **R-574** | Process & tooling | P4 | **[P3-LOW] `web/handler_debug.go` mixes page copy with JSON payload, so neither half could be converted safely.** FOUND 2026-09-18 by localisation slice 2 release A (R-557): the file holds 39 Hungarian literals and the inventory classifies them by STATEMENT, not by data flow (`I18N-INVENTORY-2026-09-17.md` §4), so which are section headings the debug page renders and which are values inside a diagnostic dump the operator copies out is not established. Converting a dump value would change what an operator pastes into a report; leaving a heading Hungarian leaves a half-English page. **Fix shape:** walk the file once and label every literal page-copy or payload in the same table slice 2 release A used, then convert only the page-copy half. Belongs to slice 2 release B or C. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: debug page is operator-facing.** | — | — | CC | | **R-576** | Process & tooling | P4 | **[P3-LOW] `i18n_go_parity.py` cannot see a call site that LOST text; it only checks that the text a key carries is real.** FOUND 2026-09-18 by localisation slice 2 release B, and found the hard way: the bulk converter silently dropped the continuation of a multi-line concatenation (`fmt.Errorf("a: "+ "b: %s", x)` kept only `"a: "`), damaging **7** producers — and the gate stayed GREEN throughout, because every surviving fragment WAS a byte-equal base-commit literal. Its question ("is this text real?") was answered yes while the CALL had lost half its sentence and its arguments. Two behaviour tests caught it (`TestR356_ScenarioC_UndeployedAppIsStillRefused`, `TestR379_ScenarioA_RollbackSucceeds_AppComesBack`), because they assert the sentence a customer READS. **Fix shape:** the gate learns a second question — for every `util.MsgError("key", …)` call site, the count of its arguments must equal the count of printf verbs in the key's Hungarian value, and no key-naming literal may be adjacent to a `+`. Both are cheap and would have convicted all 7. **The general lesson, worth keeping whatever is built: a structural gate over the TEXT cannot see a defect in the CALL.** | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: gate improvement; tooling only.** | — | — | CC | | **R-579** | Process & tooling | P4 | **[P3-LOW] Five page shells loaded `style.css` with NO cache-buster, so a browser holding an older copy kept being served CSS that did not know about the newest UI.** FOUND 2026-09-18 by the operator's screenshot of the v0.254.0 language globe: it rendered as a bare, unstyled `
` — a stray triangle and two plain words outside the card — because `login.html`, `claim.html`, `recovery.html`, `launcher_shared.html` and `launcher_share_password.html` requested `/static/style.css` with no `?v=`, while `layout.html` has used `?v={{.Version}}` since v0.166.0. **`.Version` was also absent from three of those five data maps.** FIXED AND CLOSED in the same session (controller v0.255.0): the parameter on all five, and `Version` set once in `executeTemplateLang` so a new shell cannot miss it; `TestGlobeOnAnonymousShells` now refuses an absent or EMPTY `?v=`. **The general form, which is the part worth keeping: a template that loads a versioned asset WITHOUT its version is invisible to every test that reads markup — the markup is correct and the browser fetches the wrong file.** A gate over "every stylesheet/script link in a template carries `?v=`" would catch the class; not built, because 8 first-boot-wizard templates would fail it and R-554 deletes them. | **DEFERRED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the gate "every stylesheet/script link carries `?v=`" was deliberately not built until R-554 deletes the old wizard; nothing else owns it) — **CLOSED 2026-09-18 - controller v0.255.0** | — | — | CC | -| **R-591** | Process & tooling | P4 | **[P3-LOW] `Stack.Copy()` is a deep copy with one shallow field, and the field is new.** FOUND 2026-09-20 while adding the catalog's English overlay: `controller/internal/stacks/manager.go` `Copy()` deep-copies `Meta.DeployFields` (with nested `Options`), `Meta.OptionalConfig` (with nested `Fields`), `Meta.Integrations`, `Meta.HealthCheck` and `Meta.InitialCreds` — and does NOT copy the new `Meta.I18n` map, which the struct assignment leaves shared between the original and the "copy". **It is safe TODAY and that is exactly the shape worth filing:** `Metadata.For` reads the overlay and never writes to it (pinned by `TestForDoesNotMutateTheReceiver`), so nothing can observe the sharing yet. The hole is in the CONTRACT — a function whose whole purpose is "a snapshot the caller may mutate" now has a field that is not one, and the next person to write through an overlay will find a bug with no failing test in front of it. **Fix shape:** deep-copy `I18n` in `Copy()` and pin it with a test that mutates the copy's overlay and asserts the original is unchanged. Alternatively state in `Copy()`'s comment that `I18n` is deliberately shared and immutable, and pin THAT with a test. Either is fine; silence is not. | **READY - rank P3-LOW; owner: CC (controller)** **Re-ranked 2026-10-03: P3->P4: latent contract gap, safe today.** | — | — | CC | -| **R-603** | Process & tooling | P4 | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `'`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.** **Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** **Re-ranked 2026-10-03: P3->P4: test tooling.** | — | — | CC | | **R-624** | Process & tooling | P4 | **[P3-LOW] Three of the catalog's apps cannot be seeded by ANY headless route, and for two of them that is a deliberate security decision — so the upgrade harness has a permanent ceiling nobody has written down.** FOUND 2026-09-21 while widening R-462 from 3 apps to . **`vaultwarden` and `zipline` close self-registration ON PURPOSE** — vaultwarden by `SIGNUPS_ALLOWED=false` (R-512, *„a stranger who guesses vault. must not be able to register"*), zipline by answering `E1037: User registration is disabled` — and neither ships a CLI that could make an account instead. So there is no route to a first account without the admin secret, and **that is correct**: the harness must not be the reason a customer-facing app accepts strangers. **`gitea` is a different and fixable case:** the template sets no `INSTALL_LOCK`, so a fresh instance sits in its web-installer state and `gitea admin user create` refuses (`MustInstalled() [F] Unable to load config file for a installed Gitea instance`); POSTing the installer form first would work and was simply not written tonight. **Why this is a row rather than three notes:** `09` §3 decision 6 says the upgrade test goes to **all** apps, and R-462 is costed as if every app is reachable given enough fixture work. It is not. There is a class — *apps whose only account-creating route the catalog deliberately closes* — for which the honest maximum is `inconclusive` unless the harness is given the app's admin secret at deploy time, which is a decision nobody has taken. **Needs:** the class named in R-462's scope so the remaining count is honest; a decision on whether the harness may hold an app's admin secret (it already holds the ones IT generates — see R-619); and, separately and cheaply, a gitea installer-form fixture. Evidence: `audits/update-night-2026-09-21/apps/{vaultwarden,zipline,gitea}/verdict.json`. **— NIGHT 2026-09-23:** `gitea` stayed inconclusive on both venues (its installer); `vaultwarden`/`zipline` not attempted (closed sign-up, by design); `code-server`, `outline`, `rallly` have no front-door seed route (listed, not moved); `bentopdf`, `glance`, `crafty-controller`, `wger`, `wanderer` (meilisearch) and `uptime-kuma` have no fixture tonight (listed, not moved). **-- 2026-09-30: the ceiling was WRONG for three of the six named apps.** outline has a front-door first-run route (`POST /api/installation.create` — its self-hosted setup, refused once a team exists), rallly signs up through its own better-auth API with the e-mail code read from its own table in place of a mailbox, and zipline 4's first-run route is `POST /api/setup` (the fixture had tried `/api/auth/register` and `/api/auth/setup`, which are not it). Fixtures for outline and rallly are in `upgrade_fixtures_box.py` and both apps moved on both venues; zipline's fixture now tries `/api/setup` first (measured on the bench: a SUPERADMIN made, the login works). **What remains in the class:** vaultwarden (closed sign-up by design, no CLI), code-server (a browser IDE). `audits/pg-last-six-2026-09-30/`. **-- 2026-09-30 (evening): gitea's fixable case is DONE** — the fixture posts gitea's own first-run installer form (the form's own defaults + an admin; no CSRF on that form), measured on the bench and through the setup gate on 9202; gitea 1.27.0 → 1.27.3 published `d7ba60c`. What remains in the class: vaultwarden, code-server. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (catalog harness)** **Re-ranked 2026-10-03: P3→P4: an upgrade-test coverage limit; no household meets it.** | — | — | CC | | **R-652** | Process & tooling | P4 | **[P3-LOW] The memory watch counted the kernel's file cache as the app's memory.** MEASURED 2026-09-23 night on the bench: nextcloud 34.0.4 read **100 %** of its 1 GiB and immich's PostgreSQL **100 %**, each with **0** kernel `oom_kill`s — `memory.peak` includes page cache, which the kernel drops before it kills anything. Under the watch as built (R-635 follow-up) both would be marked `memory_tight`, and the gate would demand a raised `mem_limit` — a customer-box capacity figure — for cache. **Done the same night (09 §3 decision 22, CC — operator may reverse):** the watch samples the app's own memory (`anon` of `memory.stat`) every 15 s; the mark and the ladder's `memory_peak_pct` read it where measured; the cgroup peak stays beside it (`memory_cgroup_peak_pct`). **Open:** romm's backfilled entry carries M1's 80.9 % cgroup peak (measured before the anon sample existed) — re-measure it on its next step; and decide whether an app whose anon is low but whose cgroup stays pinned at its limit (cache thrash) should be marked at all. Evidence: `audits/night-2026-09-23/apps/nextcloud/bench-1024M/`, `apps/immich/bench-noanon/`. | **READY — P3; owner: CC (catalog harness)** **Re-ranked 2026-10-03: P3→P4: the main fix is done; what is left is harness tuning, no household meets it.** | — | — | CC | | **R-693** | Process & tooling | P4 | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** **Re-ranked 2026-10-03: P3→P4: a harness judgement problem; no household meets it.** | — | — | CC | diff --git a/documentation/tests/golden-0.297.0-2026-10-05/02-round-trip.txt b/documentation/tests/golden-0.297.0-2026-10-05/02-round-trip.txt new file mode 100644 index 00000000..91bca118 --- /dev/null +++ b/documentation/tests/golden-0.297.0-2026-10-05/02-round-trip.txt @@ -0,0 +1,5 @@ +== round trip 2026-10-05T18:26:49Z: anonymous GET .../generic/felhom-golden/0.297.0/golden.tar.zst +HTTP 200 +bytes 648208028 +sha256 8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad +printed 8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad diff --git a/documentation/tests/golden-0.297.0-2026-10-05/README.md b/documentation/tests/golden-0.297.0-2026-10-05/README.md new file mode 100644 index 00000000..4ac343ce --- /dev/null +++ b/documentation/tests/golden-0.297.0-2026-10-05/README.md @@ -0,0 +1,52 @@ +# Golden 0.297.0 — bake + publish + vouch, 2026-10-05 (night, burn-down round 2) + +Procedure: `documentation/runbooks/RUNBOOK-manual-build.md` §4.0 and §4.1 steps 1–5, in the drill VM on DooPlex. + +| | Previous (`../golden-0.296.0-2026-10-05/`) | This bake | +|---|---|---| +| `build-golden.sh` | sha256 `645b3b659cba…` | same file, unchanged (agent repo `configs/build-golden.sh`) | +| Controller | `felhom-controller:0.296.0` | **`felhom-controller:0.297.0`** (MinAgent 0.131.0, unchanged) | +| Docker engine | the approved set `os-docker-20261004-142842` | same pinned set (same `GOLDEN_DOCKER_PKGS` as the 0.296.0 bake) | +| Guest packages | template | template — `GOLDEN_GUEST_PKGS` EMPTY | + +## Launch + +- No qemu running before; drill VM reverted to `virgin`, cold-booted per §4.0; `pveversion` = `pve-manager/9.2.2`. +- `pveam update` → `update successful`; `pveam available` listed `debian-13-standard_13.6-1_amd64.tar.zst` (downloaded). +- `/root/bake-run.sh` reads the token from the file; transient unit `golden-bake`. Token copied file → file (`scp`); + `systemctl show golden-bake -p Environment -p ExecStart | grep -c -F ` = **0**. + +## Pass markers (from `bake.log`, this folder) + +``` + docker OK (overlay2; data-root /var/lib/docker) +INFO: including mount point rootfs ('/') in backup +INFO: including mount point mp0 ('/var/lib/felhom') in backup +[golden] upload OK (HTTP 201) +GOLDEN_VERSION=0.297.0 +GOLDEN_SHA256=8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad +``` + +No `excluding` and no `FATAL` in the log (grep count 0). + +## Round trip — `02-round-trip.txt` + +Anonymous GET of `…/generic/felhom-golden/0.297.0/golden.tar.zst`: HTTP 200, 648 208 028 bytes, sha256 equals the printed one. + +## Secrets + +Saved-log leak grep for the literal token: **0**; positive control (a throwaway copy with the token appended): **1**, +copy shredded. + +## Vouch (step 5) + +`POST /configuration/artifacts` (Basic + `X-Felhom-Operator`, hub v0.137.0): agent **0.147.0**, golden **0.297.0**, +`min_agent` **0.131.0** → `303 flash=artifacts_set`; hub log `Artifact manifest set: agent=0.147.0 golden=0.297.0 +min_agent="0.131.0" … bundle_sha="326527d0…"`. Per-customer floors 0.297.0 (declared MinAgent 0.131.0) for demo-hp, +demo-felhom, tester-1 (`../../audits/burndown2-2026-10-05/delivery/vouch-golden-floors.txt`). The global floor unchanged. + +## Teardown + +`pct destroy 9100 --purge` (rc 0); `shred -u` of the token, runner script, bake script and log in the VM (log copied off +first); `poweroff`; qemu gone (`ps -eo comm | grep -c qemu-system-x86` = 0); `qemu-img snapshot -a virgin`. Host: +nothing provisioned. diff --git a/documentation/tests/golden-0.297.0-2026-10-05/bake.log b/documentation/tests/golden-0.297.0-2026-10-05/bake.log new file mode 100644 index 00000000..3fe4e169 --- /dev/null +++ b/documentation/tests/golden-0.297.0-2026-10-05/bake.log @@ -0,0 +1,339 @@ +[golden] build-golden.sh v3.2.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.297.0 +[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) … + Logical volume "vm-9100-disk-0" created. + Logical volume pve/vm-9100-disk-0 changed. +Creating filesystem with 8388608 4k blocks and 2097152 inodes +Filesystem UUID: 3f3e66a1-e05c-42f3-91af-3ff9fc5c8f65 +Superblock backups stored on blocks: + 32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208, + 4096000, 7962624 + Logical volume "vm-9100-disk-1" created. + Logical volume pve/vm-9100-disk-1 changed. +Creating filesystem with 6291456 4k blocks and 1572864 inodes +Filesystem UUID: 9327fbdd-d6c5-4299-a41f-3cb56b1238b3 +Superblock backups stored on blocks: + 32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208, +extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst' +Total bytes read: 553512960 (528MiB, 115MiB/s) +Detected container architecture: amd64 +Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ... +done: SHA256:6hCAi5WjsL3daO1FsrQR2bko1QAWr49AoJUmsNXkGSY root@felhom-golden +Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ... +done: SHA256:STGSy09tGfb2oiy9XwC4UHzFcSawljx6IhtUka8tiJI root@felhom-golden +Creating SSH host key 'ssh_host_rsa_key' - this may take some time ... +done: SHA256:rh5fSFDYSXH/CINtl163usIqQy+m082joU7xzA5LENU root@felhom-golden +[golden] starting + installing Docker (official repo, trixie channel) … +[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie docker-buildx-plugin=0.37.1-1~debian.13~trixie docker-ce=5:29.8.2-1~debian.13~trixie docker-ce-cli=5:29.8.2-1~debian.13~trixie docker-ce-rootless-extras=5:29.8.2-1~debian.13~trixie docker-compose-plugin=5.6.0-1~debian.13~trixie +apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct! +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = (unset), + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to the standard locale ("C"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_MESSAGES to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory +apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct! +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = (unset), + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to the standard locale ("C"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_MESSAGES to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory + installed: containerd.io 2.3.6-1~debian.13~trixie + installed: docker-buildx-plugin 0.37.1-1~debian.13~trixie + installed: docker-ce 5:29.8.2-1~debian.13~trixie + installed: docker-ce-cli 5:29.8.2-1~debian.13~trixie + installed: docker-ce-rootless-extras 5:29.8.2-1~debian.13~trixie + installed: docker-compose-plugin 5.6.0-1~debian.13~trixie +[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a +[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): 49 +[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation … +[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds … +[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) … +Unable to find image 'hello-world:latest' locally +latest: Pulling from library/hello-world +4f55086f7dd0: Pulling fs layer +4f55086f7dd0: Download complete +4f55086f7dd0: Pull complete +Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8 +Status: Downloaded newer image for hello-world:latest + docker OK (overlay2; data-root /var/lib/docker) + live-restore: on + /var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4 + /mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4 + both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576 +[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.297.0 (no registry cred at deploy) … + +WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'. +Configure a credential helper to remove this warning. See +https://docs.docker.com/go/credential-store/ + +0.297.0: Pulling from admin/felhom-controller +774043ccc8cc: Pulling fs layer +ab6b448d4be9: Pulling fs layer +23a5bfa58353: Pulling fs layer +862a57157567: Pulling fs layer +6db4169d1fd9: Pulling fs layer +167f80584563: Pulling fs layer +862a57157567: Waiting +6db4169d1fd9: Waiting +167f80584563: Waiting +774043ccc8cc: Verifying Checksum +774043ccc8cc: Download complete +862a57157567: Verifying Checksum +862a57157567: Download complete +23a5bfa58353: Verifying Checksum +23a5bfa58353: Download complete +167f80584563: Verifying Checksum +167f80584563: Download complete +6db4169d1fd9: Verifying Checksum +6db4169d1fd9: Download complete +ab6b448d4be9: Verifying Checksum +ab6b448d4be9: Download complete +774043ccc8cc: Pull complete +ab6b448d4be9: Pull complete +23a5bfa58353: Pull complete +862a57157567: Pull complete +6db4169d1fd9: Pull complete +167f80584563: Pull complete +Digest: sha256:23e4e0ffd9c9e2df28196e55bce8b89d4dfc1f9d4f8e52ca3ae2872e365073a0 +Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.297.0 +gitea.dooplex.hu/admin/felhom-controller:0.297.0 +[golden] asking the controller which infra images it manages … +[golden] baking infra images (4): traefik:v3.7.13 cloudflare/cloudflared:2026.9.3 gtstef/filebrowser:1.5.6-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 … +v3.7.13: Pulling from library/traefik +e2de96513ba9: Pulling fs layer +b686a4f73445: Pulling fs layer +78cb21c375ca: Pulling fs layer +acb2f33459b1: Pulling fs layer +acb2f33459b1: Waiting +e2de96513ba9: Verifying Checksum +e2de96513ba9: Download complete +b686a4f73445: Download complete +acb2f33459b1: Verifying Checksum +acb2f33459b1: Download complete +e2de96513ba9: Pull complete +78cb21c375ca: Verifying Checksum +78cb21c375ca: Download complete +b686a4f73445: Pull complete +78cb21c375ca: Pull complete +acb2f33459b1: Pull complete +Digest: sha256:24841fe2de7304c149343d877d2923b4c8800a38ba015dea9174c23b20e344a0 +Status: Downloaded newer image for traefik:v3.7.13 +docker.io/library/traefik:v3.7.13 +2026.9.3: Pulling from cloudflare/cloudflared +2cc7ee286bf3: Pulling fs layer +c172f21841df: Pulling fs layer +218cf840d0d9: Pulling fs layer +f6069939f718: Pulling fs layer +d6b1b89eccac: Pulling fs layer +2780920e5dbf: Pulling fs layer +7c12895b777b: Pulling fs layer +3214acf345c0: Pulling fs layer +52630fc75a18: Pulling fs layer +dd64bf2dd177: Pulling fs layer +b839dfae01f6: Pulling fs layer +ebddc55facdc: Pulling fs layer +c4bc6f35ff5e: Pulling fs layer +b96fe2995f90: Pulling fs layer +58c0c263dc73: Pulling fs layer +bd8962e29291: Pulling fs layer +cac2ae0193cb: Pulling fs layer +f0383d5ebc47: Pulling fs layer +3214acf345c0: Waiting +52630fc75a18: Waiting +dd64bf2dd177: Waiting +b839dfae01f6: Waiting +ebddc55facdc: Waiting +c4bc6f35ff5e: Waiting +b96fe2995f90: Waiting +58c0c263dc73: Waiting +bd8962e29291: Waiting +cac2ae0193cb: Waiting +f0383d5ebc47: Waiting +f6069939f718: Waiting +d6b1b89eccac: Waiting +2780920e5dbf: Waiting +7c12895b777b: Waiting +2cc7ee286bf3: Download complete +218cf840d0d9: Verifying Checksum +218cf840d0d9: Download complete +c172f21841df: Verifying Checksum +c172f21841df: Download complete +f6069939f718: Verifying Checksum +f6069939f718: Download complete +2cc7ee286bf3: Pull complete +d6b1b89eccac: Download complete +2780920e5dbf: Verifying Checksum +2780920e5dbf: Download complete +7c12895b777b: Verifying Checksum +7c12895b777b: Download complete +3214acf345c0: Verifying Checksum +3214acf345c0: Download complete +52630fc75a18: Verifying Checksum +52630fc75a18: Download complete +dd64bf2dd177: Download complete +c172f21841df: Pull complete +b839dfae01f6: Verifying Checksum +b839dfae01f6: Download complete +ebddc55facdc: Verifying Checksum +ebddc55facdc: Download complete +c4bc6f35ff5e: Verifying Checksum +c4bc6f35ff5e: Download complete +58c0c263dc73: Verifying Checksum +58c0c263dc73: Download complete +bd8962e29291: Verifying Checksum +bd8962e29291: Download complete +b96fe2995f90: Verifying Checksum +b96fe2995f90: Download complete +cac2ae0193cb: Verifying Checksum +cac2ae0193cb: Download complete +218cf840d0d9: Pull complete +f0383d5ebc47: Verifying Checksum +f0383d5ebc47: Download complete +f6069939f718: Pull complete +d6b1b89eccac: Pull complete +2780920e5dbf: Pull complete +7c12895b777b: Pull complete +3214acf345c0: Pull complete +52630fc75a18: Pull complete +dd64bf2dd177: Pull complete +b839dfae01f6: Pull complete +ebddc55facdc: Pull complete +c4bc6f35ff5e: Pull complete +b96fe2995f90: Pull complete +58c0c263dc73: Pull complete +bd8962e29291: Pull complete +cac2ae0193cb: Pull complete +f0383d5ebc47: Pull complete +Digest: sha256:072c067d25ccbe61d46e18f0d0723255f2bb5304f7317caa95b27031520ff92c +Status: Downloaded newer image for cloudflare/cloudflared:2026.9.3 +docker.io/cloudflare/cloudflared:2026.9.3 +1.5.6-stable: Pulling from gtstef/filebrowser +55afa1ecc21d: Pulling fs layer +8ed8f35f8d4f: Pulling fs layer +989b226a579c: Pulling fs layer +660aeead31d5: Pulling fs layer +4f4fb700ef54: Pulling fs layer +adce24567e4c: Pulling fs layer +f17ea56b313b: Pulling fs layer +6b6f3b3efe88: Pulling fs layer +4ed1ca4f3fce: Pulling fs layer +e6fc9c6a5757: Pulling fs layer +d47782d1182a: Pulling fs layer +660aeead31d5: Waiting +6b6f3b3efe88: Waiting +4ed1ca4f3fce: Waiting +e6fc9c6a5757: Waiting +d47782d1182a: Waiting +4f4fb700ef54: Waiting +adce24567e4c: Waiting +f17ea56b313b: Waiting +55afa1ecc21d: Verifying Checksum +55afa1ecc21d: Download complete +660aeead31d5: Verifying Checksum +660aeead31d5: Download complete +8ed8f35f8d4f: Verifying Checksum +8ed8f35f8d4f: Download complete +4f4fb700ef54: Verifying Checksum +4f4fb700ef54: Download complete +55afa1ecc21d: Pull complete +989b226a579c: Verifying Checksum +989b226a579c: Download complete +f17ea56b313b: Verifying Checksum +f17ea56b313b: Download complete +6b6f3b3efe88: Verifying Checksum +6b6f3b3efe88: Download complete +4ed1ca4f3fce: Verifying Checksum +4ed1ca4f3fce: Download complete +adce24567e4c: Verifying Checksum +adce24567e4c: Download complete +d47782d1182a: Verifying Checksum +d47782d1182a: Download complete +8ed8f35f8d4f: Pull complete +e6fc9c6a5757: Verifying Checksum +e6fc9c6a5757: Download complete +989b226a579c: Pull complete +660aeead31d5: Pull complete +4f4fb700ef54: Pull complete +adce24567e4c: Pull complete +f17ea56b313b: Pull complete +6b6f3b3efe88: Pull complete +4ed1ca4f3fce: Pull complete +e6fc9c6a5757: Pull complete +d47782d1182a: Pull complete +Digest: sha256:7c5d7ac8ffda31294d278063cf9d2e04303b39e6dce1f4c691342240ca7703b8 +Status: Downloaded newer image for gtstef/filebrowser:1.5.6-stable +docker.io/gtstef/filebrowser:1.5.6-stable +1.1.0: Pulling from admin/felhom-samba +897d797d2723: Pulling fs layer +3051591aa250: Pulling fs layer +ce57a3f93416: Pulling fs layer +fb94eeec2fe1: Pulling fs layer +fb94eeec2fe1: Waiting +ce57a3f93416: Verifying Checksum +ce57a3f93416: Download complete +fb94eeec2fe1: Verifying Checksum +fb94eeec2fe1: Download complete +897d797d2723: Verifying Checksum +897d797d2723: Download complete +897d797d2723: Pull complete +3051591aa250: Verifying Checksum +3051591aa250: Download complete +3051591aa250: Pull complete +ce57a3f93416: Pull complete +fb94eeec2fe1: Pull complete +Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10 +Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0 +gitea.dooplex.hu/admin/felhom-samba:1.1.0 +[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'. +[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'. +[golden] baking the first-boot SSH host-key regeneration unit (F3) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'. +[golden] identity-clean + minimize … +[golden] stop + archive … +INFO: including mount point rootfs ('/') in backup +INFO: including mount point mp0 ('/var/lib/felhom') in backup +INFO: archive file size: 618MB +INFO: Finished Backup of VM 9100 (00:00:29) +[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_10_05-20_24_09.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive) +[golden] publishing golden (648208028 bytes, sha256 8cebc42e15b091f0…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.297.0/golden.tar.zst +[golden] pre-delete existing: HTTP 404 (404/204 expected) +[golden] upload OK (HTTP 201) +GOLDEN_VERSION=0.297.0 +GOLDEN_SHA256=8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad +[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.297.0 / 8cebc42e15b091f0c2f8bee1170ec66b97de1a6bf14b7aa20dca82443c7e79ad +[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge)