diff --git a/REUSE.md b/REUSE.md index 482d9c6c..5b5d0d45 100644 --- a/REUSE.md +++ b/REUSE.md @@ -148,6 +148,7 @@ | `store.GuestID` | hub/internal/store/store.go (~L1268) | `(hostID string, vmid int) string` | Canonical guest primary key | Never hand-concatenate host+vmid. | | `(*Store).GetHostReportsSince` + `GetFirstHostReportAt` + `monitor.newestBackupEvidence` | hub/internal/store/store.go, hub/internal/monitor/deadline.go | `(customerID, since) ([]HostReportRow, error)`; `(customerID) (time.Time, error)`; `(rows, now) (time.Time, bool)` | **Asking "when did the hub last SEE evidence of X?" instead of "what does the latest report say?"** — the R-81 anchor. The agent's reporters are point-in-time and forget across a restart; the hub retains ~90 d of host-reports and does not. | The three go together: window scan + first-contact anchor + a bounded lookback (`backupEvidenceLookback`). **Never judge a report-derived absence on the LATEST report alone** — that is the bug class R-81 fixed for the third time. The scan early-exits on sufficiently-fresh evidence, so don't reorder rows away from newest-first. | | `scheduleDaily` | hub/cmd/hub/main.go (~L449) | `(ctx, name, "HH:MM", fn, logger)` | Daily jobs in Europe/Budapest (prune etc.) | Blocking — run as goroutine. `parseHM` returns 0,0 (midnight) on bad input. | +| `hu_grep.py` | scripts/hu_grep.py | `PATTERN --anchor ASCII [--negative TEXT] PATH…` | ANY search for Hungarian (accented) text in files — prints file:line hits, and refuses to report a zero unless an ASCII anchor hits and a negative control misses (R-364) | Reads bytes in Python, never via a shell; refuses a pattern that arrived transformed (octal escapes, U+FFFD). Exit 0 found · 1 tested zero · 2 refused | ## 2. Canonical patterns (copy structure from THE named file) diff --git a/documentation/architecture/02-controller-module-map.md b/documentation/architecture/02-controller-module-map.md index d093b927..aa5079de 100644 --- a/documentation/architecture/02-controller-module-map.md +++ b/documentation/architecture/02-controller-module-map.md @@ -370,6 +370,23 @@ own; every caller that is not the customer must decide for itself whether the ap | `storage_handlers.go` (1600 L) | **DELETE (→agent)** | Format/attach/mount/disconnect/migrate-drive/decommission disk UI. Any survivor is a **thin client calling the agent API** (e.g. per-volume placement requests). | hazard | | `templates/` (HTML, non-Go) | **PORT** | Remove disk-wizard + DR pages; keep app/deploy/backup/settings pages. | needs-rework | +#### Alert placement — inline under the storage bars, or the top banner (R-571) + +**[FACT, read from source 2026-10-05, felhom-controller `114ff27`, `controller/internal/web/alerts.go`]** Every +dashboard alert is an `Alert` with two placement fields: `PageOnly` (the pages it may appear on; empty = every +page) and `Inline` (rendered by the page template in place, not by the layout's banner). `GetAlerts` returns +the list (endpoint-drift, agent-channel and dead-app alerts first, then the rest sorted error > warning > info, +capped at five plus an overflow line), and `layout.html` paints only those that are not `Inline` and match the +page; `GetInlineAlerts(page)` hands the dashboard and monitoring pages their inline ones. + +**Exactly one warning is inline today:** the „storage is not on a separate drive" health warning. It is +`PageOnly: dashboard, monitoring` and `Inline: true`, so it sits quietly under the storage bars; **every other +warning, including every off-site failure (`07` §6.7), renders in the top banner on every page.** The choice +is made by the warning's KIND (`monitor.WarnKindStorageNotSeparate`, read with `report.WarningKindAt`), never +by its words — the earlier Hungarian-substring test would have moved the warning to the red banner on every +page the day the sentence was translated (R-553, pinned by `TestR553_DiskWarningPlacementSurvivesWordingChange`). +A new inline warning needs its own kind, not a text match. + ### `scripts/` | File | Class | Reason | Risk | |---|---|---|---| diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 552d9a3f..15c4599e 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -249,7 +249,8 @@ the identity bundle's shape is `{tunnel_token, pbs_token, wg_private_key, restic parts a host-loss recovery would read (INV Part D2.3): `hosts.dr_record_json` is `{}` on all three hosts; `host_escrow.directive_json` is `{}` on both escrowed hosts; `dr_recipe.host_half.drives` is `[]` on every customer including two with enrolled data drives; and `dr_recipe.host_half.pbs.namespace` -reads `"root"` while the real namespaces are `demo-felhom` / `demo-hp`. → **R-105**, **R-106**. +reads `"root"` while the real namespaces are `demo-felhom` / `demo-hp`. → **R-105**, **R-106**. *Since agent v0.147.0 (R-124) a genuine root namespace is +recorded as `""` (PBS's own spelling) beside `namespace_state: resolved`; the word `root` is no longer written.* --- @@ -795,6 +796,32 @@ digests all resolve today (`audits/version-travel-2026-09-26/A7/`). Options are it; older images of the app are deleted. A restore that needs an older version re-pulls it — as before; the limit above is unchanged, and kept data (decision 40) is not touched by the image clean-up. +### 6.7 Why an off-site run failed — the failure classifier (R-571) + +**[FACT, read from source 2026-10-05, felhom-controller `114ff27`]** When an off-site (restic) run fails, the +controller names ONE cause before it writes the note the customer's page shows days later +(`ClassifyOffsiteFailure` in `controller/internal/backup/offbox.go`; the head line is the bundle key +`note.offsite.fail_`, followed by the run time and the sanitised error). The classes, in the order +they are tested: + +| Class | Decided by | What it means for the household | +|---|---|---| +| `orphaned` | our sentinel `ErrOffboxOrphaned` | the remote store was made with a key this box no longer has; nothing new reaches it until the operator acts | +| `quota` | our sentinel `ErrOffsiteQuota` (the pre-run soft-quota gate, R-553) | the backup did not fit the remote space; the run was refused before upload | +| `locked` | our sentinel `ErrOffsiteLocked`, or restic's lock text (R-104) | an interrupted earlier run left the store locked and both self-heal layers failed | +| `no_units` | text: „produced no snapshots" | there was nothing to send — no chosen app had a backup on any drive | +| `no_repo` | restic's text: „unable to open config file" / „is there a repository…" | nothing exists at the remote location | +| `transport` | ssh/restic/rclone text: connection refused/reset, timeout, permission denied, host key, handshake, DNS, unreachable | the remote store could not be reached (network or sign-in) | +| `unknown` | everything else | the cause is not known — the page says so instead of guessing | + +**Two kinds of signal, and the difference matters.** The first three are OUR sentinels: they survive a +translation of our own text, which is why R-553 replaced the Hungarian-word match for `quota`. The text +signatures are **restic's, ssh's and rclone's own English output** — external strings we neither write nor +translate. A new restic or OpenSSH version that rewords an error moves that failure to `unknown`; it never +moves it to a wrong class. The order is deliberate: a cause that cannot be told apart returns `unknown` +rather than being folded into a neighbour. Where the warning is SHOWN on the dashboard is a separate rule — +`02-controller-module-map.md`, „Alert placement". + ## 7. The recovery chain (D3) — the reason this document exists **[DESIGN] 3-2-1 describes copies. It does not describe recovery.** diff --git a/documentation/audits/burndown2-2026-10-05/agent-red-proofs.txt b/documentation/audits/burndown2-2026-10-05/agent-red-proofs.txt new file mode 100644 index 00000000..79486b64 --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/agent-red-proofs.txt @@ -0,0 +1,58 @@ +===== R-118 RED (guard removed: if s.devicePresent(d.MountPath) -> if true) — 2026-10-05T18:53:22+02:00 +$ go test ./internal/localapi/ -run TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity -v -count=1 +=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity +=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/absent + disks_device_presence_test.go:212: absent drive advertises total=33697107968 used=4727169024 frac=0.140 — that is the filesystem UNDER the bare mountpoint, not the drive (R-118) +=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/present +--- FAIL: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity (0.00s) + --- FAIL: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/absent (0.00s) + --- PASS: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/present (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-agent/internal/localapi 0.010s +FAIL +rc=1 +===== R-118 GREEN (fix restored) +$ go test ./internal/localapi/ -run TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity -v -count=1 +=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity +=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/absent +=== RUN TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/present +--- PASS: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity (0.00s) + --- PASS: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/absent (0.00s) + --- PASS: TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity/present (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-agent/internal/localapi 0.009s +rc=0 + +===== R-269 RED (tokenstore.go reverted to HEAD: reload-on-MISS-only Lookup) — 2026-10-05T18:53:55+02:00 +$ go test ./internal/localapi/ -run TestTokenStore_RotatedOutTokenRejectedFirst -v -count=1 +=== RUN TestTokenStore_RotatedOutTokenRejectedFirst + tokenstore_test.go:247: rotated-out token still authorizes vmid 130 on its first presentation after rotation — Mint's 'any previous token for this guest is revoked' is false across processes (R-269) +--- FAIL: TestTokenStore_RotatedOutTokenRejectedFirst (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-agent/internal/localapi 0.008s +FAIL +rc=1 +===== R-269 GREEN (fix restored) +$ go test ./internal/localapi/ -run TestTokenStore_RotatedOutTokenRejectedFirst -v -count=1 +=== RUN TestTokenStore_RotatedOutTokenRejectedFirst +--- PASS: TestTokenStore_RotatedOutTokenRejectedFirst (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-agent/internal/localapi 0.010s +rc=0 + +===== R-317 RED (probe reverted to the pre-fix dnsmasq-base binary path usr/sbin/dnsmasq) — 2026-10-05T18:54:37+02:00 +$ go test ./internal/lanresolver/ -run TestEnsureDnsmasq_BinaryWithoutUnitInstalls -v -count=1 +=== RUN TestEnsureDnsmasq_BinaryWithoutUnitInstalls + ensure_dnsmasq_test.go:80: install was skipped on a dnsmasq-base-only host (binary present, unit absent) — the enable that follows targets a missing unit (R-317). calls: ["/usr/local/sbin/felhom-priv-apply dnsmasq /tmp/felhom-resolver-3076255408.conf felhom-resolver-base.conf" "systemctl enable --now dnsmasq" "systemctl restart dnsmasq"] +--- FAIL: TestEnsureDnsmasq_BinaryWithoutUnitInstalls (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-agent/internal/lanresolver 0.006s +FAIL +rc=1 +===== R-317 GREEN (fix restored) +$ go test ./internal/lanresolver/ -run TestEnsureDnsmasq_BinaryWithoutUnitInstalls -v -count=1 +=== RUN TestEnsureDnsmasq_BinaryWithoutUnitInstalls +--- PASS: TestEnsureDnsmasq_BinaryWithoutUnitInstalls (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-agent/internal/lanresolver 0.004s +rc=0 diff --git a/documentation/audits/burndown2-2026-10-05/catalog-red-proofs.txt b/documentation/audits/burndown2-2026-10-05/catalog-red-proofs.txt new file mode 100644 index 00000000..82a91f82 --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/catalog-red-proofs.txt @@ -0,0 +1,173 @@ +=== R-760 red-proof (2026-10-05T18:53:44+02:00) — catalog 29ac711 + working tree +--- UNDO: remove the R-760 '# No healthcheck' comment block from templates/vikunja/docker-compose.yml +$ python3 scripts/test_healthcheck_explained.py ++ [] : service(s) with no compose healthcheck and no '# No healthcheck' comment saying why: ['vikunja/vikunja'] + +---------------------------------------------------------------------- +Ran 2 tests in 0.011s + +FAILED (failures=1) +rc=1 +--- RESTORE +$ python3 scripts/test_healthcheck_explained.py +Ran 2 tests in 0.011s + +OK +rc=0 + +=== R-593 red-proof (2026-10-05T18:55:00+02:00) — catalog 29ac711 + working tree +--- UNDO: git show HEAD:templates/papra/.felhom.yml > templates/papra/.felhom.yml (the pre-fix papra .felhom.yml) +$ python3 scripts/test_deploy_field_descriptions.py +- ['papra SUBDOMAIN has no description', +- 'papra AUTH_SECRET carries the subdomain sentence', +- 'papra SUBDOMAIN has no description'] : ['papra SUBDOMAIN has no description', 'papra AUTH_SECRET carries the subdomain sentence', 'papra SUBDOMAIN has no description'] + +---------------------------------------------------------------------- +Ran 2 tests in 0.025s + +FAILED (failures=1) +rc=1 +--- RESTORE +$ python3 scripts/test_deploy_field_descriptions.py +Ran 2 tests in 0.025s + +OK +rc=0 + +=== R-781 red-proof (2026-10-05T18:55:52+02:00) — catalog 29ac711 + working tree +--- UNDO: drop the isolate_onboarding_clone(cat) call in onboarding_cases (the pre-fix harness) +1 +$ python3 + ok FACT: a new template with NO record rc=1 (expected 1) + ok FACT: a record missing id 1.4 rc=1 (expected 1) + ok FACT: 1.4 answered only inside an HTML comment rc=1 (expected 1) + ok FACT: done with a path that does not exist rc=1 (expected 1) + ok FACT: done with an EMPTY directory (the mkdir shape) rc=1 (expected 1) + ok FACT: done naming an absent file in the sibling repo rc=1 (expected 1) + ok FACT: n/a with an EMPTY reason rc=1 (expected 1) + ok FACT: n/a with a two-word reason rc=1 (expected 1) + ok FACT: an OPEN row rc=1 (expected 1) + ok FACT: opened: backdated before the checklist rc=1 (expected 1) + ok FACT: a checklist id the template a new app copies lacks rc=1 (expected 1) + ok FACT: an exempt app's record with a done that points nowhere rc=1 (expected 1) +FAILS 4 +FAIL: GENUINE: a complete record (catalog + sibling evidence): rc=1 expected 0; missing ['onboarding gate OK'] +FAIL: GENUINE: an id added AFTER opened: does not bind: rc=1 expected 0; missing ['onboarding gate OK'] +FAIL: GENUINE: an exempt app's record may say open: rc=1 expected 0; missing ['exempt app(s) with a record (shape-checked): wger'] +FAIL: STATED SKIP: sibling repo absent (the CI shape) - printed, not checked: rc=0 expected 0; missing ['felhom.eu/documentation/audits/onb/ +rc=1 +--- RESTORE + ok GENUINE: a complete record (catalog + sibling evidence) rc=0 (expected 0) + ok GENUINE: an id added AFTER opened: does not bind rc=0 (expected 0) + ok GENUINE: an exempt app's record may say open rc=0 (expected 0) + ok STATED SKIP: sibling repo absent (the CI shape) - printed, not checked rc=0 (expected 0) +FAILS 0 +rc=0 + +=== R-806 red-proof (2026-10-05T18:56:53+02:00) — catalog 29ac711 + working tree +--- UNDO: exercise_argv back to the pre-fix shape (always http://, no -k) — the old inline argv +578: pass # red-proof: no -k +581: return a + [f"http://{ip}:{port}{path}"] +$ python3 scripts/test_check_volume_persistence.py +FAIL: test_https_backend_gets_an_https_url_with_k (__main__.TestRoutedSchemes.test_https_backend_gets_an_https_url_with_k) +AssertionError: 'http://10.0.0.5:8443/' != 'https://10.0.0.5:8443/' +Ran 54 tests in 0.018s +FAILED (failures=1) +rc=1 +--- RESTORE +Ran 54 tests in 0.018s + +OK +rc=0 + +=== R-605 red-proof (2026-10-05T18:58:47+02:00) — catalog 29ac711 + working tree +--- UNDO: HARNESS_REFUSED = 2 in both gates (the pre-fix shared code); runner: rc 3 folded back into UNDETERMINED +scripts/check-image-resolvable.py:166:HARNESS_REFUSED = 2 +scripts/check-volume-persistence.py:935:HARNESS_REFUSED = 2 +124:VERDICT = {0: "OK", 1: "FAILED", 2: "INCONCLUSIVE"} +220: refused = [] # red-proof +$ python3 scripts/test_check_volume_persistence.py +FAIL: test_canary_failure_prints_the_harness_refused_marker (__main__.TestCheckEntryPoint.test_canary_failure_prints_the_harness_refused_marker) +AssertionError: 2 != 3 +Ran 55 tests in 0.017s +FAILED (failures=1) +rc=1 +$ python3 scripts/test_check_image_resolvable.py +Ran 19 tests in 0.011s +OK +rc=0 +$ python3 -m unittest scripts/test_catalog_gates.py SummaryTellsRefusedFromUndecided +FAIL: test_a_refused_harness_does_not_read_as_undetermined (test_catalog_gates.SummaryTellsRefusedFromUndecided.test_a_refused_harness_does_not_read_as_undetermined) +AssertionError: 'DID-NOT-RUN' not found in '\n==============================================================================\n== summary\n==============================================================================\n image-pins OK (exit 0)\n volume-persistence ERROR (exit 3)\nUNDETERMINED (it ran; some results could not be decided — never a pass): volume-persistence\n' +Ran 3 tests in 0.004s +FAILED (failures=1) +rc=1 +--- RESTORE +$ python3 scripts/test_check_volume_persistence.py +Ran 55 tests in 0.018s +OK +rc=0 +$ python3 scripts/test_check_image_resolvable.py +Ran 19 tests in 0.011s +OK +rc=0 +Ran 3 tests in 0.003s +OK +rc=0 + +=== R-605 red-proof, image-resolvable leg re-run (2026-10-05T18:58:59+02:00) — the first run above passed because the test compared against the constant itself (tautology); the tests now pin the literal 3 +--- UNDO: HARNESS_REFUSED = 2 in check-image-resolvable.py +$ python3 scripts/test_check_image_resolvable.py +FAIL: test_empty_catalog_is_an_error_not_a_pass (__main__.TestResolvabilityGate.test_empty_catalog_is_an_error_not_a_pass) +AssertionError: 2 != 3 +FAIL: test_untrustworthy_resolver_refuses_to_report (__main__.TestResolvabilityGate.test_untrustworthy_resolver_refuses_to_report) +AssertionError: 2 != 3 : a resolver that resolves the canary must abort (3, the harness refused), not pass — and not 2, which reads as 'some pins were throttled' (R-605) +Ran 19 tests in 0.013s +FAILED (failures=2) +rc=1 +--- RESTORE +Ran 19 tests in 0.012s +FAILED (failures=2) +rc=1 + +--- NOTE: the RESTORE above still failed because scripts/__pycache__ held the UNDONE module's bytecode: the undo + (3 -> 2) kept the file's size and the restore landed in the same second, so Python's mtime+size cache check + accepted the stale .pyc. Restored source confirmed HARNESS_REFUSED = 3; after `touch scripts/check-image-resolvable.py`: +$ python3 scripts/test_check_image_resolvable.py +Ran 19 tests in 0.011s +OK +rc=0 + (lesson for same-size red-proofs: run with PYTHONDONTWRITEBYTECODE=1 / clear __pycache__ between undo and restore) + +=== R-594 red-proof (2026-10-05T19:01:53+02:00) — catalog 29ac711 + working tree (PYTHONDONTWRITEBYTECODE=1) +--- UNDO 1: the register is read but never consulted (reg = None) — the pre-fix gate's behaviour +1 +$ python3 scripts/test_gate_decoys.py + ok FACT: registered for ANOTHER app - the promise still convicts, the entry is stale rc=1 (expected 1) + ok FACT: registered on the path, but the sentence was rewritten (match gone) rc=1 (expected 1) + ok FACT: a STALE entry - nothing in the English promises it any more rc=1 (expected 1) + ok FACT: a registered promise with a two-word reason rc=1 (expected 1) + ok FACT: n/a with a two-word reason rc=1 (expected 1) +FAIL: GENUINE: a REGISTERED true retrieval promise passes: rc=1 expected 0; missing ['copy-i18n: OK', '1 registered retrieval promise(s) in ALLOWLIST_EN, 1 used +rc=1 +--- UNDO 2: the STALE check removed (entries never judged live) +1 +$ python3 scripts/test_gate_decoys.py + ok GENUINE: a REGISTERED true retrieval promise passes rc=0 (expected 0) + ok FACT: a registered promise with a two-word reason rc=1 (expected 1) + ok FACT: n/a with a two-word reason rc=1 (expected 1) +FAIL: FACT: registered for ANOTHER app - the promise still convicts, the entry is stale: rc=1 expected 1; missing ['STALE entry vaultwarden'] +FAIL: FACT: registered on the path, but the sentence was rewritten (match gone): rc=1 expected 1; missing ['STALE entry privatebin'] +FAIL: FACT: a STALE entry - nothing in the English promises it any more: rc=0 expected 1; missing ['STALE entry privatebin'] +rc=1 +--- RESTORE +$ python3 scripts/test_gate_decoys.py + ok GENUINE: a REGISTERED true retrieval promise passes rc=0 (expected 0) + ok FACT: registered for ANOTHER app - the promise still convicts, the entry is stale rc=1 (expected 1) + ok FACT: registered on the path, but the sentence was rewritten (match gone) rc=1 (expected 1) + ok FACT: a STALE entry - nothing in the English promises it any more rc=1 (expected 1) + ok FACT: a registered promise with a two-word reason rc=1 (expected 1) + ok FACT: n/a with a two-word reason rc=1 (expected 1) +catalog gate decoys OK — 137 case(s), every label judged on its fact (R-421) +rc=0 + diff --git a/documentation/audits/burndown2-2026-10-05/controller-red-proofs.txt b/documentation/audits/burndown2-2026-10-05/controller-red-proofs.txt new file mode 100644 index 00000000..9e899426 --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/controller-red-proofs.txt @@ -0,0 +1,480 @@ + +==================== R-591 — red-proof ==================== +Fix undone in: internal/stacks/manager.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestDeepCopyStackI18nIsNotShared -v ./internal/stacks) # FIX UNDONE +rc=1 +=== RUN TestDeepCopyStackI18nIsNotShared + r591_copy_i18n_test.go:55: R-591: mutating the copy's data path overlay changed the ORIGINAL — the overlay is shared, not copied + r591_copy_i18n_test.go:55: R-591: mutating the copy's initial creds overlay changed the ORIGINAL — the overlay is shared, not copied + r591_copy_i18n_test.go:55: R-591: mutating the copy's description overlay changed the ORIGINAL — the overlay is shared, not copied + r591_copy_i18n_test.go:55: R-591: mutating the copy's tagline overlay changed the ORIGINAL — the overlay is shared, not copied + r591_copy_i18n_test.go:55: R-591: mutating the copy's deploy label overlay changed the ORIGINAL — the overlay is shared, not copied + r591_copy_i18n_test.go:55: R-591: mutating the copy's option label overlay changed the ORIGINAL — the overlay is shared, not copied + r591_copy_i18n_test.go:55: R-591: mutating the copy's use case overlay changed the ORIGINAL — the overlay is shared, not copied + r591_copy_i18n_test.go:55: R-591: mutating the copy's optional group overlay changed the ORIGINAL — the overlay is shared, not copied + r591_copy_i18n_test.go:55: R-591: mutating the copy's optional help overlay changed the ORIGINAL — the overlay is shared, not copied + r591_copy_i18n_test.go:55: R-591: mutating the copy's integration overlay changed the ORIGINAL — the overlay is shared, not copied + r591_copy_i18n_test.go:59: R-591: adding a language to the copy's I18n map added it to the ORIGINAL — the map is shared +--- FAIL: TestDeepCopyStackI18nIsNotShared (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestDeepCopyStackI18nIsNotShared -v ./internal/stacks) # FIX RESTORED +rc=0 +=== RUN TestDeepCopyStackI18nIsNotShared +--- PASS: TestDeepCopyStackI18nIsNotShared (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s + +==================== R-568 — red-proof ==================== +Fix undone in: internal/web/disk_health.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestDiskHealthRows_OrderIsStable -v ./internal/web) # FIX UNDONE +rc=1 +=== RUN TestDiskHealthRows_OrderIsStable + r568_disk_order_test.go:31: R-568: the same two disks render in a different order depending on the agent's order: [nvme0n1 sda] vs [sda nvme0n1] +--- FAIL: TestDiskHealthRows_OrderIsStable (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.009s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestDiskHealthRows_OrderIsStable -v ./internal/web) # FIX RESTORED +rc=0 +=== RUN TestDiskHealthRows_OrderIsStable +--- PASS: TestDiskHealthRows_OrderIsStable (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.009s + +==================== R-567 — red-proof ==================== +Fix undone in: internal/web/templates/layout.html (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestStorageWizardPages_OpenTheStorageNavGroup -v ./internal/web) # FIX UNDONE +rc=1 +=== RUN TestStorageWizardPages_OpenTheStorageNavGroup + r567_storage_wizard_nav_test.go:21: R-567 storage_init: the storage menu group is not open + r567_storage_wizard_nav_test.go:21: R-567 storage_attach: the storage menu group is not open +--- FAIL: TestStorageWizardPages_OpenTheStorageNavGroup (0.06s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.072s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestStorageWizardPages_OpenTheStorageNavGroup -v ./internal/web) # FIX RESTORED +rc=0 +=== RUN TestStorageWizardPages_OpenTheStorageNavGroup +--- PASS: TestStorageWizardPages_OpenTheStorageNavGroup (0.06s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.069s + +==================== R-363 + R-547 — red-proof ==================== +Fix undone in: cmd/controller/main.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestFillWatchRunsOnAnInterval -v ./cmd/controller) # FIX UNDONE +rc=1 +=== RUN TestFillWatchRunsOnAnInterval + r363_fillwatch_interval_test.go:56: R-363: main.go registers fillWatcher.Check on a periodic (sched.Every, fillWatchInterval) job 0 times, want 1 — without it the fill check is daily only +--- FAIL: TestFillWatchRunsOnAnInterval (0.01s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.016s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestFillWatchRunsOnAnInterval -v ./cmd/controller) # FIX RESTORED +rc=0 +=== RUN TestFillWatchRunsOnAnInterval +--- PASS: TestFillWatchRunsOnAnInterval (0.01s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.017s + +==================== R-10 — red-proof ==================== +Fix undone in: internal/appbackup/dbdump.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestDumpOneTo_SyncsTheDumpDirectoryAfterRename -v ./internal/appbackup) # FIX UNDONE +rc=1 +=== RUN TestDumpOneTo_SyncsTheDumpDirectoryAfterRename + r10_dump_dirsync_test.go:51: R-10: the dump directory was not fsynced after the rename: synced=[], want [/tmp/TestDumpOneTo_SyncsTheDumpDirectoryAfterRename3850100245/002/unit] +--- FAIL: TestDumpOneTo_SyncsTheDumpDirectoryAfterRename (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/appbackup 0.009s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestDumpOneTo_SyncsTheDumpDirectoryAfterRename -v ./internal/appbackup) # FIX RESTORED +rc=0 +=== RUN TestDumpOneTo_SyncsTheDumpDirectoryAfterRename +--- PASS: TestDumpOneTo_SyncsTheDumpDirectoryAfterRename (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/appbackup 0.008s + +==================== R-552 (a: the clear itself) — red-proof ==================== +Fix undone in: internal/backup/restore_record.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR552 -v ./internal/api) # FIX UNDONE +rc=1 +=== RUN TestR552_RemoveClearsTheInterruptedRestoreNotice + r552_remove_clears_notice_test.go:60: R-552: the removed app's interrupted-restore notice is still listed + r552_remove_clears_notice_test.go:67: R-552: the notice comes back after a restart — the clear was not persisted +--- FAIL: TestR552_RemoveClearsTheInterruptedRestoreNotice (0.02s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.023s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR552 -v ./internal/api) # FIX RESTORED +rc=0 +=== RUN TestR552_RemoveClearsTheInterruptedRestoreNotice +--- PASS: TestR552_RemoveClearsTheInterruptedRestoreNotice (0.02s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.024s + +==================== R-552 (b: the wiring in removeStack) — red-proof ==================== +Fix undone in: internal/api/router.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR552 -v ./internal/api) # FIX UNDONE +rc=1 +=== RUN TestR552_RemoveClearsTheInterruptedRestoreNotice + r552_remove_clears_notice_test.go:87: R-552: removeStack does not call clearInterruptedRestoreNotice +--- FAIL: TestR552_RemoveClearsTheInterruptedRestoreNotice (0.02s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.023s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR552 -v ./internal/api) # FIX RESTORED +rc=0 +=== RUN TestR552_RemoveClearsTheInterruptedRestoreNotice +--- PASS: TestR552_RemoveClearsTheInterruptedRestoreNotice (0.02s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.027s + +==================== R-251 — red-proof ==================== +Fix undone in: internal/backup/offbox_inventory.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR251 -v ./internal/backup) # FIX UNDONE +rc=1 +=== RUN TestR251_MarkerTagIsNotAnApp + r251_marker_tag_test.go:48: R-251: the recovery listing shows [calibre-web felhom-offbox]; want only [calibre-web] — the marker tag is not an app + r251_marker_tag_test.go:51: R-251: 2 size calls for one app, want 1 + r251_marker_tag_test.go:59: R-251: the marker tag is reported as an app with a snapshot time +--- FAIL: TestR251_MarkerTagIsNotAnApp (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR251 -v ./internal/backup) # FIX RESTORED +rc=0 +=== RUN TestR251_MarkerTagIsNotAnApp +--- PASS: TestR251_MarkerTagIsNotAnApp (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.007s + +==================== R-104 — red-proof ==================== +Fix undone in: internal/backup/offbox.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR104_SurvivingLockIsNamed -v ./internal/backup) # FIX UNDONE +rc=1 +=== RUN TestR104_SurvivingLockIsNamed + r104_lock_class_test.go:36: R-104: a lock that survived the self-heal is classed "unknown", want "locked" + r104_lock_class_test.go:40: R-104: the Hungarian message does not name the lock: "A távoli mentés ismeretlen okból nem sikerült (1m0s): offbox backup app: exit status 1 (offsite repository is still locked after the self-heal)" + r104_lock_class_test.go:44: R-104: the English message does not name the lock: "The remote backup failed for an unknown reason (1m0s): offbox backup app: exit status 1 (offsite repository is still locked after the self-heal)" + r104_lock_class_test.go:49: R-104: restic's lock text is classed "unknown", want "locked" +--- FAIL: TestR104_SurvivingLockIsNamed (0.01s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.017s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR104_SurvivingLockIsNamed -v ./internal/backup) # FIX RESTORED +rc=0 +=== RUN TestR104_SurvivingLockIsNamed +--- PASS: TestR104_SurvivingLockIsNamed (0.01s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.017s + +==================== R-619 — red-proof ==================== +Fix undone in: internal/api/router.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR619 -v ./internal/api) # FIX UNDONE +rc=1 +=== RUN TestR619_PasswordFieldIsServedAsRequired + r619_password_required_test.go:81: R-619: the password field reaches the wire as required:false, but the deploy refuses without it +--- FAIL: TestR619_PasswordFieldIsServedAsRequired (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/api 0.009s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR619 -v ./internal/api) # FIX RESTORED +rc=0 +=== RUN TestR619_PasswordFieldIsServedAsRequired +--- PASS: TestR619_PasswordFieldIsServedAsRequired (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/api 0.011s + +==================== R-362 — red-proof ==================== +Fix undone in: internal/backup/restore_dir_err.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR362 -v ./internal/backup) # FIX UNDONE +rc=1 +=== RUN TestR362_DetachedDriveIsNamed + r362_restore_drive_gone_test.go:41: R-362: the Hungarian refusal does not name the missing drive: "restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied" + r362_restore_drive_gone_test.go:44: R-362: the English refusal does not name the missing drive: "restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied" + r362_restore_drive_gone_test.go:49: R-362: a drive the registry marks disconnected is not named: "restore dir: mkdir /mnt/hdd_legacy: permission denied" +--- FAIL: TestR362_DetachedDriveIsNamed (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.009s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR362 -v ./internal/backup) # FIX RESTORED +rc=0 +=== RUN TestR362_DetachedDriveIsNamed +[WARN] [backup] restore dir /mnt/felhom-drives/hdd_1/backups/offsite-restore/app: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied — the drive /mnt/felhom-drives/hdd_1 (Kulso HDD) is not connected; reported as a missing drive (R-362) +[WARN] [backup] restore dir /mnt/hdd_legacy/x: mkdir /mnt/hdd_legacy: permission denied — the drive /mnt/hdd_legacy (Regi HDD) is not connected; reported as a missing drive (R-362) +--- PASS: TestR362_DetachedDriveIsNamed (0.01s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.017s + +==================== R-675 — red-proof ==================== +Fix undone in: internal/web/handlers.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR675 -v ./internal/web) # FIX UNDONE +rc=1 +=== RUN TestR675_RefusalNamesTheWholeCopy + r675_refusal_whole_copy_test.go:33: R-675 whole copy on the second drive: the Hungarian refusal reads "Ez a mentés nem tartalmazza az alkalmazás fájljait, ezért nem állítjuk vissza az adatbázist föléjük — a fájlok így a helyükön maradnak. A fájlok a második meghajtó másolatából állíthatók vissza: „Fájlok visszaállítása”."; want it to name "Teljes visszaállítás a másolatból" and not "Fájlok visszaállítása" + r675_refusal_whole_copy_test.go:36: R-675 whole copy on the second drive: the English refusal reads "This backup does not hold the files of the app, so we do not restore the database over them — the files stay where they are. The files can be restored from the copy on the second drive: “Restore files”."; want it to name "Full restore from the copy" +--- FAIL: TestR675_RefusalNamesTheWholeCopy (0.06s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.073s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR675 -v ./internal/web) # FIX RESTORED +rc=0 +=== RUN TestR675_RefusalNamesTheWholeCopy +--- PASS: TestR675_RefusalNamesTheWholeCopy (0.07s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.081s + +==================== R-240 — red-proof ==================== +Fix undone in: internal/backup/offbox.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR240 -v ./internal/backup) # FIX UNDONE +rc=1 +=== RUN TestR240_ZeroSelectionRunDoesNotSaySuccess +[ERROR] [backup] Database discovery failed: docker ps failed: R-650: refused to run the real "docker ps --format {{.ID}}\t{{.Names}}\t{{.Label \"com.docker.compose.project\"}}\t{{.Image}} --filter status=running" under go test (this host may be production Docker); use a seam or a stub on PATH, or set FELHOM_TEST_REAL_DOCKER=1 deliberately +[WARN] [offbox] pre-push dump leg failed (docker ps failed: R-650: refused to run the real "docker ps --format {{.ID}}\t{{.Names}}\t{{.Label \"com.docker.compose.project\"}}\t{{.Image}} --filter status=running" under go test (this host may be production Docker); use a seam or a stub on PATH, or set FELHOM_TEST_REAL_DOCKER=1 deliberately) — continuing with the existing dumps; the snapshot's DB half may be older than its files + r240_zero_selection_test.go:30: R-240: the note for a run that saved nothing still calls itself successful: "Sikeres — nincs mentésre jelölt alkalmazás" + r240_zero_selection_test.go:33: R-240: the note does not say the run saved nothing: "Sikeres — nincs mentésre jelölt alkalmazás" +--- FAIL: TestR240_ZeroSelectionRunDoesNotSaySuccess (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.008s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR240 -v ./internal/backup) # FIX RESTORED +rc=0 +=== RUN TestR240_ZeroSelectionRunDoesNotSaySuccess +[ERROR] [backup] Database discovery failed: docker ps failed: R-650: refused to run the real "docker ps --format {{.ID}}\t{{.Names}}\t{{.Label \"com.docker.compose.project\"}}\t{{.Image}} --filter status=running" under go test (this host may be production Docker); use a seam or a stub on PATH, or set FELHOM_TEST_REAL_DOCKER=1 deliberately +[WARN] [offbox] pre-push dump leg failed (docker ps failed: R-650: refused to run the real "docker ps --format {{.ID}}\t{{.Names}}\t{{.Label \"com.docker.compose.project\"}}\t{{.Image}} --filter status=running" under go test (this host may be production Docker); use a seam or a stub on PATH, or set FELHOM_TEST_REAL_DOCKER=1 deliberately) — continuing with the existing dumps; the snapshot's DB half may be older than its files +--- PASS: TestR240_ZeroSelectionRunDoesNotSaySuccess (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/backup 0.009s + +==================== R-256 (hu copy restored to the old sentence) — red-proof ==================== +Fix undone in: internal/i18n/locales/hu.json (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR256 -v ./internal/web) # FIX UNDONE +rc=1 +=== RUN TestR256_R257_OffboxRefusalsNameARoute + r256_r257_offbox_refusals_test.go:37: R-256: the Hungarian refusal names no route or still names the component: "A mentéskezelő nem elérhető." + r256_r257_offbox_refusals_test.go:45: R-256: the restore page's twin refusal differs: "A mentések kezelése most nem érhető el. Próbáld újra néhány perc múlva; ha akkor sem megy, keresd a Felhom ügyfélszolgálatát." vs "A mentéskezelő nem elérhető." +--- FAIL: TestR256_R257_OffboxRefusalsNameARoute (0.06s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.070s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR256 -v ./internal/web) # FIX RESTORED +rc=0 +=== RUN TestR256_R257_OffboxRefusalsNameARoute +--- PASS: TestR256_R257_OffboxRefusalsNameARoute (0.06s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.075s + +==================== R-257 (hu copy restored to the old sentence) — red-proof ==================== +Fix undone in: internal/i18n/locales/hu.json (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR256 -v ./internal/web) # FIX UNDONE +rc=1 +=== RUN TestR256_R257_OffboxRefusalsNameARoute + r256_r257_offbox_refusals_test.go:67: R-257: the Hungarian refusal still says "offsite": "Az offsite tároló nincs elárvult állapotban." + r256_r257_offbox_refusals_test.go:67: R-257: the Hungarian refusal still says "elárvult": "Az offsite tároló nincs elárvult állapotban." + r256_r257_offbox_refusals_test.go:71: R-257: the Hungarian refusal does not say why or where to go: "Az offsite tároló nincs elárvult állapotban." +--- FAIL: TestR256_R257_OffboxRefusalsNameARoute (0.06s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.071s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR256 -v ./internal/web) # FIX RESTORED +rc=0 +=== RUN TestR256_R257_OffboxRefusalsNameARoute +--- PASS: TestR256_R257_OffboxRefusalsNameARoute (0.07s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.078s + +==================== R-365 — red-proof ==================== +Fix undone in: internal/web/handlers.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR365 -v ./internal/web) # FIX UNDONE +rc=1 +=== RUN TestR365_OverdueCountdownIsNotFutureTense + r365_abandon_overdue_test.go:41: R-365: an overdue countdown does not say the deletion is due since 2026-10-04 + r365_abandon_overdue_test.go:44: R-365: an overdue countdown still renders a past date in the future tense +--- FAIL: TestR365_OverdueCountdownIsNotFutureTense (0.14s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 0.149s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR365 -v ./internal/web) # FIX RESTORED +rc=0 +=== RUN TestR365_OverdueCountdownIsNotFutureTense +--- PASS: TestR365_OverdueCountdownIsNotFutureTense (0.15s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 0.159s + +==================== R-564 — red-proof (gate; Python, so no go test -run) ==================== +Fix undone: SPLIT_PATTERNS emptied in scripts/retrieval_promise_gate.py; decoy harness run. +$ (cd controller && python3 scripts/test_gate_decoys.py | grep split) # FIX UNDONE + ok retrieval-promise/split-verb decoy rejected +FAIL: retrieval-promise/en-ok: rc=1, expected accept +FAIL: retrieval-promise/split-ok: rc=1, expected accept +rc=1 +--- fix restored --- +$ (cd controller && python3 scripts/test_gate_decoys.py | grep split) # FIX RESTORED + ok retrieval-promise/split-verb decoy rejected + ok retrieval-promise/split-ok genuine accepted +all 25 controller decoys behaved — labels do not satisfy these gates +$ python3 scripts/retrieval_promise_gate.py | head -1 +retrieval-promise gate OK — 42 surface(s) incl. 1 Go handler file(s), 18 registered claim(s) + 16 English, none unregistered +NOTE: the record above is NOT a valid red-proof — with SPLIT_PATTERNS emptied the seven new registrations go stale, so the gate fails for that reason, not because it caught the decoy. The valid red-proof follows. + +==================== R-564 — red-proof (valid) ==================== +Decoy planted in hu.json: launcher.link_masolasa = 'A régi mentéseid a kóddal bármikor állíthatók vissza.' (a split-verb retrieval PROMISE) +$ python3 scripts/ # FIX UNDONE +retrieval-promise gate OK — 42 surface(s) incl. 1 Go handler file(s), 11 registered claim(s) + 16 English, none unregistered +rc=0 <- the pre-fix gate PASSES the planted promise (blind) +$ python3 scripts/retrieval_promise_gate.py # FIX IN PLACE + launcher.html:61 unregistered retrieval claim (állíthatók vissza): +RETRIEVAL-PROMISE GATE FAILED: 1 unregistered, 0 stale, across 42 template(s). +rc=1 <- convicted +--- decoy removed --- +$ python3 scripts/retrieval_promise_gate.py +retrieval-promise gate OK — 42 surface(s) incl. 1 Go handler file(s), 18 registered claim(s) + 16 English, none unregistered + +==================== R-425 — red-proof (gate) ==================== +Decoy: a NEW template internal/web/templates/backups_offbox_extra.html containing 'NAS-mentés'. +$ python3 scripts/ # FIX UNDONE +offbox rename gate OK — Tier-3 is 'Tavoli mentes' everywhere customer-facing +rc=0 <- the fixed FILES list never looks at the new file +$ python3 scripts/offbox_rename_gate.py # FIX IN PLACE +internal/web/templates/backups_offbox_extra.html:1 [NAS-mentés]

A NAS-ment\xe9s be\xe1ll\xedt\xe1sa

+OFFBOX RENAME GATE FAILED: 1 customer-facing NAS-branding string(s) remain +rc=1 <- convicted +--- decoy removed --- +offbox rename gate OK — Tier-3 is 'Tavoli mentes' everywhere customer-facing (25 file(s) + 121 bundle value(s) named by them) + +==================== R-565 (the ' mp' unit the new detector found, put back) — red-proof ==================== +Fix undone in: internal/web/templates/backups_remote.html (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestI18nEnglishPages$ -v ./internal/web) # FIX UNDONE +rc=1 +=== RUN TestI18nEnglishPages + i18n_parity_test.go:709: backups_remote_full: ASCII-only Hungarian word "mp" on the English page (R-565), line 434: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_incomplete: ASCII-only Hungarian word "mp" on the English page (R-565), line 348: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_error: ASCII-only Hungarian word "mp" on the English page (R-565), line 333: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_running: ASCII-only Hungarian word "mp" on the English page (R-565), line 353: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_pending_agent: ASCII-only Hungarian word "mp" on the English page (R-565), line 339: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_pending_old: ASCII-only Hungarian word "mp" on the English page (R-565), line 339: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_stale: ASCII-only Hungarian word "mp" on the English page (R-565), line 337: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_stale_old: ASCII-only Hungarian word "mp" on the English page (R-565), line 337: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_escrowed: ASCII-only Hungarian word "mp" on the English page (R-565), line 336: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_notconf_hub: ASCII-only Hungarian word "mp" on the English page (R-565), line 285: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_notconf: ASCII-only Hungarian word "mp" on the English page (R-565), line 284: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_empty: ASCII-only Hungarian word "mp" on the English page (R-565), line 258: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" + i18n_parity_test.go:709: backups_remote_offsite_offer_fit: ASCII-only Hungarian word "mp" on the English page (R-565), line 349: "if(p.elapsed_sec > 0){ label += ' · ' + p.elapsed_sec + ' mp'; }" +--- FAIL: TestI18nEnglishPages (4.23s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/web 4.243s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestI18nEnglishPages$ -v ./internal/web) # FIX RESTORED +rc=0 +=== RUN TestI18nEnglishPages +--- PASS: TestI18nEnglishPages (4.27s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/web 4.285s + +==================== R-603 (an apostrophe planted in a Go-named English value) — red-proof ==================== +Fix undone in: internal/i18n/locales/en.json (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR603 -v ./internal/i18n) # FIX UNDONE +rc=1 +=== RUN TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping + r603_escape_test.go:126: control: a planted apostrophe in err.backup.a_pillanatkep_egy_utvonala_ervenytelen was not caught (got [err.backup.a_pillanatkep_egy_utvonala_ervenytelen note.offsite.fail_locked]) +--- FAIL: TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping (0.29s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/i18n 0.296s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR603 -v ./internal/i18n) # FIX RESTORED +rc=0 +=== RUN TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping +--- PASS: TestR603_GoNamedValuesDoNotHideBehindHTMLEscaping (0.34s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/i18n 0.343s + +==================== R-454 — the gate seen RED on the unformatted tree, before the formatting pass ==================== +$ (cd controller && python3 scripts/gofmt_gate.py) + not gofmt-clean: cmd/controller/main.go + not gofmt-clean: internal/agentapi/diskverdict.go + not gofmt-clean: internal/api/update_reason_test.go + not gofmt-clean: internal/appbackup/namespace_root_test.go + not gofmt-clean: internal/appbackup/r381_undo_naming_test.go + not gofmt-clean: internal/backup/r669_applied_meta_test.go + not gofmt-clean: internal/family/family.go + not gofmt-clean: internal/infra/infra.go + not gofmt-clean: internal/notify/r636_oom_storm_test.go + not gofmt-clean: internal/quiesce/tiers_test.go + not gofmt-clean: internal/stacks/delete.go + not gofmt-clean: internal/stacks/life_records.go +GOFMT GATE FAILED: 12 file(s) — run `gofmt -w ` (formatting only, no behaviour change) +rc=1 + +==================== R-208 (controller half: ARGs moved back above go mod download) — red-proof ==================== +Fix undone in: Dockerfile (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR208 -v ./cmd/controller) # FIX UNDONE +rc=0 +=== RUN TestR208_DockerfileVersionArgsSitBelowModuleDownload +--- PASS: TestR208_DockerfileVersionArgsSitBelowModuleDownload (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.009s +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR208 -v ./cmd/controller) # FIX RESTORED +rc=0 +=== RUN TestR208_DockerfileVersionArgsSitBelowModuleDownload +--- PASS: TestR208_DockerfileVersionArgsSitBelowModuleDownload (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.006s +NOTE: NOT CONVICTED above — the test kept the LAST declaration of each ARG, so a duplicate declaration above the download hid behind the one below. Test fixed to keep the FIRST declaration; red-proof re-run below. + +==================== R-208 (re-run after the test fix) — red-proof ==================== +Fix undone in: Dockerfile (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestR208 -v ./cmd/controller) # FIX UNDONE +rc=1 +=== RUN TestR208_DockerfileVersionArgsSitBelowModuleDownload + r208_dockerfile_args_test.go:50: R-208: ARG VERSION (line 12) is declared above `go mod download` (line 19), so every build with a new value re-downloads the modules + r208_dockerfile_args_test.go:50: R-208: ARG GIT_COMMIT (line 13) is declared above `go mod download` (line 19), so every build with a new value re-downloads the modules +--- FAIL: TestR208_DockerfileVersionArgsSitBelowModuleDownload (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.009s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestR208 -v ./cmd/controller) # FIX RESTORED +rc=0 +=== RUN TestR208_DockerfileVersionArgsSitBelowModuleDownload +--- PASS: TestR208_DockerfileVersionArgsSitBelowModuleDownload (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/cmd/controller 0.008s + +==================== R-591 follow-up (DataPaths + AfterLoad copies removed) — red-proof ==================== +Fix undone in: internal/stacks/manager.go (the fix text replaced by the pre-fix shape) +$ (cd controller && go test -count=1 -run TestDeepCopyStackMetaSharesNoReference -v ./internal/stacks) # FIX UNDONE +rc=1 +=== RUN TestDeepCopyStackMetaSharesNoReference + r591_copy_meta_alias_test.go:21: R-591: deepCopyStack leaves Meta.AfterLoad shared with the original — a write through the copy changes the stack + r591_copy_meta_alias_test.go:21: R-591: deepCopyStack leaves Meta.DataPaths shared with the original — a write through the copy changes the stack +--- FAIL: TestDeepCopyStackMetaSharesNoReference (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s +FAIL +--- fix restored --- +$ (cd controller && go test -count=1 -run TestDeepCopyStackMetaSharesNoReference -v ./internal/stacks) # FIX RESTORED +rc=0 +=== RUN TestDeepCopyStackMetaSharesNoReference +--- PASS: TestDeepCopyStackMetaSharesNoReference (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-controller/internal/stacks 0.007s diff --git a/documentation/audits/burndown2-2026-10-05/delivery/sign-agent-update.txt b/documentation/audits/burndown2-2026-10-05/delivery/sign-agent-update.txt new file mode 100644 index 00000000..7757075b --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/delivery/sign-agent-update.txt @@ -0,0 +1,13 @@ +== agent_update 0.147.0 (sha 642c4d19…) signed with felhom-op-1, ttl 45m, 2026-10-05T17:02:13Z +-- demo-hp-bb76ea +signed: op=agent_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=b05c992ea1b100af9f53df761c76d440 expires=2026-10-05T17:47:13Z +wrote envelope to /env-demo-hp-bb76ea-agent_update.json +uploaded signed op to the hub jobs queue +-- demo-felhom-8363b5 +signed: op=agent_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=67f8cbeed47a8a107d3809e230837af0 expires=2026-10-05T17:47:13Z +wrote envelope to /env-demo-felhom-8363b5-agent_update.json +uploaded signed op to the hub jobs queue +-- tester-1-d70be4 +signed: op=agent_update host=tester-1-d70be4 guest="" key_id=felhom-op-1 nonce=c0decf168b76b852d71db79378bcb795 expires=2026-10-05T17:47:13Z +wrote envelope to /env-tester-1-d70be4-agent_update.json +uploaded signed op to the hub jobs queue diff --git a/documentation/audits/burndown2-2026-10-05/delivery/sign-bundle.txt b/documentation/audits/burndown2-2026-10-05/delivery/sign-bundle.txt new file mode 100644 index 00000000..40b8559c --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/delivery/sign-bundle.txt @@ -0,0 +1,14 @@ +== System page 2026-10-05T17:13:38Z after agent_update: demo-hp, demo-felhom, tester-1 rows read Agent 0.147.0 (root files 0.146.1); Tester-2 0.142.0 → 0.147.0 (offline) +== agent_config_update 0.147.0 (bundle sha 326527d0…), 2026-10-05T17:13:38Z +-- demo-hp-bb76ea +signed: op=agent_config_update host=demo-hp-bb76ea guest="" key_id=felhom-op-1 nonce=f0d53302c1591959f9577dd2660690d1 expires=2026-10-05T17:58:38Z +wrote envelope to /env-demo-hp-bb76ea-agent_config_update.json +uploaded signed op to the hub jobs queue +-- demo-felhom-8363b5 +signed: op=agent_config_update host=demo-felhom-8363b5 guest="" key_id=felhom-op-1 nonce=c92ab23d76d00821a36ec20018c69b6a expires=2026-10-05T17:58:38Z +wrote envelope to /env-demo-felhom-8363b5-agent_config_update.json +uploaded signed op to the hub jobs queue +-- tester-1-d70be4 +signed: op=agent_config_update host=tester-1-d70be4 guest="" key_id=felhom-op-1 nonce=c51c5f6448ec01d4919408e0590dd1bd expires=2026-10-05T17:58:38Z +wrote envelope to /env-tester-1-d70be4-agent_config_update.json +uploaded signed op to the hub jobs queue diff --git a/documentation/audits/burndown2-2026-10-05/delivery/vouch-agent.txt b/documentation/audits/burndown2-2026-10-05/delivery/vouch-agent.txt new file mode 100644 index 00000000..755df8c6 --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/delivery/vouch-agent.txt @@ -0,0 +1,4 @@ +== vouch 2026-10-05T17:01:27Z: POST /configuration/artifacts (Basic + X-Felhom-Operator), agent 0.147.0, golden 0.296.0, min_agent 0.131.0 +HTTP/1.1 303 See Other +Location: /configuration?flash=artifacts_set +2026/10/05 19:01:56 [INFO] Artifact manifest set: agent=0.147.0 golden=0.296.0 min_agent="0.131.0" wrapper_sha=false bundle_sha="326527d0993c9a62df2f790c7700ca645cedbf0673dcfb6dc1768d8610b8007d" diff --git a/documentation/audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt b/documentation/audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt new file mode 100644 index 00000000..828e6d35 --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt @@ -0,0 +1,366 @@ +# felhom.eu burndown2 red-proofs, 2026-10-05 (fix undone -> test fails; fix restored -> test passes) + +### R-277 (offsite row bytes) +$ (cd . && go test ./internal/web -run 'TestOffsiteRow_SmallRepoNotZeroGB' -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestOffsiteRow_SmallRepoNotZeroGB + r277_r92_bytes_test.go:27: 162 KB repo rendered as "0.0 GB", want "162.0 KB" +--- FAIL: TestOffsiteRow_SmallRepoNotZeroGB (0.04s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.062s +FAIL +--- fix RESTORED (rc=0): +=== RUN TestOffsiteRow_SmallRepoNotZeroGB +--- PASS: TestOffsiteRow_SmallRepoNotZeroGB (0.04s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.064s + +### R-92 (PBS DR exact bytes) +$ (cd . && go test ./internal/web -run 'TestPBSDRPanel_ExactBytesShowsSmallDelta' -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestPBSDRPanel_ExactBytesShowsSmallDelta + r277_r92_bytes_test.go:55: PBS DR panel must show exact bytes: + PBS DR datastore + + + + + +
Datastorefelhom-offsite (ep0)
Capacity40.0 GB
Used8.0 GB · 20% full
+
+ + +
Polledjust now
+ + + +

+ The endpoint below IS the PBS DR host. Peer allocation and endpoint sync currently use the + lowest endpoint id (ep0); per-endpoint allocation is a future work item. +

+ + +
+

Endpoint

+

Not configured. Add one below (or via PUT /api/v1/admin/wg/endpoint, runbook: offsite-endpoint.md).

+
+ + + +
+

Add endpoint

+
+ + + + + + + +--- fix RESTORED (rc=0): +=== RUN TestPBSDRPanel_ExactBytesShowsSmallDelta +--- PASS: TestPBSDRPanel_ExactBytesShowsSmallDelta (0.07s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.093s + + +### R-581 (newest report tie-break) +$ (cd . && go test ./internal/store -run TestGetCustomers_SameSecondReportsOneRowNewestWins -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestGetCustomers_SameSecondReportsOneRowNewestWins + r581_newest_report_test.go:32: GetCustomers returned 2 rows for one customer, want exactly 1 +--- FAIL: TestGetCustomers_SameSecondReportsOneRowNewestWins (0.03s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/store 0.040s +FAIL +--- fix RESTORED (rc=0): +=== RUN TestGetCustomers_SameSecondReportsOneRowNewestWins +--- PASS: TestGetCustomers_SameSecondReportsOneRowNewestWins (0.03s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/store 0.040s + +### R-600 (delete cascade WG peer push + honest COMPLETE line) +$ (cd . && go test ./internal/web -run 'TestDeleteCascade_TriggersWGPeerPushAndSaysSo' -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestDeleteCascade_TriggersWGPeerPushAndSaysSo + customer_delete_test.go:642: WG peer-sync triggers = 0, want exactly 1 (the endpoint must drop the peer now, not on the next tick) +--- FAIL: TestDeleteCascade_TriggersWGPeerPushAndSaysSo (0.04s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.070s +FAIL +--- fix RESTORED (rc=0): +=== RUN TestDeleteCascade_TriggersWGPeerPushAndSaysSo +--- PASS: TestDeleteCascade_TriggersWGPeerPushAndSaysSo (0.05s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.072s + +### R-600 (main wiring of SetWGPeerSync) +$ (cd . && go test ./cmd/hub -run TestR600_MainWiresWGPeerSyncIntoWeb -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestR600_MainWiresWGPeerSyncIntoWeb + r600_wiring_test.go:32: cmd/hub/main.go never calls webServer.SetWGPeerSync(.Trigger) — the delete cascade cannot push the peer removal +--- FAIL: TestR600_MainWiresWGPeerSyncIntoWeb (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.024s +FAIL +--- fix RESTORED (rc=0): +=== RUN TestR600_MainWiresWGPeerSyncIntoWeb +--- PASS: TestR600_MainWiresWGPeerSyncIntoWeb (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.025s + +### R-599 (ONLINE 409 says when deletion opens; host + cascade) +$ (cd . && go test ./internal/web -run TestHostDelete_OnlineRefusalSaysWhenItOpens -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestHostDelete_OnlineRefusalSaysWhenItOpens + r544_r599_host_delete_test.go:74: host-delete 409 body missing "min ago": + Host is ONLINE — deletion is refused (a live agent would receive 401s permanently). + r544_r599_host_delete_test.go:74: host-delete 409 body missing "deletion opens at 17:42 UTC": + Host is ONLINE — deletion is refused (a live agent would receive 401s permanently). + r544_r599_host_delete_test.go:74: host-delete 409 body missing "45m0s": + Host is ONLINE — deletion is refused (a live agent would receive 401s permanently). + r544_r599_host_delete_test.go:84: cascade ONLINE refusal = 409 "Delete refused: host gone-vm is ONLINE. Decommission the box first — the cascade never deletes a live host.\n", want 409 naming the opening time +--- FAIL: TestHostDelete_OnlineRefusalSaysWhenItOpens (0.05s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.066s +FAIL +--- fix RESTORED (rc=0): +=== RUN TestHostDelete_OnlineRefusalSaysWhenItOpens +--- PASS: TestHostDelete_OnlineRefusalSaysWhenItOpens (0.04s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.065s + +### R-544 (host-delete log states demotion) +$ (cd . && go test ./internal/web -run TestHostDelete_LogSaysEscrowDemotedNotDeleted -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestHostDelete_LogSaysEscrowDemotedNotDeleted + r544_r599_host_delete_test.go:32: log still says the escrow was deleted: + [INFO] host deleted: esc-host (escrow deleted: true) + r544_r599_host_delete_test.go:35: log must state the demotion: + [INFO] host deleted: esc-host (escrow deleted: true) +--- FAIL: TestHostDelete_LogSaysEscrowDemotedNotDeleted (0.04s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.064s +FAIL +--- fix RESTORED (rc=0): +=== RUN TestHostDelete_LogSaysEscrowDemotedNotDeleted +--- PASS: TestHostDelete_LogSaysEscrowDemotedNotDeleted (0.04s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.063s + +### R-855 (start log prints effective Docker nights) +$ (cd . && go test ./cmd/hub ./internal/osupdates -run 'TestR855_StartLogPrintsEffectiveDockerNights|TestDockerNightsEffective' -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestR855_StartLogPrintsEffectiveDockerNights + r855_docker_nights_log_test.go:41: the Docker approval-nights start log must print osSvc.DockerNightsEffective(), not the raw field +--- FAIL: TestR855_StartLogPrintsEffectiveDockerNights (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.022s +=== RUN TestDockerNightsEffective +--- PASS: TestDockerNightsEffective (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/osupdates 0.005s +FAIL +--- fix RESTORED (rc=0): +=== RUN TestR855_StartLogPrintsEffectiveDockerNights +--- PASS: TestR855_StartLogPrintsEffectiveDockerNights (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/cmd/hub 0.022s +=== RUN TestDockerNightsEffective +--- PASS: TestDockerNightsEffective (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/osupdates 0.006s + +### R-134 (zone candidates strip progressively; red = the old one-label parentDomain) +$ (cd . && go test ./internal/cloudflare -run TestZoneCandidates -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestZoneCandidates + unblock_test.go:24: zoneCandidates("a.b.felhom.eu") = [a.b.felhom.eu b.felhom.eu], want [a.b.felhom.eu b.felhom.eu felhom.eu] + unblock_test.go:24: zoneCandidates("x.y.z.example.co") = [x.y.z.example.co y.z.example.co], want [x.y.z.example.co y.z.example.co z.example.co example.co] +--- FAIL: TestZoneCandidates (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/cloudflare 0.006s +FAIL +--- fix RESTORED (rc=0): +=== RUN TestZoneCandidates +--- PASS: TestZoneCandidates (0.00s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/cloudflare 0.006s + + +### R-292 (artifact sha flash names the cause; red = every failure -> artifact_sha_invalid, the old single flash) +$ (cd . && go test ./internal/web -run 'TestArtifactSave_FlashNamesTheCause' -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestArtifactSave_FlashNamesTheCause +=== RUN TestArtifactSave_FlashNamesTheCause/version_not_found + r292_artifact_flash_test.go:57: flash = "artifact_sha_invalid", want "artifact_version_missing" +=== RUN TestArtifactSave_FlashNamesTheCause/registry_failing + r292_artifact_flash_test.go:57: flash = "artifact_sha_invalid", want "artifact_unverifiable" +=== RUN TestArtifactSave_FlashNamesTheCause/no_sha_listed + r292_artifact_flash_test.go:57: flash = "artifact_sha_invalid", want "artifact_sha_missing" +=== RUN TestArtifactSave_FlashNamesTheCause/healthy +--- FAIL: TestArtifactSave_FlashNamesTheCause (0.24s) + --- FAIL: TestArtifactSave_FlashNamesTheCause/version_not_found (0.06s) + --- FAIL: TestArtifactSave_FlashNamesTheCause/registry_failing (0.07s) + --- FAIL: TestArtifactSave_FlashNamesTheCause/no_sha_listed (0.07s) + --- PASS: TestArtifactSave_FlashNamesTheCause/healthy (0.04s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.264s +--- fix RESTORED (rc=0): +=== RUN TestArtifactSave_FlashNamesTheCause +=== RUN TestArtifactSave_FlashNamesTheCause/version_not_found +=== RUN TestArtifactSave_FlashNamesTheCause/registry_failing +=== RUN TestArtifactSave_FlashNamesTheCause/no_sha_listed +=== RUN TestArtifactSave_FlashNamesTheCause/healthy +--- PASS: TestArtifactSave_FlashNamesTheCause (0.17s) + --- PASS: TestArtifactSave_FlashNamesTheCause/version_not_found (0.04s) + --- PASS: TestArtifactSave_FlashNamesTheCause/registry_failing (0.04s) + --- PASS: TestArtifactSave_FlashNamesTheCause/no_sha_listed (0.04s) + --- PASS: TestArtifactSave_FlashNamesTheCause/healthy (0.04s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.188s + +### R-725 (expired bind page points at the button) +$ (cd . && go test ./internal/web -run TestBindExpiredPage_PointsAtTheButton -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestBindExpiredPage_PointsAtTheButton + r725_bind_expired_copy_test.go:28: the expired page with a button does not carry the sentence that points at it + r725_bind_expired_copy_test.go:31: the expired page with a button still sends the household to support +--- FAIL: TestBindExpiredPage_PointsAtTheButton (0.04s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.065s +FAIL +--- fix RESTORED (rc=0): +=== RUN TestBindExpiredPage_PointsAtTheButton +--- PASS: TestBindExpiredPage_PointsAtTheButton (0.04s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.057s + +### R-725 (console banner glyph; red = the check mark restored in script + golden) +$ (cd . && python3 scripts/iso/test/test_console_glyphs.py) +--- fix UNDONE (rc=1): +FAIL: console glyphs the Latin-2 font cannot draw (R-725): +--- fix RESTORED (rc=0): +OK: 58 printf literals + 3 goldens use only console-safe glyphs + +### R-728 (in-flight create guard; red = guard removed, seam kept) +$ (cd . && go test ./internal/web -run TestConfigCreate_ConcurrentSubmitCreatesOnce -v -count=1) +--- fix UNDONE (rc=1): +=== RUN TestConfigCreate_ConcurrentSubmitCreatesOnce + r728_create_once_test.go:57: customer created 2 times for one press, want exactly 1: + [INFO] Customer config created: tester-2 + [INFO] self-bind link NOT auto-minted for tester-2 on customer creation: no mailer configured on this hub + [INFO] Customer config created: tester-2 + [INFO] self-bind link NOT auto-minted for tester-2 on customer creation: no mailer configured on this hub +--- FAIL: TestConfigCreate_ConcurrentSubmitCreatesOnce (0.04s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-hub/internal/web 0.064s +FAIL +--- fix RESTORED (rc=0): +=== RUN TestConfigCreate_ConcurrentSubmitCreatesOnce +--- PASS: TestConfigCreate_ConcurrentSubmitCreatesOnce (0.05s) +PASS +ok gitea.dooplex.hu/admin/felhom-hub/internal/web 0.065s + +### R-208 hub half (Dockerfile ARG order; red = ARGs back above go mod download) +$ (cd . && python3 scripts/test_dockerfile_arg_order.py) +--- fix UNDONE (rc=1): +FAIL: hub/Dockerfile (R-208): +--- fix RESTORED (rc=0): +OK: hub/Dockerfile declares its per-build ARGs below the module download + + +### R-819 (check_stands accepts CLOSED rows; red = rule 3 reads OPEN-ITEMS.md only) +$ (cd . && python3 scripts/check_stands.py) +--- fix UNDONE (rc=1): +CONVICTED — 34 problem(s): +--- fix RESTORED (rc=0): +check_stands: OK — every claim cites a source, every citation resolves, and every 'walked' cites a walk. + +### R-819 decoy suite (the genuine closed-row stand must pass; red = OPEN-only rule 3) +$ (cd . && python3 scripts/test_gate_decoys.py) +--- fix UNDONE (rc=1): + ok hub-confirm/subdir decoy rejected + ok manifest-bearer/subdir decoy rejected + ok observations/R-419 decoy rejected + ok observations/genuine-FILED genuine accepted + ok observations/genuine-NAF genuine accepted + ok reuse-refs/missing-go decoy rejected + ok reuse-refs/missing-md (KNOWN HOLE R-422) genuine accepted + ok golden-currency empty dir rejected AND named + ok closed-register/body-word (BY DESIGN) genuine accepted + ok closed-register/verdict-word decoy rejected + ok closed-register/unreadable-row decoy rejected + ok closed-register/duplicate-closed-id decoy rejected + ok closed-register/finished-row-in-open decoy rejected + ok closed-register/open-row-closed-word (BY DESIGN) genuine accepted + ok one-register/suffix-id-row decoy rejected +--- fix RESTORED (rc=0): + ok hub-confirm/subdir decoy rejected + ok manifest-bearer/subdir decoy rejected + ok observations/R-419 decoy rejected + ok observations/genuine-FILED genuine accepted + ok observations/genuine-NAF genuine accepted + ok reuse-refs/missing-go decoy rejected + ok reuse-refs/missing-md (KNOWN HOLE R-422) genuine accepted + ok golden-currency empty dir rejected AND named + ok closed-register/body-word (BY DESIGN) genuine accepted + ok closed-register/verdict-word decoy rejected + ok closed-register/unreadable-row decoy rejected + ok closed-register/duplicate-closed-id decoy rejected + ok closed-register/finished-row-in-open decoy rejected + ok closed-register/open-row-closed-word (BY DESIGN) genuine accepted + ok one-register/suffix-id-row decoy rejected + +### R-857 (golden gate reads the re-bake; red = the old regex + version-only sort) +$ (cd . && python3 scripts/test_golden_currency_gate.py) +--- fix UNDONE (rc=1): +CASE 17 ok (R-857): a later-dated bake of the same version wins +FAIL: CASE 16 (R-857): two bakes of one version — the gate did not report the RE-BAKE's sha +golden currency gate OK — the newest released controller has a golden (NOTE: this checks the BAKE, not the vouch — see the module docstring) +--- fix RESTORED (rc=0): +CASE 16 ok (R-857): of two bakes of one version, the re-bake's sha is the one reported +CASE 17 ok (R-857): a later-dated bake of the same version wins +golden-currency gate self-test OK — a directory name alone cannot satisfy it (R-410), and the waiver is judged on its dates, its row and its direction (2026-09-13) + + +### R-555 (wire-contract strips comments; red = whole-file tokenising as before) +$ (cd . && python3 scripts/test_gate_decoys.py > /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/d.out 2>&1; rc=$?; grep -E 'FAIL: wire|ok wire|behaved|exposed a hole' /tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/d.out; exit $rc) +--- fix UNDONE (rc=1): + ok wire-contract/genuine genuine accepted +FAIL: wire-contract/comment: rc=0, want 10 (10 = the comment-only tag convicted, 0 = passed) +--- fix RESTORED (rc=0): + ok wire-contract/comment decoy rejected + ok wire-contract/genuine genuine accepted + + +### R-364 (hu_grep refuses an untested zero; red = the anchor check disabled; no bytecode cache) +$ (cd . && PYTHONDONTWRITEBYTECODE=1 python3 -B scripts/test_hu_grep.py) +--- fix UNDONE (rc=1): +ok present accented string counted rc=0 1 line(s) +ok octal-escaped pattern REFUSED rc=2 REFUSED: the pattern arrived TRANSFORMED ('k\\303\\251rj') — octal esc +ok U+FFFD pattern REFUSED rc=2 REFUSED: the pattern arrived TRANSFORMED ('k�rj') — octal escapes or U +ok no anchor given REFUSED rc=2 REFUSED: an accented pattern needs --anchor with ASCII text known to b +ok NFD file, NFC pattern REFUSED rc=2 REFUSED: 0 for the pattern as typed, but 1 line(s) hold its NFD form — +ok tested zero is a zero rc=1 0 line(s) — a TESTED zero: anchor 'Felhom' found on 1 line(s), negativ +ok ascii pattern plain zero rc=1 0 line(s) +FAIL: + blind instrument REFUSED: rc=1 msg="0 line(s) — a TESTED zero: anchor 'NotInTheFile' found on 0 line(s), negative control 0, other normal form 0", want rc=2 containing 'anchor' +--- fix RESTORED (rc=0): +ok present accented string counted rc=0 1 line(s) +ok octal-escaped pattern REFUSED rc=2 REFUSED: the pattern arrived TRANSFORMED ('k\\303\\251rj') — octal esc +ok U+FFFD pattern REFUSED rc=2 REFUSED: the pattern arrived TRANSFORMED ('k�rj') — octal escapes or U +ok blind instrument REFUSED rc=2 REFUSED: the anchor 'NotInTheFile' was not found either — the instrume +ok no anchor given REFUSED rc=2 REFUSED: an accented pattern needs --anchor with ASCII text known to b +ok NFD file, NFC pattern REFUSED rc=2 REFUSED: 0 for the pattern as typed, but 1 line(s) hold its NFD form — +ok tested zero is a zero rc=1 0 line(s) — a TESTED zero: anchor 'Felhom' found on 1 line(s), negativ +ok ascii pattern plain zero rc=1 0 line(s) +hu_grep: OK — no untested zero for an accented pattern + +### R-587 (release build refuses a stray *.rootpw.txt; red = the guard body emptied; run under BusyBox PATH too) +$ (cd . && PATH=/tmp/claude-1000/-mnt-5-hdd-felhom-eu-git/ea20e5ca-a93a-4939-bd73-dd472bcf0590/scratchpad/bb python3 scripts/iso/test/test_rootpw_guard.py) +--- fix UNDONE (rc=1): +FAIL (R-587): +--- fix RESTORED (rc=0): +OK: a release build refuses a *.rootpw.txt in its out dir; a clean dir and a non-release build proceed diff --git a/documentation/audits/burndown2-2026-10-05/r124-red-proof.txt b/documentation/audits/burndown2-2026-10-05/r124-red-proof.txt new file mode 100644 index 00000000..3f0662ba --- /dev/null +++ b/documentation/audits/burndown2-2026-10-05/r124-red-proof.txt @@ -0,0 +1,5 @@ +### R-124 red-proof: PBSRootNamespace back to "root" + dr_recipe_test.go:478: root namespace on the wire = "root", want "" (PBS's spelling; no namespace is named "root") +--- FAIL: TestR124_RootNamespaceOnTheWireIsPBSSpelling (0.00s) +FAIL +FAIL gitea.dooplex.hu/admin/felhom-agent/internal/hub 0.008s diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 485ba3b6..c32a0557 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -75,6 +75,34 @@ The full text of every row below: `git show e8c56c44:documentation/backlog/OPEN- | **R-704** | **[P3-LOW] The box's crash-loop stop (decision 28) outlives the app: after remove and reinstall, the new install is still held.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: A fresh install drops an old update hold: fixed in v0.278.0 with tests; not seen live. Fix would cost: a live install with a leftover hold. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-706** | **[P3-LOW] Removing an app "with its backups" leaves its off-site verification copy on the drive.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: Removing an app with its backups also deletes its off-site test copy: fixed in v0.279.0 with tests; not seen live. Fix would cost: a live remove with an off-site copy. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | | **R-723** | **[P3-LOW] A fresh box sends the operator two mails on day one that describe nothing wrong.** (P3) | CLOSED 2026-10-05 — ACCEPTED BY THE OPERATOR (ruling 2026-10-05 18:23, the burn-down list) | Reason on the list: No 'box recovered' alarm in a new box's first hour: fixed in hub v0.126.0 with tests; not seen at a real first install. Fix would cost: watch the next real install. If never: the tests stay the proof. Evidence of the verdict: `audits/burndown-2026-10-05/partA-table.md`. | +| **R-277** | **Three hub surfaces jointly present a HEALTHY off-site tier as an absent one — and it produced a wrong operator statement during this run.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): `hub/internal/web/offsite_box.go` — a repo under 1 GB is not „0.0 GB"; `TestOffsiteRow_SmallRepoNotZeroGB`; red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-581** | **[P2-MED] `ORDER BY received_at` cannot answer "the newest report" — the column has SECOND granularity.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): `GetCustomers` joins on `MAX(id)`; `TestGetCustomers_SameSecondReportsOneRowNewestWins` (old query: 2 rows); red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-600** | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): delete cascade triggers the WG peer push and says so; wired in main; `TestDeleteCascade_*`, `TestR600_MainWiresWGPeerSyncIntoWeb`; red-proofs `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-544** | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): host-delete log states the escrow effect; `TestHostDelete_LogSaysEscrowDemotedNotDeleted`; red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-855** | **The hub's start log prints "after -1 healthy ring-0 night(s)" for the TEST override `OS_DOCKER_APPROVE_NIGHTS=0`** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): `DockerNightsEffective` printed in the start log; `TestDockerNightsEffective`, `TestR855_StartLogPrintsEffectiveDockerNights`; red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-134** | **Two zone-resolvers disagree on depth.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): `cloudflare.zoneCandidates`; `TestZoneCandidates`; red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-92** | Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): PBS-DR panel exact bytes; `TestPBSDRPanel_ExactBytesShowsSmallDelta`; red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-292** | **The artifact-save flash conflates three different facts, and a failing test found it rather than a reading.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): artifact save flashes name the cause (new `artifact_version_missing`); `TestArtifactSave_FlashNamesTheCause` +2; red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-599** | **[P3-LOW] A drill's teardown is blocked for 30 minutes by design, and nothing says so.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): 409 bodies say last report + when deletion opens; `TestHostDelete_OnlineRefusalSaysWhenItOpens`; target-selection names the wait; red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-725** | **[P3-LOW] Small copy slips on the first-hour path.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): `bind.invalid.body_resend` (hu+en) points at the button; ISO glyph removed (ships with the next ISO); `TestBindExpiredPage_PointsAtTheButton`, `iso/test/test_console_glyphs.py`; red-proofs `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-728** | **[P3-LOW] A customer created with one press was created TWICE, and the first of its two connect mails holds a dead link.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): per-customer-ID in-flight guard; `TestConfigCreate_ConcurrentSubmitCreatesOnce` (old: two creates for one press; passes under -race); red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-571** | **[P3-LOW] The off-site failure classifier and the dashboard's alert-placement rules are described in no architecture document.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): `07` §6.7 the six off-site failure classes and what each means for the household; `02` alert placement. Docs only. | +| **R-819** | **`scripts/check_stands.py` is red and runs in no runner.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): `check_stands.py` rule 3 accepts CLOSED-ITEMS ids; `stands` gate registered with 3 decoys (the status-agreement half not built — too vague); red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-857** | **Baking a golden twice under the SAME version leaves two stale facts.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): golden gate reads suffixed bake dirs and the newest bake log; runbook re-vouch line; CASE 16/17; red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-555** | **[P3-LOW] The wire-contract gate counts a field as received when its name appears in a Go COMMENT on the receiving side.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): wire gate strips comments before matching; 6 surfaced fields allow-listed (2 → R-888); decoy, exemption removed; red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-364** | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): `scripts/hu_grep.py` + `test_hu_grep.py`; pointer in `REUSE.md`; red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-587** | **[P3-LOW] Two root-password files sit in the directory the public ISO is published FROM.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): release ISO build refuses with a `*.rootpw.txt` in its out dir; skill publish block stops on one; `iso/test/test_rootpw_guard.py` (BusyBox PATH); red-proof `audits/burndown2-2026-10-05/felhom-eu-red-proofs.txt` | +| **R-129** | **Every doc says demo-hp has "no baked SSH key"** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom.eu (hub v0.137.0 commit, this one): demo-hp authenticates with DooPlex's own key (`ssh -o BatchMode=yes demo-hp` → ok; authorized_keys holds `SHA256:pgQh228R…`, the same as DooPlex's `~/.ssh/id_ed25519`, deliberate); `operations/nodes.md` „Access", `target-selection.md`, memory corrected; G1 stays the fallback. | +| **R-593** | **[P3-LOW] papra's session-signing key is described as „the app's subdomain".** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | app-catalog `4828dc7` (CI run 1368, re-run success): papra SUBDOMAIN/AUTH_SECRET descriptions; `test_deploy_field_descriptions.py`; red-proof `audits/burndown2-2026-10-05/catalog-red-proofs.txt` | +| **R-760** | **[P3-LOW] vikunja's compose has no healthcheck, and nothing says why.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | app-catalog `4828dc7`: vikunja compose says why there is no healthcheck (comment only); `test_healthcheck_explained.py`; red-proof catalog-red-proofs.txt | +| **R-594** | **[P3-LOW] The catalog copy gate can CONVICT a retrieval promise but has no way to REGISTER a true one.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | app-catalog `4828dc7`: `check-copy-i18n.py` English allow-list (`copy_freeze/allowlist_en.json`); 5 decoys; red-proof catalog-red-proofs.txt | +| **R-605** | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | app-catalog `4828dc7`: a refusing harness exits 3, the runner prints DID-NOT-RUN; decoys both ways; red-proof catalog-red-proofs.txt. (The CLAUDE.md exit-code line was NOT updated — permission check refused the instruction-file edit.) | +| **R-781** | **[P3-LOW] The catalog's `scripts/test_gate_decoys.py` fails 4 of its own "genuine" onboarding cases — on the untouched tree.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | app-catalog `4828dc7`: onboarding decoys run on a scratch clone without the real records; 137 decoy cases pass; red-proof catalog-red-proofs.txt | +| **R-124** | **The recipe spells PBS's root namespace `"root"`, but the PBS API spells it `""`** (P4) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-agent v0.147.0 (`f1b9b41`, tag v0.147.0, sha256 `642c4d19…`; delivered by signed jobs 2026-10-05 — `audits/burndown2-2026-10-05/delivery/`): the recipe's root namespace is `""` beside `namespace_state: resolved`; `TestR124_RootNamespaceOnTheWireIsPBSSpelling` (red-proof: "root" → FAIL, `r124-red-proof.txt`); runbook `ep0-datastore-copy.md` step 2 says how to read it. Operator ruling 2026-10-05 18:23: fix it. | +| **R-118** | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-agent v0.147.0 (delivered): capacity read only while the device is present; `TestDisks_UnionPath_AbsentDeviceReportsNoRootCapacity`; red-proof `audits/burndown2-2026-10-05/agent-red-proofs.txt` | +| **R-269** | **A rotated-out per-guest local-API token still authorises, and the test that appears to pin the opposite passes only because of its lookup ORDER.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-agent v0.147.0 (delivered): token store reloads on growth before answering; `TestTokenStore_RotatedOutTokenRejectedFirst`; red-proof agent-red-proofs.txt | +| **R-317** | **The agent decides whether to install dnsmasq by stat-ing a file the OTHER package owns.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down round 2) | felhom-agent v0.147.0 (delivered): dnsmasq install probed by its service unit; `TestEnsureDnsmasq_*`; red-proof agent-red-proofs.txt | +| **R-350** | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** (P3) | CLOSED 2026-10-05 — DUPLICATE of R-132 (its unique facts moved there) | Same credential (hub operator password HUB_PW), same mechanism (curl -w '%{redirect_url}' re-renders Basic-auth into the URL), same single action (operator decides to rotate). R-132 already folded R-580 (third occurrence 2026-09-18) on 2026-10-03; R-350 is the 2026-08-20 occurrence. | --- diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index db9472ed..d56a68f3 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -114,7 +114,7 @@ match what the reader sees is how an instrument stops being believed (R-421). It stopping line that lies. -## Install & onboarding — 14 rows (P3 9, P4 5) +## Install & onboarding — 12 rows (P3 8, P4 4) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -124,23 +124,20 @@ stopping line that lies. | **R-250** | Install & onboarding | P3 | **A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time.** Found 2026-08-07 creating the fifth walk's venue. `POST /configs/new` with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — **fail-closed by design** (`offsite.go:111-121`, *"don't serve a descriptor the controller can't verify"*), with `defaultScanBackoff` = 2+4+8+16+30 ≈ **60 s**, sized by its own comment to *"the observed DNS propagation lag"*. **Measured, both halves:** the first create exhausted the ladder — five `no such host`, then `dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable` (the name had just begun resolving, **AAAA-first, into a pod with no IPv6 route**) — and the hub logged `[ERROR] offsite provision for walk5`. An **identical second POST succeeded** ~70 s later on its own final rung (`shared already provisioned … (subaccount 285351)`). **Total settle ≈ 100 s against a 60 s budget.** The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, `SSH-2.0-OpenSSH_9.6p1`. **Two distinct things are wrong and should not be merged:** (1) the budget is sized against DNS *existence*, but what actually bit is the **AAAA-before-A** window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) **the operator is told nothing actionable** — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — `one_time_secrets` 1, one sub-account, no double-provision) **but that safety is invisible to the person deciding whether pressing again will double-charge them.** **Severity LOW-MEDIUM:** self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | **READY** — owner Viktor | — | — | operator | | **R-283** | Install & onboarding | P3 | **After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page.** `customer_claims` for demo-hp still read `claimed_at 2026-07-21 16:29:25`, `generation 2`, `issued_at 2026-08-03` while the freshly provisioned guest — whose `settings.json` is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ **R-282**), and any previously issued code fails with *„Hibás vagy lejárt kód"* — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of **R-214/R-235** (an already-paired box still told to pair itself) | **READY (S) — NEW 2026-08-09** | — | Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page | CC | | **R-306** | Install & onboarding | P3 | **`--preflight-only` says "no state written" and writes state — with an answer that can be wrong.** `_state_put` short-circuits on `DRY_RUN` only (`felhom-host-install.sh:418`), so a preflight-only run creates `/var/lib/felhom-install/state.json`. Observed live: after a run whose banner read `PRE-FLIGHT PASS (mode=byo) — no state written, no install step executed`, the file existed containing `{"completed": [], "dnsmasq_preexisting": "yes"}`. Both the banner and the flag's own comment at line 226 assert the opposite. **The harm is not the file, it is the value**: the runbook recommends preflight-only first, then the same command without the flag, so on a box carrying a Felhom leftover the wrong ownership answer is baked in before the real install begins | **READY (S) — NEW 2026-08-12, RANK 3** | R-300, R-305 | Either make `_state_put` a no-op under `PREFLIGHT_ONLY` (and record ownership at install instead), or correct both claims. A comment asserting an invariant needs a test pinning it | CC | -| **R-317** | Install & onboarding | P3 | **The agent decides whether to install dnsmasq by stat-ing a file the OTHER package owns.** `EnsureDnsmasq` (`felhom-agent/internal/lanresolver/lanresolver.go:105`) does `os.Stat("/usr/sbin/dnsmasq")` and skips the apt install when it exists — but that path is shipped by **`dnsmasq-base`**, while the systemd unit comes from **`dnsmasq`** (confirmed on the box: `dpkg -S /usr/sbin/dnsmasq` → `dnsmasq-base`; `dpkg -S /usr/lib/systemd/system/dnsmasq.service` → `dnsmasq`). So on any host carrying `dnsmasq-base` without `dnsmasq`, the agent skips the install and then runs `systemctl enable --now dnsmasq` against a unit that is not there: the resolver never comes up and the failure is a retried WARN in the journal rather than anything a customer or the install sees. **Pre-existing, NOT introduced by R-316** — but R-316 makes the shape reachable, because a host whose `dnsmasq-base` pre-dated Felhom now keeps it while `dnsmasq` is removed. R-316's uninstall says so explicitly instead of leaving it to be found from a silent resolver. **Ranked 2 (costs time), not 1:** the box installs fine, only LAN name resolution is missing | **READY (S) — NEW 2026-08-13** | R-316 | Probe what is actually needed — the unit or the `dnsmasq` package — rather than a path a sibling package owns. One-line change in the agent; deliberately NOT made here to keep this session to one repo | CC | | **R-516** | Install & onboarding | P3 | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** | — | — | CC | | **R-554** | Install & onboarding | P3 | **[P3-LOW] Delete the first-boot setup wizard — obsolete by design, still reachable.** OPERATOR DECISION 2026-09-17 (localisation starter, decision 4: „out of scope, obsolete"). `02-controller-module-map.md` L56 calls `internal/setup/` obsolete; `cmd/controller/main.go` L322 still enters it when `setup.NeedsSetup(cfg)` — `customer.id` empty after bootstrap ingestion, or a `.needs-setup` marker (`internal/setup/setup.go` L17-25). Ingestion leaves `customer.id` empty on a missing/invalid `bootstrap.json`, a failed hub pull, or a failed merge/write/reload (`internal/bootstrap/bootstrap.go` L109-163) — so a box whose first boot cannot reach the hub shows a household an 8-page wizard (95 Hungarian strings, its own template set and CSRF). **Fix shape:** decide what such a box shows instead (a single „cannot reach Felhom yet, retrying" page — needs no decision beyond copy), then delete `internal/setup/` and `runSetupMode`; red-proof that a failed ingestion renders the waiting page, not a 404. **Check first** whether any drill/golden path still relies on `.needs-setup`. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-310** | Install & onboarding | P4 | **Two small edges on the installer, neither costing more than a moment.** (1) The R-297 operator-named refusal states the vouched version twice in consecutive sentences (*"…but the vouched golden is 0.213.0. The vouched golden is 0.213.0."*). (2) `--uninstall` reads its typed vmid confirmation from `/dev/tty` and `--force` deliberately does **not** bypass it, so teardown cannot be scripted without a pty — correct for an irreversible destroy, but undocumented; it surfaces as `line 891: /dev/tty: No such device or address` and an rc=1 that looks like a failure rather than a refusal to proceed unattended | **READY (S) — NEW 2026-08-12, RANK 4** | R-297 | Drop the duplicated sentence; add one runbook line naming the pty requirement | CC | | **R-494** | Install & onboarding | P4 | **NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (`architecture/01-topology-and-trust.md`).** *Original finding, kept:* **[P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead.** MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention **I1**): the claim mail points at `https://felhom.drill0242.felhom.eu`; that name has **no A and no AAAA** record (`dig @1.1.1.1`, control `felhom.enkisfelhom.hu` resolves); the hub has **no tunnel- or DNS-creation code** (`hub/internal/cloudflare/` holds only geo-rule removal; `cf_tunnel_token` is a pasted, optional form field, `configs.go:1478`) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered `google.com` but not the dashboard name at 13:27:39Z; the agent applied the record at **13:27:44Z** (`lanresolver: applied split-horizon record … ip=192.168.0.158`, 3 m 46 s after the controller started), so the box CAN answer the name — **but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that**; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (`curl --resolve …:443:192.168.0.158`). **A volunteer could not have done that.** **What it needs:** an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: operator-side automation; the operator creates the tunnel by hand per the day-0 runbook.** | — | — | CC | | **R-504** | Install & onboarding | P4 | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** **Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic.** | — | — | operator | -| **R-725** | Install & onboarding | P4 | **[P3-LOW] Small copy slips on the first-hour path.** MEASURED 2026-09-29: the self-bind PAGE says the passphrase was received „a beállításkor" while the mail and console say „a Felhom üzemeltetőjétől" (R-497 unified the mail and console, not the page); the recovery-code wizard addresses the household formally („Írja fel", „adja meg") while every other screen says „te"; the console's linked banner ends „a doboz össze van kötve. V" (a stray glyph); a gated app answers a phone app's API call with English JSON „this app is waiting for its first setup" (the browser gets the Hungarian gate page). **FIXED 2026-09-30:** the recovery wizard speaks „te" (controller v0.283.0; formal ceiling 18 → 17); the bind page says „a Felhom üzemeltetőjétől kaptál" (hub v0.126.0). **NARROWED — remaining:** the console's stray „V" (the installer/agent's banner, not these repos' text); the gate's English JSON to a phone app (the app shows its own error; left, deliberately); and the expired bind page still says „kérj újat az ügyfélszolgálattól" ABOVE the new „Új linket kérek" button (hub copy, next hub release). | **NARROWED — three small copy items; owner: CC** **Re-ranked 2026-10-03: P3→P4: three small copy slips left; nothing blocks the household.** | — | — | CC | | **R-881** | Install & onboarding | P4 | **The installer's uninstall does not remove `/usr/local/sbin/felhom-priv-apply`** (added to the bundle by agent v0.146.1, R-861), and its comment still says the guest hook is agent-installed at runtime (`scripts/felhom-host-install.sh` ~line 1679) — since v0.146.1 the hook is a bundle file. Found 2026-10-05 while checking the installer against the new bundle. | READY | — | Add the file to the uninstall list and fix the comment at the next installer tag | CC | -## Apps & catalog — 25 rows (P3 12, P4 13) +## Apps & catalog — 23 rows (P3 11, P4 12) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-498** | Apps & catalog | P3 | **[P3-LOW] The „Első lépések" on 52 of 53 app pages tell a customer to open a literal `wiki.DOMAIN` — the placeholder is never filled in.** MEASURED 2026-09-14 on a fresh 0.242.0 box (drill step 5): BookStack's app page renders „Nyisd meg a **wiki.DOMAIN** címet a böngészőben", PrivateBin's „Nyisd meg a **paste.DOMAIN** címet". `grep -rl '\.DOMAIN c[íi]m' app-catalog-felhom.eu/templates/*/.felhom.yml` → **52 of 53** templates carry it in `first_steps`; the controller renders the string as text (`internal/stacks/metadata.go`). A stranger reading their first instruction meets a word that is not an address. **Fix shape:** substitute the stack's real `SUBDOMAIN.DOMAIN` at render time (one place in the controller), with a render test per template that fails on a literal `DOMAIN`. | **READY — rank P3-LOW; owner: CC** | — | — | CC | | **R-562** | Apps & catalog | P3 | **[P3-LOW] Dates and sizes are not formatted for any locale — and the Hungarian pages disagree with themselves.** FOUND 2026-09-17 by the i18n inventory §2.8: the two template date layouts differ (`2006. 01. 02. 15:04` Hungarian vs `2006-01-02 15:04` ISO); 10 layout literals in `internal/web` Go and 25 elsewhere pick formats ad hoc; sizes print a decimal POINT (`%.1f GB`, 4 helpers) where Hungarian uses a comma; `timeAgo`/`nextRunLabel`/`pruneLabel` produce Hungarian words outside the three converted pages. Not changed by v0.247.0 (Hungarian bytes are frozen by the parity rule). **Fix shape:** one date and one size formatter per language in `internal/i18n`, the Hungarian output deliberately changed in ONE reviewed release with the parity fixtures re-captured for that release only and the change named in its CHANGELOG. Needs an operator word on the Hungarian format (comma, date style). | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-575** | Apps & catalog | P3 | **[P3-LOW] The soft memory-overcommit warning is returned as a Hungarian STRING, so it renders Hungarian on an English page.** FOUND 2026-09-18 by localisation slice 2 release B (R-557, controller v0.253.0): `memoryVerdict` now returns its REFUSAL as an error carrying a key (`util.MsgErrorf(ErrNotEnoughMemory, …)`), but its WARNING is a plain string with no error to carry one, and no language is known where it is built — so it goes through `msgHU` and is always Hungarian. The copy is in the bundle (a translator sees it; `i18n_go_parity.py` pins it), only the render is fixed to `hu`. The named helper and this consequence are stated in `internal/stacks/deploy_errors.go`. **Fix shape:** the caller carries the key the way `Alert` and `UpdateRefusal` now do — `memoryVerdict` returns `(refusal error, warningKey string, warningArgs []any)` and the deploy answer renders it — not a helper guessing a language it cannot know. | **READY - rank P3-LOW; owner: CC** | — | — | CC | -| **R-593** | Apps & catalog | P3 | **[P3-LOW] papra's session-signing key is described as „the app's subdomain".** FOUND 2026-09-20 translating the catalog (R-560 slice 5, batch 1). `templates/papra/.felhom.yml` `deploy_fields[AUTH_SECRET].description` reads **„Az alkalmazás aldomainje"** — the sentence that belongs on `SUBDOMAIN`, on a field that signs sessions. `SUBDOMAIN` itself has NO description at all in that file, so this is a copy-paste that landed one field too low and took the original with it. A customer opening papra's install page reads a wrong explanation under a key they must not regenerate. **A localisation release may not change Hungarian bytes (10 §1)**, so it was not fixed there; and translating a wrong sentence faithfully would have shipped the error in a second language, so **that one field was left untranslated** — it falls back to the Hungarian exactly as today, and it is the ONE string keeping papra at 13/14 and the catalog ceiling off zero. **Fix shape:** move the sentence to `SUBDOMAIN` and give `AUTH_SECRET` its own („A munkamenetek aláírásához használt kulcs" or similar), re-capture that app's two entries in `scripts/copy_freeze/hu.json` in the same commit with the reason, then add the English. Owned by R-516 as a Hungarian-words change. | **READY - rank P3-LOW; owner: CC (catalog)** | — | — | CC | | **R-612** | Apps & catalog | P3 | **[P1-HIGH] `wishlist` cannot be signed up to on a fresh Felhom install, the deploy reports SUCCESS, and the error the customer sees is a LIE.** MEASURED 2026-09-21 on guest 9202 while seeding for the power-cut drill. The image's first-boot `pnpm prisma db seed` is **`Killed` — OOM at the catalog's `mem_limit: 128M`**. Without it the `Role` and `Group` rows are absent, so **every** signup fails. **The message the user is shown is `User with username or email already exists`** while the container log says the real cause: `FOREIGN KEY constraint violated`. A household would conclude the account already exists and try to recover a password that was never created. **The controller reports the app running and HEALTHY throughout, and the deploy reported successful** — so nothing on the box says anything is wrong. Repaired on the scratch guest only, to unblock seeding: memory raised to 512 M, the image's own seed re-run, memory put back to 128 M. **The catalog was NOT changed** — the fix is a memory-limit question for the catalog and is deliberately left to a session that can measure the real ceiling rather than guess it. **Needs: the actual peak RSS of that seed, then a `mem_limit` that clears it, plus a check that the seed's failure is not silent.** **— FIXED 2026-09-23 night (catalog `a5a729a`): 512M.** Measured on the bench: the seed peaks at 312–345M and was OOM-killed at 128M on every run (kernel `oom_kill` 3); running, the app sits at ~115M (90 % of the old limit alone). On 9202 at 128M the kill came AFTER the Role/Group rows this time, so the sign-up lie did NOT reproduce — the kill is timing-dependent, the fault is memory (the brief's claim held). At 512M the seed completes (`The seed command has been executed`), and wishlist then moved v0.66.0 → v0.67.1 with its test record. **Still open:** a failed first-boot seed is invisible to the box — nothing reads the seed's exit. | **READY — P3, narrowed to making a failed seed visible; owner: CC (catalog)** **Re-ranked 2026-10-03: P1->P3: the memory fix shipped (catalog a5a729a); only the silent-seed detection remains.** | — | — | CC | | **R-613** | Apps & catalog | P3 | **[P2-MEDIUM] `uptime-kuma` parks on its setup wizard with no login and no monitors, and the box tells the household it is HEALTHY.** MEASURED 2026-09-21 on guest 9202. On first boot uptime-kuma 2.4.0 sits at `SETUP-DATABASE` (`Waiting for user action...`) and its main socket.io server never starts. **The controller's `http :3001` probe sees the wizard's 302 and records the app as running and healthy.** So a monitoring app that cannot be logged into, and is monitoring nothing, is presented to the customer as fine. Passed through its own front door for the drill with `POST /setup-database {"dbConfig":{"type":"sqlite"}}`. **This is the health-check class the catalog skill already warns about — a probe that proves the PORT answers, not that the APP works** — and it is worth a row because the failure direction is a false GREEN, which no alarm will ever catch. **Needs: a healthcheck for this template that fails while the wizard is up** (the catalog `REUSE.md` maps the families), and a sweep for other templates whose probe would pass on a setup wizard. **— UPDATE NIGHT 2026-09-21:** the update night could not seed `uptime-kuma` for the same reason and left it out rather than faking it. **— FIXED 2026-09-23 night for uptime-kuma (catalog `a5a729a`):** `UPTIME_KUMA_DB_TYPE=sqlite` — the database choice is made, the wizard never appears, the real server starts (`/api/entry-page` → `entryPage`, `/metrics` 401), the database lands in `/app/data` (the backed-up volume). Red-proofed on 9202 through the product: before, the box read `running` over `setup-database`; after, the real server. (A first "after" run used a stale template — R-607's lag — and is kept.) **Still open:** the sweep for other templates whose probe passes on a setup wizard. | **READY — P3, narrowed to the sweep; owner: CC (catalog)** **Re-ranked 2026-10-03: P2->P3: uptime-kuma fixed; only a sweep of other templates remains.** | — | — | CC | | **R-676** | Apps & catalog | P3 | **[P3-LOW] Watch: immich's first start restarted 12 times — decision 28's crash-loop stop (6 in 10 min) would stop it.** From the 2026-09-17 chaos night (DB connection dropped during the first-start geocoding import on a 6 GB guest; it did not recover that night). No healthy app in any drill evidence restarts on a first start (1831 samples, 40 live containers), so the threshold stands; this row exists so the first immich install under v0.269.x is watched. `audits/night-2026-09-24/A3/40-first-start-restarts.txt` **2026-09-25 night (read from source, v0.271.0): a DEPLOY's first start is NOT covered by decision 28's suppression** — `Deploying` clears when `compose up -d` returns (`deploy.go` "Clear deploying flag"), and `ObserveUnhealthy` then samples the app; an automatic update's step, verify and undo ARE covered (`Updating`, pinned by `TestD28_NoCrashLoopStopDuringAnAutomaticStep`). So a first start that restarts ≥ 6 times in 10 min is stopped — which R-676 already accepts for a broken first start; a healthy slow first start would be stopped too. **-- 2026-09-30: the first-start restarts are explained.** immich's first-start geodata import OOM-kills its database at 512M on a guest with no swap (R-732, measured: 61–104 kills); the 2026-09-17 chaos-night case (DB connection dropped during the import on a 6 GB guest) fits it. Fixed in the catalog (`56c4888`, 768M). The watch itself (decision 28 on a DEPLOY's first start) is unchanged. | **OPEN — P3; owner: CC (watch)** | — | — | CC | @@ -156,7 +153,6 @@ stopping line that lies. | **R-644** | Apps & catalog | P4 | **[P3-LOW] `gokapi` on scratch guest 9202 is crash-looping — 329 restarts by 2026-09-23 07:51 UTC, *password does not appear to be a SHA-1 hash* — and the controller still lists it deployed.** OBSERVED at the start of the 2026-09-23 session, not caused by it. The twenty-eight walk's teardown (2026-09-22) removed a `gokapi` container left by R-633 by name; a `gokapi` is running again, recorded `deployed: true`. Not investigated (scope). Likely the R-633/R-634 shape — a restore-then-remove race leaving a record — and a scratch-box fact, not a customer one; filed so the next drill does not read it as its own. | **OPEN — P3; owner: CC; investigate before the next drill on 9202** **Re-ranked 2026-10-03: P3→P4: seen only on a scratch test guest; the row itself calls it not a customer fact.** | — | — | CC | | **R-707** | Apps & catalog | P4 | **[P2] 37 apps still start with a login a stranger can take (`09` §3 decision 45).** Audit of all 53 apps: `app-catalog-felhom.eu/FIRST-ADMIN.md` (class, fix route, status, measured or read). Open: **3 hard-coded defaults** — calibre-web (`admin / admin123`, measured working on demo-hp and 9202; its own `cps.py -s` route needs a generated password WITH a special character — our generator is letters+digits, a controller change), mealie (`changeme@example.com / MyPassword`), wger (`admin / adminadmin`); **34 open first-run screens** (the first visitor creates the admin: actualbudget, adventurelog, audiobookshelf, calcom, docmost, emby, ghost, gitea, gramps-web, home-assistant, homebox, immich, jellyfin, komga, n8n, navidrome, opengist, outline, papra, plant-it, radarr, rallly, recipe-importer, romm, seerr, sonarr, sparkyfitness, tandoor, termix, uptime-kuma, vikunja, wanderer, wishlist, zipline). **Stale notes:** romm's `default_creds` `admin / admin` answers 401 on demo-hp (like a wrong password) — the page now warns with a login that does not exist; zipline's looks stale too. **Measured on demo-hp 2026-09-28 (read-only):** bookstack's default still logs in on the INSTALLED app (the fix is for new installs; the page now warns). Each fix: route (a) env or (b) the app's own CLI/API via `after_install:`, proven on 9202 with the default failing and the generated password working; route (c) a page sentence. Several sessions (operator, 2026-09-28). **2026-09-29 (controller v0.280.0, catalog `d0e7e2e`):** every class-3 app fixed — mealie, wger, calibre-web by `after_install` (calibre-web with the new `password:24:special`), proven on 9202 fresh installs (`audits/login-gate-2026-09-29/D/`); the setup gate (decision 46, spike PASSED) built and live on immich, n8n, audiobookshelf (probes measured) and uptime-kuma (button) (`…/C/`); romm's and zipline's stale notes removed. **Left: 30 class-4 apps** — gate each (probe measured on 9202 where one exists — 11 upstream candidates listed in `…/B/B-VERDICT.md` §3; the button otherwise). **2026-09-29 afternoon (controller v0.281.0, catalog `6faf432`):** 28 more class-4 apps gated — 32 of 34 — each proven on 9202 (`audits/gate-rollout-2026-09-29/`B): stranger → gate page / 401, household reached the first-setup screen, the gate opened (9 by a measured probe, the rest by the press), the app answered after. seerr, outline, rallly: gated, their opening needs a media server / e-mail (not proven). **Left:** wanderer (R-714); plant-it is not installable. | **NARROWED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — seerr, outline and rallly are gated, but the gate OPENING is not proven (needs a media server / e-mail)) — **CLOSED — 2026-09-29 (the rest → R-714)** | — | — | CC | | **R-718** | Apps & catalog | P4 | **[P3-LOW] "Close sign-up now" restarts an app with its own switch, and the card does not say so.** MEASURED 2026-09-29 on demo-hp: pressing it on adventurelog recreated its backend (~30 s, one 500 on its login page). The window's card says the app restarts; the close card does not. **Fix direction:** the close card and the gate-open moment say "the app restarts once" where `after_setup.env` exists. **ALSO MEASURED 2026-09-29 (new-household drill):** the gate-open press on a fresh vikunja restarted it for its own switch — the front door answered 404 for ~2 s and nothing said so. | **OPEN — P3; owner: CC** **Re-ranked 2026-10-03: P3→P4: a short unannounced restart; copy polish only.** | — | — | CC | -| **R-760** | Apps & catalog | P4 | **[P3-LOW] vikunja's compose has no healthcheck, and nothing says why.** FOUND 2026-10-01 by `scripts/onboarding_gaps.py` (row 4.1): two services in the catalog have no compose `healthcheck:` — `adventurelog-frontend` (deliberate, a comment cites R-655: the image brings its own) and `vikunja` (no comment). Not measured: whether the vikunja image declares its own `HEALTHCHECK`. The controller's probe still runs (`healthcheck.checks` in `.felhom.yml`), so the badge is not blind; Docker's own health state is. **Needs:** read the image's config; either a compose healthcheck of the family the image has (REUSE.md §2), or a comment saying why none. `app-catalog-felhom.eu/onboarding/EXISTING-APPS-GAPS.md` | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: the box's own health probe still runs; only Docker's health state is blind.** | — | — | CC | | **R-764** | Apps & catalog | P4 | **[P3-LOW] wger sends no mail: its mail backend is the console, so a password-reset mail never leaves the box, and the template maps no SMTP.** READ 2026-10-01 inside the app on 9202 (checklist row 7.1): `EMAIL_BACKEND django.core.mail.backends.console.EmailBackend`; the settings read `ENABLE_EMAIL`, `EMAIL_HOST`, `EMAIL_PORT`, `FROM_EMAIL` …; the template carries no `smtp_mapping`. The household's admin can reset another member's password in the app; a member who forgets theirs and asks wger by e-mail gets nothing, and the page does not say so. Not measured: what wger shows after a reset request. **Needs:** an `smtp_mapping` (vaultwarden's shape, a fresh install with mail OFF booting — REUSE.md §2), or the page saying mail is not available. `audits/new-app-checklist-2026-10-01/C/C2-static-reads.txt` | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: wger is hidden and no box runs it.** | — | — | CC | | **R-769** | Apps & catalog | P4 | **[P3-LOW] Pinchflat is not built: upstream is paused (no release in 2026, last push 2025-12-16) and its last release has no image tag.** READ 2026-10-01: ghcr's newest version tag `v2025.6.6`, `latest` amd64 only; runs as root by default; an unanswered 30 GB yt-dlp memory report (#866); third parties call it unmaintained (community-scripts #15968). Forks with images exist (Pinchflat-NGX, MorganKryze). **Needs:** the operator's word on a fork (a new upstream), after the YouTube sentence (R-767). `audits/new-apps-2026-10-01/FIT.md` | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** **Re-ranked 2026-10-03: P3→P4: a new-app idea waiting on a decision.** | — | — | operator | | **R-770** | Apps & catalog | P4 | **[P3-LOW] Invidious — fit check only; the recommendation is not to build it.** READ 2026-10-01: playback needs `invidious-companion` (rolling `latest`, no version tags); PostgreSQL 14 (EOL 2026-11); `registration_enabled: true` by default; upstream: a bot check means „your IP is blocked from YouTube”, a 429 can last 24 h, triggered by „someone on your network” — on our boxes that IP is the household's. One bad period in 2026 (March, ~2 weeks). No report found of a family's other devices being bot-checked (inference). **Needs:** the operator's go / no-go. `audits/new-apps-2026-10-01/FIT.md` | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator** **Re-ranked 2026-10-03: P3→P4: a new-app idea waiting on a decision.** | — | — | operator | @@ -178,7 +174,7 @@ stopping line that lies. | **R-621** | App updates | P4 | **[P2-MEDIUM] A held update DESTROYS the evidence of why it failed: `failAndHold` runs `compose down`, the failing containers are removed, and their output is gone before anyone — household, operator or the next session — can read it.** MEASURED 2026-09-21 on guest 9202 on a REAL upstream edge: `adventurelog v0.12.1 → v0.13.0`. The new backend applied **nine Django migrations successfully** and then never listened; the update held after the full 5-minute health wait. **`Manager.failAndHold` (`stacks/update.go:723`) calls `updateCompose(dir, env, "down")`**, which removes the containers rather than stopping them, and nothing captures their logs first. Within seconds the box's own log recorded `Logs result for adventurelog: 0 bytes returned (empty)` and `docker ps -a` held nothing at all. **What survives is the WHAT and not the WHY:** the controller line `update adventurelog FAILED after the new version was started: not healthy: not healthy within 5m0s (last: state unhealthy)` and the household's sentence, both of which say the app did not come up and neither of which says the migrations ran and the server then failed to bind. **This is R-320 ("evidence off the machine before the teardown") as a PRODUCT behaviour rather than a session habit** — the teardown here is the product's own, it is correct to perform (a half-started new version must not keep running), and it happens before anyone can look. **Why it matters beyond a drill:** the hold sentence sends the household to a restore, and after the restore the only remaining question is *should I press Update again?* — which nobody can answer, because the one artefact that would say so no longer exists. It also makes every future held update unreportable to an upstream project. **Fix shape:** capture `compose logs --no-color --tail N` into the stack directory (beside `applied-compose.yml`, which already travels with the stack) IMMEDIATELY before the `down`, and surface it on the app page's hold panel or at least through the existing `/api/stacks//logs` fallback. Bounded size, written once per hold. A test that holds an app and asserts the captured file is non-empty — it fails today. Evidence: `audits/update-night-2026-09-21/apps/adventurelog/why-it-failed.txt`, `state-after-hold.txt`. **FIXED in v0.262.0.** `failAndHold` now writes each service's log into `/hold-logs//compose-logs.txt` (`compose logs --no-color --tail 400`) **before** the `down` that destroys them, best-effort by design: a hold must never fail because its evidence could not be written. This is what `adventurelog` cost — nine migrations ran, the app never bound its port, and the only log that could have said why was gone before anyone looked. | **VERIFY** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the fix shape also said to show the captured hold log on the app page's hold panel; the row does not say that was done) — **CLOSED 2026-09-22 — v0.262.0: the hold keeps the app's own log before stopping it** | — | — | CC | | **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator | -## Backup & restore — 48 rows (P2 8, P3 19, P4 21) +## Backup & restore — 47 rows (P2 8, P3 19, P4 20) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -213,7 +209,6 @@ stopping line that lies. | **R-91** | Backup & restore | P4 | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | | **R-99** | Backup & restore | P4 | Server-side prune **never removes** a phantom snapshot. Confirmed it does NOT count them toward `keep-last` (dry-run kept 2 real + the phantom) so there is **no retention/data-loss bug** — but one accumulates per aborted upload, forever | READY (S) | — | Decide a cleanup path. Deletion on a **customer** datastore is a separate ruling — detection shipped, removal deliberately not automated | CC | | **R-104** | Backup & restore | P4 | **An interrupted offsite run leaves an exclusive restic lock the existing self-heal cannot reach.** `resticStep` has `unlock --remove-all` (`internal/backup/offbox.go:634-648`) but `ensureOffboxRepo`'s probe fails first, `classifyResticProbe` (`:77-93`) has no lock case → `"other"` → fail-fast; `ClassifyOffsiteFailure` likewise, so the operator is told *„A távoli mentés ismeretlen okból nem sikerült"* for a precisely-known, self-healable condition **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-28, size S, roadmap state `READY — 2026-07-28`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. ****PARTLY STALE, checked against live source 2026-08-22 — migrated as written per the task rule, with the staleness named rather than edited away.** The self-heal this row calls unreachable was built: `resticStep` escalates to `unlock --remove-all` and retries once (`internal/backup/offbox.go:763-768`), and `unlockStale` runs before every off-site run and restore (`:1274`, `offbox_restore.go:261`). Its premise that the probe fails first is also doubtful: the probe is `restic cat config`, a read that takes no lock. **What REMAINS true:** `ClassifyOffsiteFailure` (`offbox.go:179-193`) still has no lock case, so if a lock ever did survive both layers the customer would still be told an unknown reason. Re-rank on that basis, not on the original text.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | **Was C9-F3.** Reachable by any interruption — container restart, OOM, network drop, host reboot mid-backup. The tier stays dead until a human runs `restic unlock --remove-all`. Flips: the offsite row in map §C; `07` §8 row 15 | CC | -| **R-124** | Backup & restore | P4 | **The recipe spells PBS's root namespace `"root"`, but the PBS API spells it `""`** and no namespace is literally named `root` — an operator pasting the field into `pct restore --ns root` gets a failure | READY (XS) **Operator ruling 2026-10-05 18:23: kept OPEN and fixed in burn-down round 2** — a recovery step that fails during a real recovery is the worst time to find it. | — | Pre-existing wire convention (`ToHub` has normalised empty→`"root"` since slice 6), deliberately NOT changed under R-106 so the field's meaning did not shift mid-fix. Documented at `hub.PBSRootNamespace`. Affects only a box with no `namespace` line — **no real customer today**, all three are per-customer. Fix = emit `""` + rely on `namespace_state`, or emit a `--ns`-ready form | CC | | **R-164** | Backup & restore | P4 | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | | **R-213** | Backup & restore | P4 | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC | | **R-246** | Backup & restore | P4 | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor **Folded R-248 2026-10-03** (the same ruling: give `stale_at` a visible, evidence-bearing setter, or retire it). | — | — | operator | @@ -231,11 +226,10 @@ stopping line that lies. | **R-832** | Backup & restore | P4 | **ep0's copy in a place outside both Hetzner and the operator's home (roadmap).** Today DooPlex (the operator's home) holds it (decision 71). A Hetzner Storage Box would share a provider with ep0 and with every household's file backups, and cannot run PBS, so the copy could not be verified or restored from directly. | **DEFERRED — later, if the product grows** | — | — | operator | | **R-878** | Backup & restore | P4 | **A catch-up (R-871) runs the database-dump leg in the DAY, and that leg stops an app with a volume for its copy — the household may notice the stop, and a large volume makes it longer.** MEASURED 2026-10-05 on demo-felhom: the catch-up at 08:25:02 stopped opengist, copied 182.5 KB, started it again — about 1 s, then a few seconds of `health: starting`; the night does exactly the same, unseen. Nothing measured for a large volume. Fix direction (if it matters): skip the volume copy of a running app in a DAYTIME catch-up and leave it to the next night, or warn. `audits/catchup-2026-10-05/partA/live-demo-felhom.txt` | **READY — owner: CC** | — | measure a large volume first | CC | -## Storage & devices — 10 rows (P3 6, P4 4) +## Storage & devices — 9 rows (P3 5, P4 4) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-118** | Storage & devices | P3 | **An absent drive's union row advertises the ROOT filesystem's capacity as its own.** In the absent-state payload the registry-union row reports `total_bytes: 49675956224 / used_bytes: 4584579072` — **byte-identical to the `local` row** (`durable_id: path:/var/lib/vz`, i.e. `pve-root`) in the same response. The real drive is **4 GB** | **READY (XS) — NEW 2026-07-30** | — | Cause: `statfsCapacity(d.MountPath)` (`disks.go:335-338`) statfs's `/mnt/cel`, which with the device gone is a **bare directory on the root filesystem**. `observe.go:176-183`'s comment warns about exactly this trap and guards the Observe path ("*an unmounted removable dir-storage's mountpoint reverts to a bare directory on root … catastrophic DR mis-id*"); **the union path has no equivalent guard.** **Not a DR mis-id** — `durable_id` on that row is still the correct `uuid:…`, so re-attach identity is safe. It is a **false capacity** reaching every consumer of `total_bytes`/`used_fraction` (fill monitors, storage cards): a detached 4 GB drive advertises 46 GiB at 9.2 % used. Same class as `role.go:180-181` — an absent drive's fields decaying to the root filesystem's. Evidence: `audits/DIAG-r116-disks-payload-2026-07-30.md` §12 | CC | | **R-298** | Storage & devices | P3 | **The `/storage` page's unregistered list is filtered by `role==='user-data'`, so a drive that is also the backup target can never be registered from it.** `storage.html:363` routes anything not `user-data` into the read-only protected group with NO actions. On the rebuilt `demo-hp` the NVMe is deliberately BOTH the user-data drive and the `felhom-backup` target (`/etc/pve/storage.cfg`: `dir: felhom-backup` → `/mnt/nvme-1tb`), so it renders locked. **This is the SECOND reason that page was empty** during the reinstall rehearsal, independent of R-280's candidate-source defect, and R-280's fix does not touch it — attaching is non-destructive, so the format-wizard protection is the wrong gate for a REGISTER action | **READY (S) — NEW 2026-08-10** | R-280 | Split the role gate: `user-data` keeps destructive actions; any mounted role may be REGISTERED | CC | | **R-330** | Storage & devices | P3 | **Disk health Phase 2 — the three SMART attributes the wire does not carry.** The failing drive's most telling counter was **187 `Reported_Uncorrect`**, sitting at normalized **1** against threshold **0** with a raw count of **1001** — one point from failing and structurally unable to get there. Also wanted: **199 `UDMA_CRC_Error_Count`** (cabling) and **188 `Command_Timeout`**. None are on the agent→controller wire today, so v0.215.0's ladder could not use them. Phase 2 also persists periodic SMART **samples** (the right home is `metrics.MetricsStore`, NOT the Phase-1 state file, which is one record per disk and must stay that way). **This is a declared WIRE change, so under the G-1 gate the hub must model the new fields in the SAME session** — that is precisely why it was kept out of Phase 1, where it would have turned a one-word severity fix into a three-repo change | **READY (M) — NEW 2026-08-14** | R-328 (closed) | Add 187/199/188 to the agent's `SmartSummary` + hub model in one session; then persist samples | CC | | **R-332** | Storage & devices | P3 | **The new Hiba-from-counters path has never fired on real hardware.** v0.215.0's whole point is a verdict the product could not previously reach, and it is proven only against the committed fixture's values in unit tests (12 scenario groups, 11 of 12 red-proofs failing as required). The live validation on demo-hp proved the **negative** — three healthy disks still read Rendben across the deploy, no false alert — and the **severity wire** end to end, but no live disk has actually reached Hiba. **This is the honest gap and it must not be closed by pointing at the fixture tests**: the drive that produced the fixture is in DooPlex, which is Tier 2 and never a drill target, and the demo boxes are all-flash and healthy | **WATCHING — NEW 2026-08-14, NARROWED same day.** One item originally in this gap is now PROVEN LIVE: the **persisted state surviving a controller restart**. The v0.215.0→v0.216.0 redeploy destroyed and rebuilt the container, and the new one read back a `changed_at` written by the PREVIOUS version (`2026-08-14T07:23:14.640216851Z`, still intact at 09:31:35Z) instead of re-baselining — Scenario L on real hardware, not just the production-path unit test. **What remains unproven is the verdict itself, plus the stronger restart half: an already-ALERTED disk not re-alerting** | a real degrading disk, or an injection harness | **Closing condition:** a live disk reaching Hiba from counters, OR a deliberate injection through the REAL pipeline (agent `/disks` → controller check → hub event), not a hand-set verdict | CC | @@ -246,25 +240,22 @@ stopping line that lies. | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | | **R-568** | Storage & devices | P4 | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: `/dashboard` fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (`audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt`). `diskHealthRows` (`disk_health.go` L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. **Fix shape:** sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic.** | — | — | CC | -## Security & access — 28 rows (P2 2, P3 23, P4 3) +## Security & access — 24 rows (P2 2, P3 20, P4 2) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-777** | Security & access | P2 | **[P2-MEDIUM] Emby and Jellyfin treat every internet visitor as being on the LAN — users with "remote access" off can sign in from the internet, IP filters and remote limits are skipped.** READ in source (`audits/visitors-2026-10-01/A/sweep/sweep-1.md`, `sweep-2.md`), not measured live: Jellyfin with `KnownProxies` empty uses the TCP peer (traefik, private) → "LAN"; Emby reads the leftmost XFF (its chain is now removed by the R-753 reset, so it sees cloudflared's private address → "LAN", as before). True before R-753 too; R-753 neither caused nor fixed it. **Needs:** measure on 9202 (a user with remote access off, through the simulated tunnel); Jellyfin: `KnownProxies` `172.16.0.0/12` in `network.xml` (no env — an `after_install` or a seed file); Emby: no setting fixes it (its `LocalNetworkSubnets` still counts private ranges) — a page sentence or a decision. | **READY — rank P2-MEDIUM; owner: CC (measure), operator (Emby route)** | — | — | CC + operator | | **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC | | **R-126** | Security & access | P3 | **A `.fab` bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | READY (S) | — | Split out of R-108, which closed without it: this is an explicit customer-chosen **export destination**, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (`07` §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share | CC | -| **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | +| **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | | **R-136** | Security & access | P3 | **Rename `hub_session` → `__Host-hub_session`** — makes cookie tossing structurally impossible | READY (XS, one line) | — | Verified on the live production response that all three prefix preconditions already hold: `Path=/`, `Secure`, no `Domain`. **Caveat for the ticket:** browsers reject a `__Host-` cookie without `Secure`, and `isSecure` is conditional on `r.TLS`/`X-Forwarded-Proto`, so plain-HTTP *browser* access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: `r.Cookie` returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 | CC | | **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | | **R-138** | Security & access | P3 | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) | READY (S) | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | | **R-255** | Security & access | P3 | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator | -| **R-269** | Security & access | P3 | **A rotated-out per-guest local-API token still authorises, and the test that appears to pin the opposite passes only because of its lookup ORDER.** `localapi.TokenStore.Mint` documents *"last-write wins — any previous token for this guest is revoked"*. Across processes that is FALSE until something unrelated forces a reload: the long-lived agent serves `Lookup` from an in-memory index and re-reads the store **only on a miss** (the B3 reload-on-miss optimisation, `tokenstore.go`), so a superseded token is a direct map **hit** and returns `(vmid, true)`. **Red-proved twice.** (a) A unit probe — `TestTokenStore_ReloadOnMiss_RemintCoherence` with the two lookups swapped, i.e. present the rotated-out token FIRST — **fails**; the shipped test passes only because it looks up the NEW token first, and that miss is what evicts the old hash. (b) On hardware, 2026-08-09: after the on-disk rotation the old token returned **HTTP 200**, then 401 only once a new-token lookup had forced the reload, and reliably 401 after `systemctl restart felhom-agent`. **This is the `CLAUDE.md` invariant-comment case exactly** — the comment reads as settled and the test that looks like its pin is order-dependent. **Fix options:** pin the reversed order with a test, or make eviction not depend on an unrelated miss. Until then, **an operator rotating a leaked token MUST restart the agent** — the runbook step is not optional | **READY (S) — NEW 2026-08-09** | — | Found by doing R-268's rotation rather than reading about it | CC | | **R-270** | Security & access | P3 | **R-268's stated rotation recipe is incomplete: the controller never re-reads `bootstrap.json`'s `local_api`, so a rotation leaves the agent channel dead across restarts.** `bootstrap.ensureLocalAPI` returns early when `cfg.LocalAPI.Endpoint != ""` — by design it FILLS an absent block and never refreshes a present one — so the token the controller uses lives in its own `controller.yaml`, not in the mount. Proved live 2026-08-09: two controller restarts after a correct `bootstrap.json` rotation, still `HTTP 401`; the channel came up only once `local_api.token` was written into `controller.yaml`. The neighbouring `DetectEndpointDrift` compares the ENDPOINT and deliberately does not compare the token (*"a token mismatch is a different failure"*), so this shape is knowingly unmodelled. Parent question — which file is authoritative — is **R-78** | **READY (S) — NEW 2026-08-09** | — | Either teach the drift detector the token, or make the rotation path write both files. Correct the R-268 row's recipe either way | CC | | **R-275** | Security & access | P3 | **`--uninstall` leaves five 0600 `agent.json.*` credential backups, and the reinstall hands them to the new service account.** `/etc/felhom-agent/` survives with `agent.json.{campaign8-before,campaign9-before,campaign9-prev,pre-e-target-move,pre-prunegate.bak}`, each carrying a 64-char `hub.api_key` and a 59-char `proxmox.token`. The teardown claims to remove *"config (+ its .bak backups)"* and `scripts/CHANGELOG` F1 records *"uninstall now purges the agent config's `.bak*` siblings (one held a live hub api_key)"* — **that fix does not match the filenames in use, and it misses `agent.json.pre-prunegate.bak`, a file that literally ends in `.bak`.** **Exposure assessed, not assumed:** these are SUPERSEDED — the orphaned key hashes to `a5d2222a…`, the hub's current demo-hp key to `8c59d1b6…`, and the Proxmox token was deleted by the same uninstall. **But the reinstall recreates `felhom-agent` at uid 999, the same uid the deleted account had**, so three of the backups become the new account's files — verified readable as `felhom-agent`. A fresh install's service account inherits read access to the prior install's credentials; superseded today, live if the backups were recent (R-179's precedent). **Also left, undeclared:** `/etc/felhom/{.bootstrap-done,appliance-pairing-code}`, `felhom-bootstrap.service` + `/usr/local/sbin/felhom-bootstrap.sh`, the `vmbr9` stanza in `/etc/network/interfaces`, and `/etc/sudoers.d/felhom-agent.bak-pre-e2a` (21 KB — **INERT: sudo skips dotted filenames, verified with `sudo -l -U felhom-agent`; `visudo -c -f` parsing it OK is NOT evidence sudo loads it**) | **READY (S) — NEW 2026-08-09** | — | Purge by directory, not by glob; and do not let a new service account reuse a uid that owns old secrets | CC | | **R-276** | Security & access | P3 | **RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing.** After `--uninstall` on demo-hp, `wg-quick@wg-felhom` is **enabled and active**, `/etc/wireguard/wg-felhom.conf` present, handshake to `167.233.158.164:443` **52 s old**, counters 5.86 GiB in / 2.48 GiB sent. It appears in **neither** the WIPED nor the KEPT list, though the BYO install disclosure names it prominently on the way in (*"an OUTBOUND WireGuard tunnel to the Felhom hub"*). A host told to leave Felhom retains a live network path into Felhom infrastructure, its hub-side peer registration intact, and nobody is told | **READY (S) — NEW 2026-08-09** | — | Tear the tunnel down and deregister the peer, or list it under KEPT with the reason and the removal command | CC | | **R-338** | Security & access | P3 | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | -| **R-350** | Security & access | P3 | **SECURITY — the hub operator password was printed in cleartext into a session transcript by CC, 2026-08-20. Rotation recommended.** **What happened:** vouching the artifact manifest used `curl -w '%{redirect_url}'` for confirmation. The hub answers `POST /configuration/artifacts` with a **303**, and curl renders the redirect target **with the basic-auth credentials re-attached** — so the URL it printed contained `http://:@10.43.52.34:8080/configuration?flash=artifacts_set`. The password was never read aloud from the credentials file, never echoed deliberately, and every other call in the session correctly printed only `${#HUB_PW}`; it arrived through curl's own output formatting, which is why the usual discipline did not catch it. **Blast radius, stated precisely rather than minimised:** the value is **not** in git, not in `CHANGELOG.md`/`REPORT*.md`/any committed file (checked), and not in the evidence directory — it is in the Claude Code session transcript under `~/.claude/projects/` on DooPlex, which is operator-readable and persists across sessions. The hub UI is reachable only on the k3s ClusterIP and via the operator's own routes, not from the internet. **The value is deliberately not recorded here; it is stored out-of-band in the usual credentials file.** | **READY (S) — NEW 2026-08-20** | — | **Operator decides whether to rotate.** The hub's own `/configuration` password form does it (`current_password`/`new_password`/`confirm_password`), and per `hub-password-ui-2026-07-13` the DB override wins over the ConfigMap, which stays break-glass. CC can perform the rotation **file-to-file without printing the new value** (the `operator-present-one-time-secrets` convention) if asked — it did not do so unilaterally, because rotating a credential the operator holds in their own head or notes is their call, not CC's. **The reusable half, which matters more than this one password:** never use curl's `%{redirect_url}` (or `-v`, or `--libcurl`) against a basic-auth endpoint — all three re-render the credential. Confirm a redirect with `%{http_code}` and read the flash from a follow-up GET. | **Viktor decides**, CC executes | -| **R-587** | Security & access | P3 | **[P3-LOW] Two root-password files sit in the directory the public ISO is published FROM.** FOUND 2026-09-18 running the ISO release gate's credential criteria for slice 4: `/mnt/5_hdd/felhom.eu/felhom-iso/out/` holds `felhom-pve-9.2-1-v1.24.0-nested-probe-generic.iso.rootpw.txt` and `felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso.rootpw.txt` from 2026-07-22/23, mode 0600, left by two appliance-mode builds. **They are NOT on the bucket** — both `https://iso.felhom.eu/` return **404**, against a control (`felhom-installer-1.28.0-pve9.2-1.iso.sha256`) that returns 200, so the check is real and not a dead probe. **The only thing keeping them off a PUBLIC bucket is the `--include "felhom-installer-*"` pattern in the publish command**, and they are named `felhom-pve-*`, so the pattern misses them by an accident of naming rather than by design. A publish typed without the include, or an include widened to `felhom-*`, uploads root passwords to a world-readable bucket. **Fix shape:** shred the two files (they are three months old and their VMs are long gone), and make the release build refuse to run — or the publish step refuse to start — while any `*.rootpw.txt` exists in the out directory. A pattern that protects by coincidence is not a control. **The two files were SHREDDED 2026-09-18** from the publish source directory, immediately before the 1.29.0 upload ran from it. The directory now holds no `*.rootpw.txt`. **The mechanism half is still open**: nothing stops the next appliance build leaving one there, and nothing refuses a publish while one exists — the `--include` pattern still protects by coincidence. | **OPEN** — **PARTLY DONE 2026-09-18 (files gone; the guard is not built) - rank P3-LOW; owner: CC** | — | — | CC | | **R-616** | Security & access | P3 | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** | — | — | CC + operator | | **R-717** | Security & access | P3 | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). | **OPEN — P3; owner: CC** | — | — | CC | | **R-747** | Security & access | P3 | **[P3-LOW] A stranger can lock the household out of mealie with five wrong logins.** MEASURED 2026-09-30 on 9202 (the R-741 proof): after the install hold opened, a stranger's default-login tries were refused (401) and after five of them mealie answered 423 (locked) to every login — the generated, correct password included. mealie's own brute-force guard, on an app published on the internet; the setup gate and the install hold do not cover an app after its first setup. Not measured: how long the lock lasts. **Needs:** measure the lock's length; decide whether the page tells the household what to do. `audits/night-rulings-2026-09-30/C/C3-mealie-poll.txt` **-- 2026-10-01:** Measured at v3.28.0 (source + 9202): 5 wrong logins lock the ACCOUNT (not the IP) for `SECURITY_USER_LOCKOUT_TIME` hours (default 24); the lock is lifted by an hourly job; an admin can unlock others via `POST /api/admin/users/unlock`, but the household's only admin is the locked account. Both login names are public (`admin`, `changeme@example.com`). **Fixed** (`09` §3 decision 57, decided by CC unattended — operator may reverse): `SECURITY_USER_LOCKOUT_TIME=1`, catalog `a4597cd`; on 9202 the right password answered 423 for 120 min, then 200; a wrong one still 401 (`audits/rulings-2026-10-01/C/`). **Left:** a stranger can renew the lock every hour (per account, public name — option (d) in decision 57 would end that); the page does not tell the household why it is locked; an INSTALLED mealie takes the new setting only when its compose is rendered again (not measured which act does that). | **NARROWED — the lock is 1–2 h; renewal and the page remain; owner: CC** | — | — | CC | @@ -274,7 +265,6 @@ stopping line that lies. | **R-782** | Security & access | P3 | **[P3-LOW] Two side observations of the R-753 sweep, inferred, not measured:** glance's seeded `glance.yml` has no `auth:` block (the dashboard is public to anyone with the address), and homepage's `/api/*` refuses a Host not in `HOMEPAGE_ALLOWED_HOSTS`, which the template does not set (widgets may 400). **Needs:** measure both on 9202; glance: decide whether a public link dashboard is intended (the setup gate does not cover it after setup). | **READY — rank P3-LOW; owner: CC (catalog)** | — | — | CC | | **R-831** | Security & access | P3 | **The Hetzner storage API token (`HETZNER_TOKEN`, the storage project's token in `Secret/storagebox`) was printed into the 2026-10-03 session transcript** — CC read the gitignored `manifests/storagebox.secret.yaml` and its redaction pattern missed the quoted value. It can create, reset and delete Storage Box sub-accounts. Not rotated by the operator's choice (decision 73). **Rotation, whenever chosen (3 steps):** create a new token in the storage project in the Hetzner console → patch `Secret/storagebox` key `HETZNER_TOKEN` in `felhom-system` and `kubectl rollout restart deployment/hub` → delete the old token in the console. Rule for sessions: never print a file that holds secrets — read the one field needed. | **WAITING-ON-OPERATOR — rotation is his call** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator | | **R-870** | Security & access | P3 | **Tester 1's two Cloudflare credentials — the zone API token (`infrastructure.cf_api_token`) and the tunnel token (`infrastructure.cf_tunnel_token`) of the hub's `customer_configs` row `tester-1` — were printed into the 2026-10-04 night session's transcript** (not into any file): a read-only query selected `substr(config_json,1,400)`, and both values sit in the first 400 characters. Tester 1 is CC's disposable test customer (`enkicsifelhom.hu`). **Not rotated, by the operator's ruling of 2026-10-05 06:49 (option B).** **Rotation, whenever chosen (3 steps):** in the Cloudflare dashboard create a new API token for the `enkicsifelhom.hu` zone with the same permissions and refresh the Tester 1 tunnel's token (Zero Trust → Networks → Tunnels → the tunnel → refresh token) → hub → Configs → `tester-1` → Edit → the two Cloudflare fields → Save, then confirm on the box that cloudflared reconnected (`docker ps` health `healthy`) → delete the old API token. Rule for sessions (as R-831): never select a whole config row — name the fields, and never `config_json` without `json_extract` of a non-secret field. | **WAITING-ON-OPERATOR — rotation is his call (ruled: not now)** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator | -| **R-134** | Security & access | P4 | **Two zone-resolvers disagree on depth.** The controller strips labels progressively (`controller/internal/cloudflare/zone.go:18`); the hub's `resolveZone` tries the exact name then `parentDomain`, which strips exactly ONE label (`hub/internal/cloudflare/unblock.go:115,136`) | READY (XS) | — | For a one-label Felhom-issued subdomain both work; for anything deeper the hub silently fails to find the zone while the controller succeeds — the geo-unblock would then no-op with a "no active zone found" error. One concept, two implementations. Same audit §2.6 | CC | | **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | | **R-879** | Security & access | P3 | **A copy of hub.db still holds secrets readable without any key:** each box's hub API key (`hosts.api_key`), each household's owner passphrase and API key (`customer_configs.retrieval_password`, `api_key`), and the PBS-DR token values (`host_pbs_secrets.value`, kept after use). Found 2026-10-05 while answering "what does a hub database backup now contain" (R-133 sealed the console passwords; the off-site passwords were sealed by R-821). A stolen database copy lets an attacker report as any box and read every owner passphrase. `05` §16.2 | READY | — | Seal or hash each with the same seal (the box keys and the passphrases are compared, so a hash may fit; the PBS token is served once, so it could be deleted after use) | CC | @@ -324,7 +314,7 @@ stopping line that lies. | **R-886** | Monitoring & notifications | P3 | **DooPlex's Alertmanager cannot write its own state since the Longhorn restart of 2026-10-05 13:20Z** — every 15 min `Running maintenance failed … open /alertmanager/nflog.…: permission denied` (and the same for `silences`), 8 times by 14:21Z; the pod was recreated 13:20:48Z by that restart. Mail still goes out (`alertmanager_notifications_total{integration="email"}` 5 → 6, `failed_total` 0, 14:23Z), but a silence set now and the record of what was already sent do not survive the next pod restart — so a restart can re-send every active alarm or drop a silence. Likely collateral of the restart (volume ownership on re-attach), not measured. | **OPEN** | — | Compare the volume's file owner with the pod's `securityContext` (`fsGroup`/`runAsUser`); fix in homelab-manifests; prove with a silence that survives a pod restart | operator | | **R-884** | Monitoring & notifications | P4 | **ArgoCD app `monitoring` shows `Deployment/prometheus` OutOfSync** (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. | **OPEN** | — | `argocd app diff monitoring` (or the CR's resource diff) to see what differs, then decide git or live | operator | -## Hub & operator — 21 rows (P2 1, P3 9, P4 11) +## Hub & operator — 13 rows (P2 1, P3 5, P4 7) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -332,23 +322,15 @@ stopping line that lies. | **R-30** | Hub & operator | P3 | **[P2-HIGH] Liveness presence should come from the wait channel, not the report clock.** The box was powered off at the start of the rehearsal, yet the hub carried it as healthy until the staleness threshold expired ~30 min later (`host_stale` 16:05:24 "no report for 30m"; cleared 16:33:24 "was stale for 27m"). The host-delete guard compounds it: RESET refuses while any host row exists, so a stale-but-"Online" host stalls a forced teardown. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-side presence delay; alarms still fire after 30 minutes and no household data is at risk.** | — | Direction: derive presence from **Dir-2 long-poll connectedness (~90 s grace)**, decoupled from notification hysteresis (the hysteresis is right for *alerting*, wrong for *presence*); an agent/ep0 analog can follow. Pairs with R-13/R-23 — the transport already exists, this is about believing it. *(Discussed in-session as "R-29"; that number was already taken by the gate-rot item earlier the same day, so it is R-30.)* | CC | | **R-31** | Hub & operator | P3 | **[P2-HIGH] Offsite provisioning is synchronous with no status affordance.** Save runs the Hetzner sync in-request, so the request can hit the nginx 504 **while succeeding server-side**: the operator cannot tell failed from slow, and a retry races the first attempt. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size M, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **Re-ranked 2026-10-03: P2→P3: operator-only; a known workaround (click once, wait, verify) exists.** | — | Direction: make it async + a status card, reusing the proven **awaiting-card/poll idiom** (v0.138.0 escrow card). **Interim mitigation belongs in R-3 as an operator note: click once, wait, verify — do not re-click.** | CC | | **R-244** | Hub & operator | P3 | **The customer DELETE cascade leaves `app_log_issues` behind, and it is systematic across every venue ever torn down.** Found **2026-08-07** while verifying the `finalwalk` teardown with a **full census** (every table, every column) rather than a per-table query. After a cascade that logged `COMPLETE … full teardown`, **61 rows still matched `finalwalk`**. Four of the five sources are **deliberate and correct** — the cascade's own header states *"Provenance/events are NEVER wiped — audit outlives every tier"*: `events` 16, `notification_log` 14, `host_deletions` 1, `customer_resets` 1. **The fifth is a gap:** `app_log_issues` 29 rows, which the residue purge does not touch (its logged leg covers `reports`/`app_telemetry`/`app_log_tails`/`log_tail_requests`/`notif_prefs`/`selfbind_tokens`/`appliance_registrations` — not this table). **It is not a `finalwalk` quirk:** rows still reference **`c11` 40, `rewalk` 20, `part4` 24** — all three torn down 2026-08-06, whose ledger recorded *"0 occurrences"*. **That prior claim was measured with a narrower query and does not survive a full census; the correction is recorded rather than the measurement quietly redone.** **Why it was probably never written, established rather than assumed:** the table is a **fleet-wide aggregate** keyed on `app_name`+`fingerprint` with an `affected_customers` JSON list — of the 29 `finalwalk` rows, **12 reference only `finalwalk`** (orphans, safely deletable) and **17 are shared with LIVE customers** (`demo-felhom`, `peti-felhom`, …) and **must not be deleted, only de-referenced.** A naive `DELETE … WHERE customer LIKE` would destroy a live customer's issue history — which is very likely why the leg does not exist, and is the reason this is not a one-line fix. **Severity is LOW and stated plainly: no secret material is involved** — app name, fingerprint, message text, counts, timestamps. What survives is a deleted customer's *identifier* inside an aggregate row. **Proposed shape:** a residue leg that (a) removes the customer id from `affected_customers`/`context_customer`, and (b) deletes rows whose `affected_customers` becomes empty; plus a one-off sweep for the four already-torn-down venues. **The general lesson is the reusable part:** *a per-table absence query is not a census.* The teardown verification is now a full-schema sweep, and that is what found this. **Not fixed** — a cascade change needs its own red-proof and this session was scoped as a spike plus two operations. Evidence: `tests/teardown-finalwalk-2026-08-07.md`. **⚠ STILL OWED, AND NOW MEASURED RATHER THAN ESTIMATED (2026-08-08 census, read-only, no truncation).** `app_log_issues` holds **1309 rows**; **71 reference a torn-down venue** (`finalwalk`, `c11`, `rewalk`, `part4`); of those **44 are ORPHANS** — they name only torn-down customers and are safely deletable — and **27 are SHARED with a live customer** (`demo-felhom`, `peti-felhom`, …) and **must be de-referenced, never deleted**. 1238 rows are untouched. **The 27 are exactly why the leg was never written**, and why a `DELETE … WHERE customer LIKE` would destroy a live customer's issue history. **What it needs, precisely:** a cascade leg that (a) removes the customer id from `affected_customers` / `context_customer`, and (b) deletes only rows whose `affected_customers` becomes empty; plus a one-off sweep for the four venues already gone. **Why it was NOT done on 2026-08-08:** the fix is hub code, and that session's scope forbade a hub version bump; a hand-run SQL mutation over 71 rows — 27 of them needing surgical de-referencing — with no tested code path and no red-proof is precisely the shape that goes wrong on a live database. **It accumulates one venue at a time, so the next walk adds to it**; the numbers above mean the next session starts from data rather than a guess. **⚠ IT GREW AGAIN, AS PREDICTED — walk5 teardown, 2026-08-08.** The fifth walk's venue was torn down with a full-schema census taken **before and after**: **168 rows → 67**. Of the 67, **37 are by design** (`events` 21, `notification_log` 14, `host_deletions` 1, `customer_resets` 1) and **30 are `app_log_issues`** — this row's gap, and the count was **predicted in the pre-run enumeration rather than discovered afterwards**, which is the difference from the ledger that once recorded *"0 occurrences"* from a narrower query. **The running total across torn-down venues therefore rises from 71 to ~101 rows** (`finalwalk`, `c11`, `rewalk`, `part4`, now `walk5`) — the shared-with-a-live-customer subset must still be **de-referenced, never deleted**. **It accumulates one venue at a time and it did so again.** Evidence: `tests/walk5-r201-2026-08-07/teardown-walk5-2026-08-08.md`. **2026-09-25:** `peti-felhom` is no longer a live customer (deleted through the cascade, journal #20); 8 `app_log_issues` rows still name it — the same gap. | **READY** — owner Viktor | — | — | operator | -| **R-277** | Hub & operator | P3 | **Three hub surfaces jointly present a HEALTHY off-site tier as an absent one — and it produced a wrong operator statement during this run.** For demo-hp on 2026-08-09 the box was pushing off-site daily without a gap (18 restic snapshots, `last_status: ok`), yet: (a) the customer page's Backup panel read `Snapshots 0 / Repo Size 0 MB / Integrity Unknown` — it renders the **local disk tier**, while the healthy `offsite` object sits **in the same report** unrendered on that panel; (b) the Offsite page read `0.0 GB` — true, but a 162 KB repo rounds to nothing; (c) a stale `offsite_delivery_stuck` event from **2026-08-07 10:19** (not recurring) reads as current state. **Three independent surfaces agreeing on a wrong picture is how a working backup gets "fixed".** It did exactly that here: the rehearsal reported a fleet-wide off-site outage to the operator and had to retract it. **Note the true half:** demo-felhom IS genuinely stuck (`offsite.state=needs_credential`, no run has ever succeeded) → **R-278** | **READY (S) — NEW 2026-08-09** | — | Render the offsite object on the offsite row; show bytes not rounded GB; distinguish a live alarm from event history | CC | -| **R-581** | Hub & operator | P3 | **[P2-MED] `ORDER BY received_at` cannot answer "the newest report" — the column has SECOND granularity.** FOUND 2026-09-18 building hub v0.118.0 (R-558): `Store.CustomerLanguage` read the household's language from the newest report ordered by `received_at`, and `TestNewestReportedLanguageWins` failed — four reports written in the same test tick all carry the same `datetime('now')` string, so the "newest" was whichever row SQLite felt like returning. On a real box the same shape appears whenever two reports land in one second (a settle burst, a restart race), and the symptom would have been a household switching language on their dashboard and getting the old language back at random. FIXED in the same release by ordering on the autoincrement `id`, which is the real insertion order. **What is still open, and it is the reason this is a row rather than a note: `GetCustomers()` has the same shape** — `INNER JOIN (SELECT customer_id, MAX(received_at) …)` with no tie-break — and it is what the whole operator dashboard and `countBoxesBelowFloor` read. A same-second tie there picks an arbitrary report's health, version and vitals. Not observed in the wild; not looked for either. **Fix shape:** tie-break every newest-report query on `id DESC`, or give `reports` a monotonic ordering column and use it everywhere; then a test that writes two reports in one tick and asserts which one wins. | **READY - rank P2-MED; owner: CC** **Re-ranked 2026-10-03: P2->P3: operator dashboard only; the household-facing language case was fixed; never observed.** | — | — | CC | -| **R-600** | Hub & operator | P3 | **[P2-MEDIUM] "Full teardown" is logged while the deleted box's WireGuard peer is still configured on ep0.** FOUND 2026-09-20 by the slice-6 drill's teardown, **measured on ep0 rather than inferred from the hub**. The customer delete cascade finished at 19:41:54 with `customer DELETE cascade COMPLETE for drill-en-0920 (journal #18) — full teardown`, and every hub-side row was gone (0 configs, 0 hosts, 17 residue rows purged, PBS tenancy deprovisioned, escrow demoted). **Three minutes later `wg show wg0 allowed-ips` on ep0 still listed `10.77.0.5/32`** — the drill box's peer — because `wgsync` pushes on its own cycle. **Watched to the end rather than assumed: the peer was gone by 17:47:46Z — it outlived the *full teardown* line by about 6 minutes.** (My first estimate said ~35, read off the gap between two log lines; the sync runs oftener and only LOGS when something changes. That is the second time in this session that a period inferred from two log lines was wrong — the other was the delete's own staleness window. **A period read off two log lines is not a measurement.**) **The 2026-09-14 drill's findings say "the teardown removes it through the host delete"; measured, the host delete removes the hub's RECORD and the peer goes on the next push.** The mechanism is not broken — it is asynchronous, and the log line claims a completeness it does not yet have. **Why it is P2 rather than P3:** a session that tears down, reads *full teardown*, and leaves is the normal case; the peer outlives it by minutes, and the third teardown layer is the one the workspace rules single out as the one that gets forgotten. Six minutes is short — but the session that reads *full teardown* and leaves has no way to know it is six and not six hours. **Fix shape (smallest first):** the cascade triggers a wgsync push before it logs COMPLETE, or the log line says what is still pending and when ("wg peer removal queued; next push in N min"). A session should not have to read ep0 to know whether a teardown finished. **2026-09-25 (Peti's retirement):** nothing to remove for `peti-felhom-86d37d` — its host was deleted 2026-07-15 and ep0's live `wg show` carries no peer beyond the demo boxes', drill-r50 and the operator OOB (`audits/retire-peti-2026-09-25/A1-ep0-before.txt`); whether that July delete removed a peer, or none existed, is not recorded. **-- 2026-09-28: the Day-0 test install's peer (`drill-g0276`, key `Ly0yjK…`, 10.77.0.5) was checked on ep0 the day after its customer DELETE: gone from `wg show`, absent from `/etc/wireguard/*.conf` (control: a live key found), no namespace or file names the customer — the hub's sync removed it; nothing was removed by hand.** `audits/logins-nvme-2026-09-28/E/`. **MEASURED AGAIN 2026-09-30 (a HOST delete, not a customer delete; new-household drill):** the host was deleted 07:23:03Z and `wgsync: pushed 4 peers` at 07:23:51Z — `10.77.0.5/32` gone from ep0 48 s later, the other four peers unchanged (`audits/evidence-drill-new-household-2026-09-30/teardown/layer4-ep0.txt`). | **READY - rank P2-MEDIUM; owner: CC (hub)** **Re-ranked 2026-10-03: P2->P3: measured as seconds to minutes and asynchronous by design; operator-only log wording.** | — | — | CC | -| **R-728** | Hub & operator | P3 | **[P3-LOW] A customer created with one press was created TWICE, and the first of its two connect mails holds a dead link.** MEASURED 2026-09-30 on `Tester-2`: the hub logged `Customer config created: Tester-2` twice in the same second and two self-bind mints (hashes `c40df008…`, `6a1cbef4…`); a mint replaces the previous link (single-active), so one of the two identical mails the tester received answers „expired". Cause not established (a double form submit, or the handler run twice). **Fix direction:** make the create idempotent within a few seconds (or disable the button on submit), and pin it. The workaround for the tester is in STATUS. | **READY — rank P3-LOW; owner: CC (hub)** | — | — | CC | | **R-882** | Hub & operator | P3 | **Longhorn on DooPlex could not grow a volume online: its `instance-manager` (116 days up) called a host process that no longer existed** — `nsenter: cannot open /host/proc/196610/ns/mnt` on every expansion retry, and an offline growth was blocked by the expansion's own attachment ticket (found 2026-10-05 growing `hub-data` to 2 Gi). A restart of the instance-manager (operator-approved) fixed it: 77/77 volumes back `attached/healthy` in 110 s. **Why the cached PID went stale was not established** (likely a containerd/k3s or iscsid restart after the instance-manager started), so it will recur after the next such restart and stay invisible until a volume needs to grow. `audits/hub-db-offsite-2026-10-05/partA/step1-*.txt` | **OPEN** | — | Find which host process the PID was and whether Longhorn 1.10.x re-resolves it; until then, before growing any volume, check the instance-manager's age against the last k3s/containerd restart | operator | | **R-883** | Hub & operator | P3 | **8 DooPlex workloads run an image by a moving tag (`:latest` or none), so any pod restart is a silent upgrade.** Measured 2026-10-05: the Longhorn restart restarted zipline on `ghcr.io/diced/zipline:latest` (pull Always), which pulled 4.8.0; 4.8.0 refused its database (`cannot safely migrate from prisma to drizzle: expected migration 20260508022000 … was not applied`) and crash-looped. Fixed for zipline by pinning `4.7.0` (homelab-manifests `90f60e4`, `4c8ec7a`; 4.7.0 applied the four missing migrations; a dump from before is kept out-of-band). The other 7 were counted, not named or changed. | **OPEN** | — | List the 7 (`kubectl get deploy,sts -A` images without a fixed tag), pin each to the running version, and let Renovate move them | operator | -| **R-92** | Hub & operator | P4 | Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable | READY (XS) | — | Widen precision when retention becomes customer-visible | CC | | **R-264** | Hub & operator | P4 | **Twenty-one facts the boxes report that the hub can now decode nowhere, each allowlisted with a reason rather than silently skipped — and for these the reason is "no consumer today, and one is arguably owed".** Split out of R-260 on 2026-08-08 so that closing the CLASS (gated) and fixing its sharpest instance (`operator_key_configured`) could not be mistaken for having decided what the hub should do with the rest. **The list, grouped by what a consumer would be for.** **(a) Guest-network health — `guest_net` and its seven children** (`checked_at`, `has_route`, `dhclient_alive`, `heal_succeeded`, `heals_last_hour`, `last_heal_at`, `damped`). The R-54 watchdog reports per-guest network state and self-heal counts every cycle and the hub — the component that emails the operator — models none of it. There is a live incident in this project's own record where a killed `dhclient` took a tunnel down for 1 h 15 m (`audits/INCIDENT-guest-dhclient-killed-2026-07-20.md`); a recurring-heal signal is exactly what would have surfaced it. **This is the strongest candidate of the twenty-one.** **(b) `selfupdate_pending` + `selfupdate_pending_version`** — an agent that has flipped its binary and never committed reports pending on every heartbeat so that "the operator sees WHY the version isn't advancing", and no operator can see it. **(c) `mgmt_plane.healed_recently`** — bounded: the hub DOES alarm on the `privsep_healed_at` timestamp beside it, so the recurring-clobber signal is not lost, only this flag. **(d) `restore_tests.mount_parity` + `mount_inventory`** — R-262's subject; the verdict is not lost (a mismatch fails before `Pass` is set) but the hub cannot tell a full-fidelity pass from a boot-only one. **(e) `pbs_dr.applied_at`.** **(f) Controller-side: `config_hash`, `reporting_disabled`, `stacks`, `storage.migrated_to`, `backup.last_db_dump`, `backup.last_integrity_check`** — the last two are backup-integrity timestamps, which is the "presence is not success" neighbourhood. **For each the question is the same and is NOT answered here: is it wanted? If the hub should act on it, model it and name what consults it. If it should not, the honest end is that the emitter stops sending it** — a fact emitted forever and consumed nowhere is a future false green waiting for someone to write a check against it. **Deliberately not decided in the G-1 session**, whose scope was the gate plus the operator-access instance; unilaterally removing emitters would also break the byte-identical cross-repo host-report golden and is a coordinated two-repo change. **⚠ DISPOSITIONS RECORDED 2026-08-13 — and the first thing to say is that they were NOT in this register before today.** The rulings were made on 2026-08-12; this row still read **READY — owner Viktor** and the gate's twenty entries still all said *"arguably owed"*, so a session told to *"re-read the dispositions from the register"* would have found none. They are written down now, which is the point of writing them down. **THE COUNT WAS ALSO WRONG:** this row says *twenty-one*; the gate's allowlist held **twenty**, measured. Twenty is the number the dispositions below account for, exactly. **(1) BUILD A READER — four groups, fourteen facts.** (a) guest-network health, (b) the staged-update-pending pair, (d) the restore-test depth pair, (f-part) the two backup-integrity timestamps. **(2) NO READER WANTED — five facts, now recorded as `not consumed, DELIBERATELY` with the ruling and its date** in `scripts/wire_contract_gate.py`, each with its own reason rather than a bare refusal: `mgmt_plane.healed_recently` (the hub already alarms on the timestamp beside it), `pbs_dr.applied_at` (`pbs_dr.state` is the verdict; the timestamp alone is the attempt-read-as-result trap), `config_hash` (the hub authors the config and knows its own generation), `stacks` (the app view is built from the purpose-built `app_telemetry` wire), `storage.migrated_to` (box-local bookkeeping with no hub-side intent to reconcile against). **The emitters are deliberately left alone** — removing one is a coordinated two-repo change and breaks the host-report golden; the honest end here is a recorded decision, not a deletion. The gate grew a THIRD entry kind to carry them, because an undecided fact and a decided one must not read alike. **(3) `reporting_disabled`, decided on its own merits: RECLASSIFIED `redundant`** — `health.status = "disabled"` travels in the same minimal report, is decoded into `reports.health_status`, and IS rendered. **The decision surfaced a real defect the flag would not have fixed: the staleness checker is age-only, so a deliberately-silent box still alarms → R-321.** **PROGRESS, 2026-08-13: the first reader is BUILT — guest-network health, R-319.** Its eight allowlist entries are **removed** (an allowlisted tag is skipped, so leaving them would have meant the new reader's fields were never checked); the gate's checked-tag count rose **182 → 190** and skipped fell **88 → 80**, which is the positive control that the wiring is real. **WHERE THE TWENTY NOW STAND: 8 read · 5 deliberately unread · 1 redundant · 6 still owed a reader** (`selfupdate_pending`, `selfupdate_pending_version`, `restore_tests.mount_parity`, `restore_tests.mount_inventory`, `backup.last_db_dump`, `backup.last_integrity_check`) — counts measured from the allowlist, not estimated. **Only ONE reader was built on purpose:** four at once is a design session pretending to be an implementation, and this one now tells us what the other three cost | **OPEN — 6 of 20 still owed a reader; dispositions recorded 2026-08-13, first reader shipped (R-319)** — owner Viktor | — | — | operator | -| **R-292** | Hub & operator | P4 | **The artifact-save flash conflates three different facts, and a failing test found it rather than a reading.** `artifact_sha_invalid` reads *"the Gitea sha lookup failed (version missing / Gitea unreachable) or the manually-entered sha is invalid"* — three causes, one message, and the operator acts differently on each. It surfaced because scenario E of the new installability gate kept reporting `artifact_sha_invalid` where it expected `artifact_unverifiable`: `resolveArtifactSHA` ran first and swallowed the distinction. **Worked around in v0.102.0 by ORDERING** — the installability probes now run before the sha resolution, so an unreachable registry is reported as unreachable — **but the underlying message is untouched and still conflates on its own paths** | **READY (XS) — NEW 2026-08-09** | — | Split it into "version not found", "registry unreachable" and "invalid sha" | CC | | **R-402** | Hub & operator | P4 | **The off-site integrity verdict and its depth are published to the hub and NO hub surface reads either.** `offsite.last_integrity_ok` has been on the wire since controller v0.227.0 and `offsite.last_integrity_depth` since v0.228.0; both are allowlisted in `scripts/wire_contract_gate.py` **with their reason**, which is why the gate is green rather than silent. **The order is deliberate and is the opposite of the one that produced R-331:** publish the value first, build the display when someone decides what the screen should say. R-331 removed a hub Backup card that rendered `Integrity Unknown` for every customer forever from fields nothing wrote. **The depth is not decoration:** "checked, OK" means two different things at structure depth and at 100%, so a card showing the verdict without the depth shows the same words for a check that re-read every byte and one that only read the index. **WHAT HAPPENS IF NOBODY ACTS:** the operator can only answer "was this customer's off-site store verified, and how deeply?" by reading that box's own log. | **OPEN — SMALL, needs a HUB decision first** | — | Decide what the hub screen should say, then model both fields hub-side and delete the two allowlist entries together. `offsite.last_integrity_check` is already decodable and is not allowlisted. | Viktor decides, CC builds | | **R-451** | Hub & operator | P4 | **[P3-LOW] UPDATE ARC SLICE 7 — a fleet sweep: the operator can SEE, and MOVE, how far behind every box is.** Slices 1 and 2 make one box's state visible on that box's own pages. The operator has no fleet view, and **it is not derivable from what is already reported: the hub's report payload carries container name, state, CPU and memory, and NO image field at all** (spike §5, which is why Peti's box could only be recorded UNKNOWN). So this is a hub-side change as well as a controller one. Rank LOW today because the fleet is two enrolled boxes; it rises with the fleet. `architecture/09-update-architecture.md` §6, §8.4 **BOTH SIDES VERIFIED 2026-09-21, and it is cheaper than this row implies.** The controller's payload carries no image (`internal/report/types.go` L98–103) and the hub's `Store.SaveReport` (`hub/internal/store/store.go:965`) denormalises only container **counts** — but **the hub stores the raw report JSON whole**, so a new controller field lands there the day it is sent. What is missing is the denormalisation and the page, not the transport. Shape in `09` §6.2–6.3; the payload question is `09` §3b **Q7**. **-- RULED 2026-09-23 (`09` §3 decision 18):** the report carries, per compose service (database included), the installed reference, the catalog reference and the badge state. **Built later, when the fleet grows** — Q7's recommendation, confirmed. | **DEFERRED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the fleet sweep (report field, hub denormalisation, fleet page) is ruled and not built — deferred until the fleet grows (`09` §6.3)) — **RULED — build deferred until the fleet grows; owner: CC** | — | — | CC | -| **R-544** | Hub & operator | P4 | **[P3-LOW] The host-delete log line says „escrow deleted: true" while the documented (and actual) effect is DEMOTION to retained custody.** MEASURED 2026-09-16 during the teardown of the fresh box: an unacknowledged delete was correctly refused 409 („has key escrow (acknowledgement missing)") and the refusal text promises the acknowledgement „moves it to retained custody"; the acknowledged delete then logged `host deleted: tester-1-33b6a9 (escrow deleted: true)`. The hub's own customer page states the truth — „host deletion only demotes custody, never destroys it… recovery-key custody is demoted to retained custody, not destroyed", with the customer delete named as „the one true purge point". **Nothing is broken; the log is.** An operator reading that line during an incident would believe a household's last key had just been destroyed, and the R-304 retention exists precisely so it is not. **Fix shape:** log what happened — `escrow custody demoted to retained (host delete)` — and keep the boolean's name out of operator-facing text. | **READY — rank P3-LOW; owner: CC (hub)** **Re-ranked 2026-10-03: P3->P4: operator log wording only; no household meets it.** | — | — | CC | -| **R-599** | Hub & operator | P4 | **[P3-LOW] A drill's teardown is blocked for 30 minutes by design, and nothing says so.** FOUND 2026-09-20 tearing the slice-6 drill down. VM destroyed at 17:12Z; the hub then refused **both** `POST /configs//delete` (409, *host … is ONLINE*) and the host delete (`deletable:false`) — correctly, because an online host would receive permanent 401s. But "online" is not a liveness probe: it is a **report-staleness window**, and the window is **45 minutes** — `manifests/hub.yaml` sets `alerting.stale_threshold: "45m"`, which `hostStatus()` reads (`ok` under it, `stale` over, `down` at 2x). A machine that no longer exists therefore reads ONLINE for three quarters of an hour. **Measured the boring way, and worth recording:** this row first said 30 minutes, because `monitor/host_staleness.go`'s literal default says 30m — the DEPLOYED value is in the manifest, and the 409s kept coming after the half hour was up. Reading a default and calling it the live value is the same mistake in a smaller coat. **The consequence is not theoretical:** a session that destroys its VM and then tears down the hub side walks away believing the delete failed, or leaves the customer behind — and the 2026-09-14 drill's teardown had the same shape without recording this. **Fix shape (smallest first):** the 409 body says *how long* it will refuse ("the last report was N minutes ago; deletion opens at HH:MM"), and `runbooks/target-selection.md`'s drill section names the wait. A force flag is NOT proposed — the refusal is right, only silent about its own clock. | **READY — rank P3-LOW; owner: CC (hub)** **Re-ranked 2026-10-03: P3->P4: operator teardown comfort; refusal is correct.** | — | — | CC | | **R-719** | Hub & operator | P4 | **[P2-MEDIUM] A customer who already exists never gets a fresh self-bind link when their new box registers: the last link expires in 7 days and nothing re-sends it.** MEASURED 2026-09-29 (new-household drill, `tester-1`): the previous link went out 2026-09-17 07:25 UTC at a host delete and expired 2026-09-24; the box registered at 19:11:30 UTC and its console told the volunteer to open the link from their e-mail — there was none that worked. Hub source: the link is sent at customer creation, RESET, e-mail set on a box-less customer and host delete (`selfbind_mint.go` callers `hosts.go:908`, `configs.go:850`, `customer_reset.go:162`) — never on appliance registration. The volunteer guide says the operator needs to press nothing. The operator pressed „Send self-bind link" (the mail arrived in 1 s) — an operator step the volunteer depends on, recorded, not an intervention. **Fix direction:** send the link when an unclaimed appliance registers while a box-less customer waits with no live link (R-509's first fix shape), or the guide's operator part says: press it the day the volunteer installs. Evidence: `audits/evidence-drill-new-household-2026-09-30/` `phase0/operator-steps.txt`. **CHANGED AND BUILT 2026-09-30 (hub v0.126.0) — the brief's shape was not buildable:** a box registers UNCLAIMED (uuid, MACs, host keys, hardware — nothing of a customer), so "send the link when their box registers" would mail every waiting customer. Built instead: the expired AND used link pages offer „Új linket kérek" → a fresh link to the address registered for that link's customer, only when it has no box, ≤1/h per customer, identical answer for any token (no oracle). Live: the button on the real hub, the same page for a made-up token, no mail; the mint+send path unit-proven (RP42). Evidence: `audits/evidence-fixes-first-tester-2026-09-30/``partD/`. | **WAITING-ON-OPERATOR** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the operator has not reviewed the changed page shape ("operator may prefer another"), and mint+send is unit-proven only) — **CLOSED 2026-09-30 — hub v0.126.0 (changed shape; operator may prefer another)** | — | — | operator | | **R-814** | Hub & operator | P4 | `PBS-storage-1` (u629193, box 611421) still `status=active`, 19.9 MB | **VERIFY** (2026-10-03 triage: a July watch row with no id; given R-814. WAITING-ON-OPERATOR — no record found that the box was deleted.) — WAITING-ON-OPERATOR | operator console | Delete the box | operator | | **R-844** | Hub & operator | P4 | **The household's OS-update line exists only on the hub's customer timeline.** 2026-10-04: the box itself has no event surface for agent results (the controller UI shows no timeline), so `os_update_applied` is a hub customer event (info: recorded, never mailed). Its stored text is the hub's English sentence; the hu/en bundle text (`mail.event.os_update_applied`) is used only if it is ever mailed. Fix direction: a controller-side line (the controller already polls the agent's local API) when the box gets a household timeline. `audits/os-guest-lane-2026-10-04/partG/hub-customer-timeline-demo-hp.txt` | **READY — owner: CC** | — | — | CC | -| **R-855** | Hub & operator | P4 | **The hub's start log prints "after -1 healthy ring-0 night(s)" for the TEST override `OS_DOCKER_APPROVE_NIGHTS=0`** (the internal "none" value; the WARN line before it is right). Cosmetic, TEST configuration only. Fix: print 0. `audits/os-docker-crash-2026-10-04/partB/` (hub log 16:27:42) | **READY — owner: CC** | — | — | CC | +| **R-888** | Hub & operator | P4 | **Two fields the controller reports are never read by the hub: `stacks.deployed` and `storage.decommissioned`.** Surfaced 2026-10-05 when the wire-contract gate stopped counting a tag named only in a comment (R-555): the hub decodes no deployed-app list and no per-drive decommissioned marker from the box report. Allow-listed in `scripts/wire_contract_gate.py` so the gate stays green. **Needs a decision, not a fix:** does any hub surface (customer page, drive view) need either? If yes, decode and show it; if no, record why. | **OPEN** | — | Operator: say whether the hub needs either field | operator | ## Business & legal — 6 rows (P2 4, P4 2) @@ -361,15 +343,14 @@ stopping line that lies. | **R-89** | Business & legal | P4 | Retention as a per-customer **commercial** policy on the hub | READY (increment 2) | — | Policy object + reconciler → ep0 prune job; keep box tokens write-only | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 52 rows (P3 4, P4 48) +## Process & tooling — 43 rows (P3 4, P4 39) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| | **R-565** | Process & tooling | P3 | **[P3-LOW] The English page test sees only ACCENTED Hungarian: an ASCII-only Hungarian word left in a template passes it on the English page.** FOUND 2026-09-17 by slice 1 release C (R-556, controller v0.250.0): after the extractor and the tests were green, a by-eye review of the English renders found six Hungarian fragments still in JavaScript strings — „, majd a(z)” and „FIGYELEM:” in the storage decommission dialog, „jelenlegi:” on the drive-init list, the uptime units „mp” and „p” and the count word „ db” on the debug page. All six were converted by hand; **no test failed on any of them**, because `TestI18nEnglishPages` looks for Hungarian letters and the extractor's ASCII word list (`i18n_extract.py` `ASCII_HU`) is used by neither test nor gate. Release B's review had found more of the same kind (Konfig, Megtartva, helyi, pl., Befejezve, automatikus, jelenleg:, kedd/szerda/szombat, szint). **Fix shape:** run the ASCII word list over the English renders in `TestI18nEnglishPages` (after the data mask), with a negative control on an English sentence and a decoy planting „mp” in an English value; extend the list with the words releases B and C found. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-578** | Process & tooling | P3 | **[P3-LOW] A helper that takes the settings lock must never be called from inside a settings callback — there is no gate, only one test in one package.** FOUND 2026-09-18 the hard way, by localisation slice 2 release C introducing exactly that: `UpdateOffboxStatus` holds the settings WRITE lock while it runs its callback, `boxLang()` reads the language through the READ lock, and `sync.RWMutex` is not reentrant — so the off-site run's final status write DEADLOCKED, **holding the settings lock**, which would wedge everything else on that box that touches `settings.json`. The only symptom was `go test ./internal/backup/` going from 8 minutes to a 25-minute timeout. Fixed by hoisting the language resolution; `TestNoteHelpersAreNotCalledUnderTheSettingsLock` (internal/backup) now names the file and line in a second. **What is still open:** that test covers `internal/backup` only, and it knows only the `note`/`noteErr`/`boxLang` helpers. Any other settings-reading helper, in any other package, can make the same mistake with nothing to catch it but a hang. **Fix shape:** promote it to a gate over every package, keyed on "a call to a method that reads settings, inside a literal passed to a `settings.Update*` function"; or give `Settings` a re-entrant read path and remove the class. | **READY - rank P3-LOW; owner: CC** | — | — | CC | | **R-733** | Process & tooling | P3 | **[P3-LOW] The test bench has NO swap and the boxes have 512 MiB — so a box proof can pass on swap where the bench fails, and nobody records whether a customer guest has swap.** MEASURED 2026-09-30 (R-732): immich's first start was OOM-killed 61–104 times on the bench (swap 0) and passed on 9202 by swapping ~108 MB; the bench given 512 MiB swap passed too. demo-hp 9201, 9202 and demo-felhom 9201 all read `swap: 512`; the golden's guest config is not recorded in its bake evidence, so a customer guest's swap is NOT measured. The harness's memory watch judges `anon` against the limit and never reads `memory.swap.current`. **Needs:** the golden's `swap` read and recorded; the box walk and the harness report `memory.swap.peak` beside `anon`; a decision whether proofs run with swap off (the stricter venue, as R-732's fix was proven). | **READY — rank P3-LOW; owner: CC (harness + golden evidence)** | — | — | CC | -| **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. | — | Run the 2026-10-12 check; if a job is lost again, read Gitea's log for the assignment within the hour (it rotates) and the runner's fetch timeout | operator | -| **R-129** | Process & tooling | P4 | **Every doc says demo-hp has "no baked SSH key"** and needs the G1 break-glass password — but `ssh -o BatchMode=yes demo-hp` authenticated **by key**, first try, 2026-07-31 | READY (XS) | — | Stale in the expensive direction: a session that believes it sends itself to the hub vault for a credential it does not need. Verify who owns the key and when it landed, then correct `CLAUDE.md`, `runbooks/target-selection.md:41-42`, `runbooks/workspace-CLAUDE.md` and `felhom-agent/CLAUDE.md` together — or remove the key if it was not deliberate | CC | +| **R-887** | Process & tooling | P3 | **Some CI jobs are never run, and Gitea fails them ~10–13 minutes later with no log.** Seen 2026-10-05: felhom.eu job 1361 (commit `1122b5c`) and felhom-controller job 1357 (`114ff27`): every step reads `failure`, including the first fetch, the log API answers `file does not exist`, and the runner pod's log has no `task` line for them (its task ids are job id + 1). A re-run through the API ran the controller job normally (success in 32 s) but the felhom.eu job was again never picked up and failed after ~12 min. The runner pod (`gitea-system/act-runner`, image `felhom-act-runner:0.1.0`) had restarted 5 times ~142 min earlier, around the Longhorn instance-manager restart (R-882). Suspected, NOT measured: a stale runner registration claims jobs it never runs — the session's Gitea token cannot list runners (`read:admin` scope). Consequence: a red CI verdict that is not about the code, and **no failure mail** (the alarm step never runs either), so only the pull check sees it. **CORRECTED 2026-10-05 18:21 (operator's screenshot of Gitea → Site Administration → Runners): ONE runner only — ID 2, `felhom-gates-runner`, v0.6.1, label `felhom-gates`, Idle, last online „now". There is no old registration; the stale-registration guess (this row's first text and the reviewer's) was WRONG.** **RE-DIAGNOSED 2026-10-05 (round 2), from the logs that survive:** (1) **„lost in a runner restart" does NOT fit** — the runner pod last restarted 13:24:42Z (`restartCount 5`, all around the 13:20Z Longhorn restart), the lost attempts started 1.5–2.5 h later. (2) **FOUR attempts were lost, not two:** controller job 1357 (start 15:05:41Z → failed 15:18:38Z), felhom.eu job 1359 (15:15:37 → 15:28:38 — the previous session blamed that one on the BusyBox fault; the runner never ran it), job 1361 (15:33:21 → 15:43:38) and its API re-run (15:46:48 → 15:58:38). None has a `task` line in the runner log; every one was failed at a :38-second mark on a 5-minute step, 10–13 min after it was handed out — **the shape of Gitea's periodic „zombie task" stop** (a task assigned to a runner that never reports is failed after ~10 min; no log exists because none was written). (3) The runner's task ids are NOT job id + 1 (the controller re-run was task 1363). (4) **Gitea's own log for the window is gone** — the pod log starts 16:16:32Z (rotated), so the assignment side cannot be read. **Likely mechanism, NOT proven:** the runner's fetch-task request timed out on its side after Gitea had already assigned the task, so the task was orphaned. In that same hour this session polled Gitea's jobs API hard (15 pages every 15 s per wait loop) and Gitea logged „slow" requests — a plausible load cause, and the session's own. Mitigation taken: the session's CI waiter now polls once a minute. **Nothing changed on DooPlex.** **MECHANISM SEEN 2026-10-05 17:15–17:28Z, with Gitea's own log (round 2):** catalog run 1368 (`4828dc7`) — 17:15:14 the job is marked started; 17:15:16 `router: slow POST /api/actions/runner.v1.RunnerService/FetchTask for 10.42.0.42 (the runner), elapsed 3192ms`, then `UpdateRepoRunsNumbers … context canceled` and `GetActionWorkflow: EOF` — **the runner abandoned its fetch after Gitea had assigned the task**; the runner log has no line for task 1371; 17:28:39 `actions/clear_tasks.go:174 stopTasks() [W] Cannot transfer logs of task 1371` — Gitea's zombie-task stop. **The load at that minute:** an outside crawler (216.73.216.78) walking commit pages and `archive/*.tar.gz`, and THIS session's CI waiter, whose 15-page job listings took 13–31 s each. An API re-run passed in 7 s. **Done in-session:** the waiter now asks `GET …/actions/runs?head_sha=` once a minute (1 s). **Not done (DooPlex, the operator's):** the runner's fetch timeout and Gitea's exposure to the crawler. | **OPEN** **DATED CHECK 2026-10-12 (DUE-CHECKS):** if no job was lost since 2026-10-05 16:00Z (no completed job whose runner log has no `task` line / whose log API answers `file does not exist`), close. | — | Operator: decide whether to raise the act-runner fetch timeout and/or rate-limit the public Gitea pages the crawler walks; meanwhile re-run a lost job via `POST /repos/admin//actions/runs//rerun`. Keep the 2026-10-12 check | operator | | **R-206** | Process & tooling | P4 | **The build-cache cap and the weekly prune exist only as a hand-edited `/etc/docker/daemon.json` on DooPlex — not in Ansible, so a rebuild loses them.** The `node_housekeeping` role must also carry the prune, which today it is forbidden to run | **READY (M) — NEW 2026-08-05** | — | **The spike validated the recipe; this row builds it.** Three parts. **(a) Template `/etc/docker/daemon.json`** with the **`policy` array** form — **the flat form (`{"gc":{"reservedSpace":…}}`) is SILENTLY IGNORED**, measured: the daemon starts, logs nothing, and `docker buildx inspect` still reports the built-in defaults. **The oracle is `docker buildx inspect`, never `dockerd --validate`** — the validator returned `configuration OK` for a bogus key AND for the config that then **crashed the daemon** (`filter` takes one value per policy entry, not an array; `error initializing buildkit: filters expect only one value`). **(b) Narrow the role's Docker ban** (`node-housekeeping.sh.j2:10-14`) to permit exactly `docker builder prune -af` and nothing else — the ban's stated premise ("Docker here runs only unrelated jarr-* dev containers") is obsolete: the growth is Felhom Go build cache. **The measured prune is SYNCHRONOUS** (150.35 GB back at t+0, two consecutive polls <1 MB apart within 60 s) — **unlike containerd's image GC, so it needs no `settle_imagefs` equivalent**, but it MUST measure the filesystem rather than trust the command: `prune` claimed **156.9 GB** and the filesystem returned **150.35 GB**, the 6.5 GB gap being layers still shared with images. **(c) A restart-safety note in the role:** a bad `daemon.json` takes the daemon down AND leaves the `unless-stopped` dev containers stopped — they needed a manual `docker start` — so the role must restart-and-verify, not validate-and-assume. Recipe + every measurement: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` | CC | | **R-208** | Process & tooling | P4 | **Every Felhom Go build re-downloads its modules because `ARG VERSION` sits ABOVE the module-download layer — ~440 MB of dead cache per build, 90.5 GB of the 157 GB** | **READY (S) — NEW 2026-08-05, ROOT CAUSE PROVEN** | — | **Measured, not inferred.** All **208** retained `go mod download` records carried **`Usage count: 1`** — not one was ever reused in ~a month of builds. The mechanism was isolated by four controlled builds: an unchanged tree rebuilt with the **same** `--build-arg VERSION` → `RUN go mod download` **CACHED**; the **same** tree with a **new** `VERSION` → **executed**. `COPY go.mod ./` stays CACHED either way, which is the tell: a `COPY`'s key is content-based, while a `RUN`'s key includes the stage **environment**, and `ARG VERSION`/`ARG GIT_COMMIT` are declared *before* the download in `felhom-controller/controller/Dockerfile`. Since every real build passes a fresh version, the layer is invalidated **every single time**. **`felhom.eu/hub/Dockerfile` has the identical defect** (`ARG VERSION`/`ARG BUILD_TIME` above `COPY go.mod go.sum*` → `RUN go mod download`) — and because both Dockerfiles produce byte-identical `buildx du` description strings, the 208 records are a COMBINED count and must not be attributed to one project. **Fix shape (one line each, not applied here):** move the `ARG VERSION`/`ARG GIT_COMMIT`/`ARG BUILD_TIME` declarations down to just above the final `go build`. **Worth more than the cap and the move combined** — the cap bounds the symptom, this removes the source. `build.sh`'s `rm -rf` + `cp -a` and its host-side `go mod tidy` were **ruled out by fingerprinting**: the tree is byte-identical across runs and `tidy` is a no-op | CC | | **R-209a** | Process & tooling | P4 | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC | @@ -379,7 +360,6 @@ stopping line that lies. | **R-315** | Process & tooling | P4 | **The wire-contract gate's positive control FAILS on the new wire: it checks name-presence, not decodability.** R-311 declared `hub -> agent (GET /escrow/retained)` as a fourth ROOT, and the gate's tag count rose 174 → 182, so the fields ARE inspected. But renaming the agent-side `superseded_at` json tag to `superseded_at_RENAMED` **still passed** — because the string `superseded_at` also occurs as a map key in the agent's local-API response, and the check is a repo-wide name search. The gate documents this ("name-reachability is not use"), so it is a known limit rather than a regression — but it means **declaring this wire bought documentation, not enforcement**, and a report that claimed coverage would have been wrong. The mutation was asserted to have applied before the run | **READY (M) — NEW 2026-08-12, RANK 3** | R-311 | Make the check resolve the RECEIVER'S mirror type and compare field-by-field, or state per-root which kind of check it got. **A gate whose positive control fails is an instrument nobody has calibrated** | CC | | **R-325** | Process & tooling | P4 | **The shared copy vocabulary is imported by ONE of its two consumers, and drift-checked into the other.** `customer_copy_vocab.py` is the single list; `hub_copy_gate.py` imports it. **`felhom-controller/controller/scripts/retrieval_promise_gate.py` still carries its own `STEMS` literal**, because the session that created the shared module was under a hard end-state requirement to leave `felhom-controller` untouched — its target box was being re-deployed the same evening. **Two copies of a word list is not a theoretical risk in this project: it is the R-299 defect exactly**, where a guard asserted one inflection of a Hungarian verb and the plural walked past it. **So the gap is instrumented rather than left open: `hub_copy_gate.py` READS the controller gate's `STEMS` and FAILS if the two disagree** — single-source semantics tonight without a cross-repo edit. **Watched failing:** removing one stem from the shared list produced *"the shared vocabulary is no longer shared"* with both lists printed, and restoring it returned the gate to green. An ABSENT sibling clone is **INCONCLUSIVE (exit 2), never a pass** — the G-1 lesson. **This is a scaffold, not the destination** | **READY (S) — NEW 2026-08-13, RANK 3** | R-299, R-324 | Make `retrieval_promise_gate.py` import `felhom.eu/scripts/customer_copy_vocab.py` and delete its own literal — a felhom-controller change of a few lines, needing no bake (a gate is not shipped code). Then the drift check becomes redundant and should be removed with it, rather than left as a second mechanism nobody re-reads | CC | | **R-327** | Process & tooling | P4 | **The standing picture still describes a defect that has been fixed twice over.** Found by the first run of `unproven.py` (R-326), which is the argument for having built it. `where-felhom-stands.yaml`'s `claim.code-naming` is `status: partial` and its title reads *"The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not show"* — **both halves of which are now false.** The box side shipped 2026-08-10 (R-295), the hub half and the page-naming fix on 2026-08-13 (R-295 hub, new `reenroll` mail kind), and the third near-homograph on 2026-08-13 (R-323). **NOT MOVED BY THIS SESSION, deliberately and by the dataset's own rule:** *"A status may not be RAISED here — if the evidence supports a stronger status than the capability map records, the MAP changes first and this file follows it."* Raising it here would be the exact inversion the file's header forbids, and the map edit is a separate judgement about what "walked" means for a naming change that no customer has yet met | **READY (S) — NEW 2026-08-13, RANK 4** | R-295, R-323, R-326 | Decide the capability-map status for the naming arc, then let the dataset follow it. **Note the honest difficulty: no customer has typed „Tulajdonosi jelmondat” yet**, so `walked` would be an over-claim; `built` is probably right, and the title needs rewriting either way because it describes a defect rather than a capability | operator + CC | -| **R-364** | Process & tooling | P4 | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC | | **R-377** | Process & tooling | P4 | **`CONTEXT.md`'s standing rulings are 188 KB in one section with no sub-headings, and that is why nobody reads them.** Measured 2026-08-22: `CONTEXT.md` is 217 KB, of which **187,913 bytes — 86% — is a single `## Standing rulings` section** carrying 39 `S-` ids and 153 bullets under **one** heading. **This session deliberately did NOT compress or split it**, and the reason is the ruling itself: `PROMPT-TEMPLATE.md` §3.4 and this file's own contract say the decision log is *dated, never edited afterwards*, and it is the only place that answers *"has this been proposed before, and why did we say no?"* — compressing it destroys exactly that. With no per-ruling delimiter, any mechanical split risks cutting a live ruling from its reason, which is the failure this whole arc is correcting. **So the problem is navigational, not volumetric, and the fix is structural: give each ruling a sub-heading with its `S-` id and date.** Then it can be linked, cited and found without a single word being edited. **13 mentions of SUPERSEDED already sit inside that blob** and cannot be separated from live text safely today. | **OPEN — LOW** | R-369 | Add per-ruling sub-headings only. Do not compress, do not reorder, do not edit any ruling's text. | CC | | **R-392** | Process & tooling | P4 | **No architecture document covers the two-AI workflow.** `documentation/architecture/` holds eight documents and **all eight cover the product** — topology, host agent, control-plane authorization, hub, off-site connectivity, backup, controller modules, capability map. Nothing records how the Claude.ai / Claude Code split works, what each side owns, how skills and `.claude/rules/` are scoped, or why. **The absence was found by trying to fill the template field, not by a survey:** the task that added the five process skills (2026-08-25) had to name an owning architecture document and could not, and the template requires that be recorded rather than passed over. The exposure today is low — the split is stable and both sides work — but it lives entirely in the operator's head and in chat, which is precisely the shape of a commitment nothing enforces. | **OPEN — LOW** | — | Write one architecture document for the agent-tooling layer: which AI owns which artifact class (`TASK-*.md`, `RUNBOOK-*.md`, validation), how skills are scoped and installed, what belongs in a `CLAUDE.md` versus a skill versus a rules file, and the reasoning for each boundary. **The rules themselves already exist** in `skills/felhom-doc-authoring/SKILL.md`; what is missing is the map of who owns what. Do not restate the doc-authoring rules there — point at that skill. | CC | | **R-394** | Process & tooling | P4 | **`felhom-build-deploy/SKILL.md` is 179 lines, over the 150-line limit its own repo now enforces.** Found 2026-08-25 by `scripts/check_skills.py` on its first run — the over-length was discovered BY the new checker, on the day the limit was written down, which is the checker working as intended. **It is not edited and not trimmed here:** the task that introduced the limit explicitly scoped the four pre-existing skills out, and trimming a build-and-deploy skill without exercising its commands is how a wrong command ships to a live host. **It is a named single-entry exception in `GRANDFATHERED` in `scripts/check_skills.py`, printed as a WARN on every run**, so it cannot fade; a NEW skill over the limit is convicted normally, and growing the set requires editing that file in a commit with a row to name. **The rationale for the limit** — attention thins across the excess, so the lines that matter are not the ones that survive — is in `skills/felhom-doc-authoring/SKILL.md` §5. | **OPEN — LOW** | — | Trim `felhom-build-deploy/SKILL.md` under 150 lines in a session that can VERIFY the commands it keeps, then delete its `GRANDFATHERED` entry in the same commit. The likely trim is the per-artifact command blocks moving behind a pointer to the runbooks, keeping the gotchas inline — but that is a judgement for a session with a build to run, not a line-count exercise. | CC | @@ -393,30 +373,23 @@ stopping line that lies. | **R-492** | Process & tooling | P4 | **[P3-LOW] `cfg.Paths.HDDPath` is empty on every box and still has readers; delete it.** R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, `systemInfo`, the same fallback. The global now carries no information on any box and its deletion was deferred twice. **Fix shape:** remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: dead-code cleanup; every reader already falls back correctly.** | — | — | CC | | **R-502** | Process & tooling | P4 | **[P3-LOW] The bootstrap regression harness is run by NO gate and NO CI — and it had never exercised the pairing banner.** MEASURED 2026-09-14: `grep -rn bootstrap-modes scripts/*.py .gitea/workflows` → nothing; the harness is run by hand in `felhom-iso-assistant:trixie`. While adding the R-496 checks, its fake hub's `register` reply turned out to carry no `pairing_code`, so `print_pairing_banner` returned early in every run: the banner that a household reads first had no test at all until ISO v1.27.0. **Fix shape:** register the harness in `repo_gates.py` behind a container-available check that reports INCONCLUSIVE (never skip-as-pass) when docker is absent, with a decoy. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: test harness coverage; no household meets it directly.** | — | — | CC | | **R-507** | Process & tooling | P4 | **[P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen.** MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): `qm sendkey 332 tab` did not move focus (both password copies landed in one field), `mouse_move 1237 772` + `mouse_button 1` did not move the cursor or press Next, while `alt-n` did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. **Fix shape:** measure QEMU `input-send-event` with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: release-test tooling; an operator click-through covers it.** | — | — | CC | -| **R-555** | Process & tooling | P4 | **[P3-LOW] The wire-contract gate counts a field as received when its name appears in a Go COMMENT on the receiving side.** FOUND 2026-09-17 by CC adding the report's `language` field (controller v0.247.0): `scripts/wire_contract_gate.py` passed WITHOUT an allowlist entry, because `receiver_tokens()` tokenises whole files and the word „language" occurs hub-side only in a comment (`hub/internal/web/configs.go:558`, „this page's existing language"). The shape is the gate's own named failure class (name-for-fact, R-421): any English tag name that also appears in hub prose passes unread. The field was allowlisted by hand with this row named. **Fix shape:** strip `//` and `/* */` comments (and template `{{/* */}}`) before tokenising; add the decoy „a tag whose name appears only in a receiver comment must convict"; expect a handful of currently-passing tags to surface — each is a finding, not noise. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: gate precision; tooling only.** | — | — | CC | | **R-564** | Process & tooling | P4 | **[P3-LOW] The retrieval-promise gate's Hungarian stems cannot see a SPLIT verb — „csak akkor állíthatók vissza", „hozod vissza" — so those Hungarian sentences were never scanned; the English translation exposed them.** FOUND 2026-09-17 by slice 1 release B (R-556): after the gate learnt English (`EN_PATTERNS`), seven English retrieval phrases on `backups_remote`, `backups_escrow`, `backups_restore` and `backups_restore_wizard` had NO Hungarian registration, because their Hungarian carries the verb particle after the verb („A távoli mentések csak akkor állíthatók vissza …", „a távoli mentések CSAK ezzel a kóddal állíthatók vissza", „csak a hiányzó fájlokat hozod vissza"). The stems (`visszaállíthat`, `visszaszerezhet`, `visszahozhat`, `visszanyit`) match only the joined form. The seven were registered in English with reasons (none is a false promise: two are preconditions, five describe the action on the same page). **Fix shape:** add split-form patterns to the Hungarian scan (`állíthatók? vissza`, `(hoz|szerez|nyit)\w* vissza`), register the Hungarian occurrences found, decoy with a planted split-verb promise. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: copy gate precision; the seven found sentences were reviewed and are true.** | — | — | CC | | **R-569** | Process & tooling | P4 | **[P3-LOW] Four more API handlers pick their status code by matching ENGLISH words in an error — the same shape as R-553, one language over.** FOUND 2026-09-17 while fixing R-553: `controller/internal/api/router.go` matches `"protected"`, `"not found"`, `"not deployed"`, `"still running"`, `"not orphaned"` in `err.Error()` at the stop/start, remove and orphan-cleanup handlers (three separate blocks). These strings are internal English, so localisation does not move them — the risk is a reworded internal error, not a translation, which is why this is P3 and was NOT folded into R-553's release. **Fix shape:** the same `util.KindErrorf` sentinels in `internal/stacks` (`ErrProtectedStack`, `ErrStackNotFound`, `ErrStillRunning`, …), a `statusFor` helper per handler family, and one table test per family passing a reworded message. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: internal robustness; only a future rewording would break it.** | — | — | CC | -| **R-571** | Process & tooling | P4 | **[P3-LOW] The off-site failure classifier and the dashboard's alert-placement rules are described in no architecture document.** FOUND 2026-09-17 while fixing R-553: `07-backup-architecture.md` and `02-controller-module-map.md` grep clean for `ClassifyOffsiteFailure`, `Inline` and `PageOnly`, so the six failure classes (quota / orphaned / no-repo / no-units / transport / unknown), the head lines they pick and the rule that one warning renders inline under the storage bars while every other renders in the top banner exist only in code. `10-localisation.md` §9 now names the SIGNALS each decision reads; the behaviour itself still has no home. **Fix shape:** a short section in `07-backup-architecture.md` for the classifier (its classes, what each means for the customer, and that restic/ssh text signatures are external) and one in `02-controller-module-map.md` for alert placement. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: documentation only.** | — | — | CC | | **R-574** | Process & tooling | P4 | **[P3-LOW] `web/handler_debug.go` mixes page copy with JSON payload, so neither half could be converted safely.** FOUND 2026-09-18 by localisation slice 2 release A (R-557): the file holds 39 Hungarian literals and the inventory classifies them by STATEMENT, not by data flow (`I18N-INVENTORY-2026-09-17.md` §4), so which are section headings the debug page renders and which are values inside a diagnostic dump the operator copies out is not established. Converting a dump value would change what an operator pastes into a report; leaving a heading Hungarian leaves a half-English page. **Fix shape:** walk the file once and label every literal page-copy or payload in the same table slice 2 release A used, then convert only the page-copy half. Belongs to slice 2 release B or C. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: debug page is operator-facing.** | — | — | CC | | **R-576** | Process & tooling | P4 | **[P3-LOW] `i18n_go_parity.py` cannot see a call site that LOST text; it only checks that the text a key carries is real.** FOUND 2026-09-18 by localisation slice 2 release B, and found the hard way: the bulk converter silently dropped the continuation of a multi-line concatenation (`fmt.Errorf("a: "+ "b: %s", x)` kept only `"a: "`), damaging **7** producers — and the gate stayed GREEN throughout, because every surviving fragment WAS a byte-equal base-commit literal. Its question ("is this text real?") was answered yes while the CALL had lost half its sentence and its arguments. Two behaviour tests caught it (`TestR356_ScenarioC_UndeployedAppIsStillRefused`, `TestR379_ScenarioA_RollbackSucceeds_AppComesBack`), because they assert the sentence a customer READS. **Fix shape:** the gate learns a second question — for every `util.MsgError("key", …)` call site, the count of its arguments must equal the count of printf verbs in the key's Hungarian value, and no key-naming literal may be adjacent to a `+`. Both are cheap and would have convicted all 7. **The general lesson, worth keeping whatever is built: a structural gate over the TEXT cannot see a defect in the CALL.** | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: gate improvement; tooling only.** | — | — | CC | | **R-579** | Process & tooling | P4 | **[P3-LOW] Five page shells loaded `style.css` with NO cache-buster, so a browser holding an older copy kept being served CSS that did not know about the newest UI.** FOUND 2026-09-18 by the operator's screenshot of the v0.254.0 language globe: it rendered as a bare, unstyled `
` — a stray triangle and two plain words outside the card — because `login.html`, `claim.html`, `recovery.html`, `launcher_shared.html` and `launcher_share_password.html` requested `/static/style.css` with no `?v=`, while `layout.html` has used `?v={{.Version}}` since v0.166.0. **`.Version` was also absent from three of those five data maps.** FIXED AND CLOSED in the same session (controller v0.255.0): the parameter on all five, and `Version` set once in `executeTemplateLang` so a new shell cannot miss it; `TestGlobeOnAnonymousShells` now refuses an absent or EMPTY `?v=`. **The general form, which is the part worth keeping: a template that loads a versioned asset WITHOUT its version is invisible to every test that reads markup — the markup is correct and the browser fetches the wrong file.** A gate over "every stylesheet/script link in a template carries `?v=`" would catch the class; not built, because 8 first-boot-wizard templates would fail it and R-554 deletes them. | **DEFERRED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the gate "every stylesheet/script link carries `?v=`" was deliberately not built until R-554 deletes the old wizard; nothing else owns it) — **CLOSED 2026-09-18 - controller v0.255.0** | — | — | CC | | **R-591** | Process & tooling | P4 | **[P3-LOW] `Stack.Copy()` is a deep copy with one shallow field, and the field is new.** FOUND 2026-09-20 while adding the catalog's English overlay: `controller/internal/stacks/manager.go` `Copy()` deep-copies `Meta.DeployFields` (with nested `Options`), `Meta.OptionalConfig` (with nested `Fields`), `Meta.Integrations`, `Meta.HealthCheck` and `Meta.InitialCreds` — and does NOT copy the new `Meta.I18n` map, which the struct assignment leaves shared between the original and the "copy". **It is safe TODAY and that is exactly the shape worth filing:** `Metadata.For` reads the overlay and never writes to it (pinned by `TestForDoesNotMutateTheReceiver`), so nothing can observe the sharing yet. The hole is in the CONTRACT — a function whose whole purpose is "a snapshot the caller may mutate" now has a field that is not one, and the next person to write through an overlay will find a bug with no failing test in front of it. **Fix shape:** deep-copy `I18n` in `Copy()` and pin it with a test that mutates the copy's overlay and asserts the original is unchanged. Alternatively state in `Copy()`'s comment that `I18n` is deliberately shared and immutable, and pin THAT with a test. Either is fine; silence is not. | **READY - rank P3-LOW; owner: CC (controller)** **Re-ranked 2026-10-03: P3->P4: latent contract gap, safe today.** | — | — | CC | -| **R-594** | Process & tooling | P4 | **[P3-LOW] The catalog copy gate can CONVICT a retrieval promise but has no way to REGISTER a true one.** FOUND 2026-09-20 translating batch 3 (R-560 slice 5). Vaultwarden's invite step and its sign-up setting both ended „…can open an account", and the English retrieval-promise pattern reads `can … open` as the claim that sealed backups can be opened. **The conviction was a FALSE POSITIVE** — opening an account is not opening a backup — and the two sentences were reworded to „can sign up", which is also the better copy, so nothing is blocked today. **The gap is structural.** The shared vocabulary this gate copies (`scripts/customer_copy_vocab.py`) states the design explicitly: these stems are NOT banned, because each carries a claim that is sometimes TRUE, and *"an occurrence must be REGISTERED with a reason in the consuming gate's allowlist"*. The hub gate has `ALLOWLIST_EN`; `app-catalog-felhom.eu/scripts/check-copy-i18n.py` has none, so the only ways past it are to reword or to bypass the gate — and a catalog app whose English genuinely says a file can be restored (a backup app, a versioned document store) has no honest third option. **Fix shape:** an `ALLOWLIST_EN` of `(app, path, reason)` in the gate, a decoy proving a REGISTERED occurrence passes and an unregistered one still convicts, and a check that every entry still matches something (a stale allowlist entry was R-299's shape, and the retrieval gate has gone red on stale entries before). | **READY - rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3->P4: gate tooling; nothing blocked today.** | — | — | CC | | **R-603** | Process & tooling | P4 | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `'`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.** **Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** **Re-ranked 2026-10-03: P3->P4: test tooling.** | — | — | CC | -| **R-605** | Process & tooling | P4 | **[P3-LOW] A catalog gate that REFUSED TO RUN and a gate that ran and could not decide print the same word, so a reader cannot tell which happened.** FOUND 2026-09-21 while answering why the chaos night's update round could not run. On 2026-09-17 `check-image-resolvable` and `check-volume-persistence` both returned INCONCLUSIVE and the drawn `update` action was replaced with `use` (`audits/DRILL-chaos-night-2026-09-17.md:692-695`). **Neither script is defective — they behaved exactly as designed**, and both headers say why: a detector that cannot prove itself must refuse to report rather than guess (`check-image-resolvable.py` cites the 2026-07-21 incident where a Docker Hub throttle read as 24 of 65 pins falsely dead). **What is missing is the DISTINCTION.** `check-volume-persistence.py`'s `self_test` refuses to evaluate ANY app when it cannot build its canary image — a harness-level refusal — while `classify()` returns a per-app UNDETERMINED for an app that wrote nothing; `check-image-resolvable.py` likewise separates a harness-level canary failure (exit 2 at `check()` L180-183) from a per-pin throttle (L121-128). **`catalog_gates.py`'s VERDICT map collapses all of them into one `INCONCLUSIVE` label**, so the operator-facing summary cannot say whether the gate ran at all. **The cost is real and already paid:** no raw stdout of the 2026-09-17 run survives in either evidence directory, so the exact triggering path is INFERRED from the code plus the documented throttle precedent, not observed — a distinct summary line would have recorded it for free. **Fix shape:** have each gate's exit distinguish "the harness refused" from "the result is undetermined" (a third exit code, or a marker line the runner matches), and have `catalog_gates.py` print the two differently. **Ships with a decoy each way (R-421): a run whose canary fails must NOT read as a per-app undetermined, and vice versa.** Small. | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3->P4: tooling report clarity.** | — | — | CC | | **R-624** | Process & tooling | P4 | **[P3-LOW] Three of the catalog's apps cannot be seeded by ANY headless route, and for two of them that is a deliberate security decision — so the upgrade harness has a permanent ceiling nobody has written down.** FOUND 2026-09-21 while widening R-462 from 3 apps to . **`vaultwarden` and `zipline` close self-registration ON PURPOSE** — vaultwarden by `SIGNUPS_ALLOWED=false` (R-512, *„a stranger who guesses vault. must not be able to register"*), zipline by answering `E1037: User registration is disabled` — and neither ships a CLI that could make an account instead. So there is no route to a first account without the admin secret, and **that is correct**: the harness must not be the reason a customer-facing app accepts strangers. **`gitea` is a different and fixable case:** the template sets no `INSTALL_LOCK`, so a fresh instance sits in its web-installer state and `gitea admin user create` refuses (`MustInstalled() [F] Unable to load config file for a installed Gitea instance`); POSTing the installer form first would work and was simply not written tonight. **Why this is a row rather than three notes:** `09` §3 decision 6 says the upgrade test goes to **all** apps, and R-462 is costed as if every app is reachable given enough fixture work. It is not. There is a class — *apps whose only account-creating route the catalog deliberately closes* — for which the honest maximum is `inconclusive` unless the harness is given the app's admin secret at deploy time, which is a decision nobody has taken. **Needs:** the class named in R-462's scope so the remaining count is honest; a decision on whether the harness may hold an app's admin secret (it already holds the ones IT generates — see R-619); and, separately and cheaply, a gitea installer-form fixture. Evidence: `audits/update-night-2026-09-21/apps/{vaultwarden,zipline,gitea}/verdict.json`. **— NIGHT 2026-09-23:** `gitea` stayed inconclusive on both venues (its installer); `vaultwarden`/`zipline` not attempted (closed sign-up, by design); `code-server`, `outline`, `rallly` have no front-door seed route (listed, not moved); `bentopdf`, `glance`, `crafty-controller`, `wger`, `wanderer` (meilisearch) and `uptime-kuma` have no fixture tonight (listed, not moved). **-- 2026-09-30: the ceiling was WRONG for three of the six named apps.** outline has a front-door first-run route (`POST /api/installation.create` — its self-hosted setup, refused once a team exists), rallly signs up through its own better-auth API with the e-mail code read from its own table in place of a mailbox, and zipline 4's first-run route is `POST /api/setup` (the fixture had tried `/api/auth/register` and `/api/auth/setup`, which are not it). Fixtures for outline and rallly are in `upgrade_fixtures_box.py` and both apps moved on both venues; zipline's fixture now tries `/api/setup` first (measured on the bench: a SUPERADMIN made, the login works). **What remains in the class:** vaultwarden (closed sign-up by design, no CLI), code-server (a browser IDE). `audits/pg-last-six-2026-09-30/`. **-- 2026-09-30 (evening): gitea's fixable case is DONE** — the fixture posts gitea's own first-run installer form (the form's own defaults + an admin; no CSRF on that form), measured on the bench and through the setup gate on 9202; gitea 1.27.0 → 1.27.3 published `d7ba60c`. What remains in the class: vaultwarden, code-server. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (catalog harness)** **Re-ranked 2026-10-03: P3→P4: an upgrade-test coverage limit; no household meets it.** | — | — | CC | | **R-652** | Process & tooling | P4 | **[P3-LOW] The memory watch counted the kernel's file cache as the app's memory.** MEASURED 2026-09-23 night on the bench: nextcloud 34.0.4 read **100 %** of its 1 GiB and immich's PostgreSQL **100 %**, each with **0** kernel `oom_kill`s — `memory.peak` includes page cache, which the kernel drops before it kills anything. Under the watch as built (R-635 follow-up) both would be marked `memory_tight`, and the gate would demand a raised `mem_limit` — a customer-box capacity figure — for cache. **Done the same night (09 §3 decision 22, CC — operator may reverse):** the watch samples the app's own memory (`anon` of `memory.stat`) every 15 s; the mark and the ladder's `memory_peak_pct` read it where measured; the cgroup peak stays beside it (`memory_cgroup_peak_pct`). **Open:** romm's backfilled entry carries M1's 80.9 % cgroup peak (measured before the anon sample existed) — re-measure it on its next step; and decide whether an app whose anon is low but whose cgroup stays pinned at its limit (cache thrash) should be marked at all. Evidence: `audits/night-2026-09-23/apps/nextcloud/bench-1024M/`, `apps/immich/bench-noanon/`. | **READY — P3; owner: CC (catalog harness)** **Re-ranked 2026-10-03: P3→P4: the main fix is done; what is left is harness tuning, no household meets it.** | — | — | CC | | **R-693** | Process & tooling | P4 | **[P3-LOW] The memory watch marks a Node app `memory_tight` at any limit — its heap sizes itself from the limit.** Measured 2026-09-25 on the bench (docmost 0.96.0, harness v4): the app's own memory (`anon`) peaked at **349 MB of 384 MB (90.9 %)**, then, with the limit raised to 512 MB, at **431 MB of 512 MB (80.4 %)** — 0 OOM kills and 0 restarts in both 10-minute watches (~12 000 requests each). So the mark (decision 22's "does not fit the memory") fires for an app that fits, and the gate's remedy (raise the limit) cannot clear it. docmost moved with the limit raised to 512 MB (decision 39). **Needs:** a basis that tells growth-to-fill from pressure (e.g. kills/restarts plus a GC-pressure signal, or a second watch at a higher limit showing the peak scales), or a per-app `memory_scales_with_limit` fact. `audits/night-2026-09-26/C/bench-run1/`, `…/bench-run2/` | **OPEN — P3; owner: CC** **Re-ranked 2026-10-03: P3→P4: a harness judgement problem; no household meets it.** | — | — | CC | | **R-731** | Process & tooling | P4 | **[P3-LOW] The catalog-currency comparison cannot see an upstream that changes its TAG SHAPE, and two newest tags need a release check.** MEASURED 2026-09-30 (`audits/catalog-currency-2026-09-30.md` §2 item 5): the same-shape rule read three apps as up to date that were not — gramps-web (`v25.6.0` → upstream dropped the `v`, at `26.9.1`), jellyfin (`10.11.11` → two-part `12.1`), kimai (`apache-2.57.0` → plain `2.67.0`; the plain tag's digest equals `apache`'s, so kimai moved on 2026-09-30). A control pass over every shape caught them. Not checked: whether `mariadb:13.0` and `gitea/gitea:28.0.0` (pushed 2026-09-30 00:15 UTC) are general releases. **Needs:** the currency script's shape-switch control made standing (it is in the audit's tools today), and the two release checks before either is walked. **-- 2026-09-30: the two release checks answered** (`audits/immich-first-start-2026-09-30/D/D2-release-checks.txt`): **gitea v28.0.0** is a general release (GitHub: not prerelease, not draft, 2026-09-29) — upstream renumbered 1.27.x → 28; **mariadb 13.0** is `Stable` but a short-term `Rolling` line (no EOL date), while 12.3 and 11.8 are the LTS lines — so a move of any of the four MariaDB apps to 13.0 would leave LTS. Not done: the shape-switch control made standing. | **NARROWED 2026-09-30 — only the standing shape-switch control is left; owner: CC** **Re-ranked 2026-10-03: P3→P4: only a standing check in a tooling script is left.** | — | — | CC | | **R-739** | Process & tooling | P4 | **[P3-LOW] The test bench cannot run wanderer at all: its web server calls the database at the PUBLIC name `https://.`, which the bench has no name or TLS for.** MEASURED 2026-09-30 (more-night-apps brief, Part B): on bench 9401 the catalog template (v0.20.0) came up `wanderer-db` healthy, `wanderer-search` healthy, `wanderer` **unhealthy** for 12 min, every page 500 („Error 0: Something went wrong"); `PUBLIC_POCKETBASE_URL=https://hike-db.gate.invalid`. So the harness can only ever answer `inconclusive` at FROM for wanderer, never about an update. Its step `v0.20.0 → v0.21.0` (web + db) and meilisearch `v1.36 → v1.54` stay untested; **the brief's question — does the search index survive meilisearch's move or is it rebuilt — is NOT measured.** **Needs:** a bench venue that gives the stack the DB name (an `extra_hosts` + plain-http override in the bench's render only, stated in the verdict), or wanderer proven on the box venue alone with the operator's word. `audits/more-night-apps-2026-09-30/B/wanderer-probe.txt` **-- 2026-09-30 late: the bench CAN run wanderer now** — `upgrade-test.py` `BENCH_ENV_OVERRIDES` points `PUBLIC_POCKETBASE_URL` at the database container on the bench only (the only address the web image reads, measured in v0.20.0 and v0.21.0), recorded in every verdict: web 200, sign-up (`PUT /api/v1/user`) 200, login 200. **The meilisearch question, answered:** v1.54.2 REFUSES v1.36's database („Your database version (1.36.0) is incompatible”); with `MEILI_UPGRADE_DB=true` it upgrades in place and all three indexes (actors, lists, trails) are there — so wanderer's step needs that switch in the template first. Not done: a fixture (creating a list answered PocketBase's „Failed to create record” — likely a rule for unverified users), whether a household's record survives in the index, the step itself. `audits/night-rulings-2026-09-30/` | **NARROWED — the bench runs it; the step needs a fixture and `MEILI_UPGRADE_DB`; owner: CC** **Re-ranked 2026-10-03: P3→P4: test bench coverage; the template switch and fixture are tooling work.** | — | — | CC | | **R-759** | Process & tooling | P4 | **[P3-LOW] wger's onboarding record (the checklist pilot) keeps rows open that no other row owns.** 2026-10-01, `app-catalog-felhom.eu/onboarding/wger.md`: **2.5** no backup → remove → restore → read back of wger exists — the box has no per-app backup press outside an Update (R-648) and wger has no newer step to carry one; **3.7** changing the password and adding a family member not measured, and the template has no `add_people` text; **6.3** no forced-fail undo for wger; **8.2** the app page not read on 9202 this session; **9.1** the runtime volume-persistence gate not re-run (last CLEAN 2026-08-02). The other open rows have their own: 1.5 (R-755), 1.7/2.8 (R-762), 3.4 (R-763), 7.1 (R-764). wger is exempt from the onboarding gate (published before the checklist), so nothing blocks; this row is what keeps the record honest. **Needs:** the five measured on 9202 — 2.5 and 6.3 ride wger's next ladder step (the update's backing-up phase is the per-app backup). `audits/new-app-checklist-2026-10-01/` | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: record-keeping for a hidden app; nothing blocks.** | — | — | CC | -| **R-781** | Process & tooling | P4 | **[P3-LOW] The catalog's `scripts/test_gate_decoys.py` fails 4 of its own "genuine" onboarding cases — on the untouched tree.** Measured 2026-10-01 (same 4 FAIL lines before and after this session's checklist edit): the harness copies the catalog into a temp tree with a fake sibling `felhom.eu`, and the REAL published records (radicale, karakeep, dawarich) name evidence that the fake sibling lacks, so a complete record reads incomplete. The gate itself (`onboarding`) is green on the real tree. **Needs:** the genuine cases build their own records, or the fake sibling mirrors the cited evidence. | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: a test-harness self-check; the real gate is green.** | — | — | CC | | **R-786** | Process & tooling | P4 | **[P3-LOW] SparkyFitness's onboarding record has six open rows** (`app-catalog-felhom.eu/onboarding/sparkyfitness.md`): 0.5 runtime internet (food search providers), 0.7 the phone app's sign-in route through traefik, 1.6 the env names the server reads, 1.7 the entrypoint read, 5.4 a second memory watch at another limit, 8.3 no logo/screenshots on felhom.eu (404). Everything else measured this session (bench + 9202). **Needs:** each row measured, or n/a with a reason. | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: onboarding record completeness; no household meets it directly.** | — | — | CC | | **R-805** | Process & tooling | P4 | **[P3-LOW] The volume-persistence gate judges an empty NAMED volume (R-788) but not an empty BIND mount.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): Grimmory's `/app/data` bind held 0 files and read CLEAN (its state is in MariaDB); komga, paperless-ngx, radarr and sonarr also carry empty binds that are not judged. **Needs:** decide whether an empty bind after the exercise is UNDETERMINED like a named volume, with a decoy. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a test-gate rule; no household meets it.** | — | — | CC | -| **R-806** | Process & tooling | P4 | **[P3-LOW] The gate's GET exercise speaks plain http to an HTTPS backend (crafty-controller :8443), and gramps-web did not answer on :5000.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): both reached data only through their fixture seed. **Needs:** read the `loadbalancer.server.scheme` label (upgrade_boxport already does); find why gramps-web's first GET got no answer. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a test-gate defect; no household meets it.** | — | — | CC | +| **R-806** | Process & tooling | P4 | **[P3-LOW] The gate's GET exercise speaks plain http to an HTTPS backend (crafty-controller :8443), and gramps-web did not answer on :5000.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): both reached data only through their fixture seed. **Needs:** read the `loadbalancer.server.scheme` label (upgrade_boxport already does); find why gramps-web's first GET got no answer. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a test-gate defect; no household meets it.** **NARROWED 2026-10-05 (app-catalog `4828dc7`): the harness GET now uses the traefik scheme label (https + -k for crafty-controller). LEFT: gramps-web :5000 not answering — a live look on a box.** | — | — | CC | | **R-807** | Process & tooling | P4 | **[P3-LOW] 13 apps stay UNDETERMINED because their seeds write only to the database, so their upload / media / cache / redis volumes stay empty.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): claper, crafty-controller, dawarich (redis), docmost, gramps-web, immich (ML cache), outline (+redis), sparkyfitness, tandoor, vikunja, wger, wishlist, zipline — plus plex (no seed route: a plex.tv claim token) and wanderer (no fixture; unhealthy under the gate). None is BROKEN: nothing was written outside a preserved folder. **Needs:** per app, a seed that uploads one file (or the volume named n/a with a reason: a redis cache, an ML model cache). | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: test coverage; nothing was found broken.** | — | — | CC | -| **R-819** | Process & tooling | P4 | **`scripts/check_stands.py` is red and runs in no runner.** Measured 2026-10-03 on `9e2786c` (before the triage): it convicts `where-felhom-stands.yaml` for citing R-273 and R-356, which are in neither `OPEN-ITEMS.md` nor anywhere it reads. After the triage it also convicts R-281, R-198 and R-201, because its rule 3 reads only `OPEN-ITEMS.md` and those rows are closed. It is in neither `repo_gates.py` nor CI, so nobody saw it — the R-29 shape. | **READY — filed 2026-10-03 (triage); owner: CC.** Let rule 3 accept an id in `CLOSED-ITEMS.md` (and check the stand's status agrees), fix the two dangling ids, then register it in `repo_gates.py` with a decoy. | — | — | CC | -| **R-857** | Process & tooling | P4 | **Baking a golden twice under the SAME version leaves two stale facts.** 2026-10-04 (golden 0.292.0 re-baked with live-restore): (1) `golden_currency_gate.py` keeps reporting the FIRST bake directory's sha (`golden-0.292.0-2026-10-04`, d6cf8b33…) — it picks one of two directories with the same version, not the newest; (2) the publish replaces the package in place (pre-delete + upload), so until the operator re-vouches, the hub vouches a sha the registry no longer holds — a fresh install in that window fails its sha check (fail-closed; none ran; re-vouched 17:48). Fix direction: the gate prefers the newest bake directory (or refuses two for one version); the runbook says "re-vouch at once after a same-version re-bake". `documentation/tests/golden-0.292.0-2026-10-04-rebake/` | **READY — owner: CC** | — | — | CC |