diff --git a/documentation/audits/burndown-2026-10-05/r262-red-proof.txt b/documentation/audits/burndown-2026-10-05/r262-red-proof.txt new file mode 100644 index 00000000..b469d4c5 --- /dev/null +++ b/documentation/audits/burndown-2026-10-05/r262-red-proof.txt @@ -0,0 +1,6 @@ +### R262-a: drop skipped from knownUnmodelled +--- FAIL: TestR262_RestoreTestFieldsAreAKnownSubset (0.00s) +FAIL +### R262-b: the hub stops decoding source_tier +--- FAIL: TestR262_RestoreTestFieldsAreAKnownSubset (0.00s) +FAIL diff --git a/documentation/audits/burndown-2026-10-05/r263-red-proof.txt b/documentation/audits/burndown-2026-10-05/r263-red-proof.txt new file mode 100644 index 00000000..c8b735a1 --- /dev/null +++ b/documentation/audits/burndown-2026-10-05/r263-red-proof.txt @@ -0,0 +1,9 @@ +### R263-a: ClearBackupTarget writes true +1700: s.StoragePaths[i].BackupTarget = true +--- FAIL: TestR263_OnlySetBackupTargetGrantsTheRole (0.24s) + r263_backup_target_writers_test.go:73: BackupTarget may be granted outside SetBackupTarget at: [../settings/settings.go:1700:3] +FAIL +### R263-b: a composite literal BackupTarget: true in another package +--- FAIL: TestR263_OnlySetBackupTargetGrantsTheRole (0.28s) + r263_backup_target_writers_test.go:73: BackupTarget may be granted outside SetBackupTarget at: [../web/zz_r263_decoy.go:5:50] +FAIL diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 850acd2a..1d33ca43 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -62,6 +62,20 @@ The full text of every row below: `git show ab2b3049:documentation/backlog/OPEN- | **R-818** | **Two changelogs cite register ids for other findings.** (P4) | CLOSED 2026-10-05 — CORRECTED | Dated correction notes under hub v0.109.0 (`hub/CHANGELOG.md`) and controller v0.224.0 + v0.225.0 (`felhom-controller/CHANGELOG.md`): those two findings never had register rows of their own — the triage's „the real ids are in CLOSED-ITEMS" was itself wrong (no closed row names hub v0.109.0 or controller v0.224.0/v0.225.0). Nothing renumbered. | | **R-755** | **[P3-LOW] wger runs Django's DEVELOPMENT server in production: `manage.py runserver`, because the template does not set `WGER_USE_GUNICORN=True`.** (P3) | CLOSED 2026-10-05 — DUPLICATE of R-762 (its unique fact moved there) | Still true: templates/wger/docker-compose.yml has no WGER_USE_GUNICORN (grep empty). R-762 (open, read) states 'Owner decides together with R-755 (same server question)' and its fix names 'the gunicorn switch of R-755'. | | **R-446** | **[P2-MEDIUM] „Naprakész" can be FALSE, and the badge that says it cannot tell.** (P3) | CLOSED 2026-10-05 — DUPLICATE of R-440 (its unique fact moved there) | felhom-controller/controller/internal/stacks/updateorder.go:96: `if len(s.CatalogDigests) == 0 // s.CatalogTestedAt.IsZero() { return false }` — blind only for apps with no ladder entry, i.e. the same 15 templates R-440 lists (app-catalog has no update_ladder for them). Both rows close by the same act: each app's first proven ladder step (R-462). | +| **R-799** | **[P3-LOW] The MeTube fixture's `POST /add` leaves out `download_type`, which upstream's validator lists as required.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | app-catalog `29ac711`: `MeTube.add_body()` sends `download_type: video`; `scripts/test_upgrade_fixtures_metube.py` (red-proof: the field removed → FAIL). Not exercised on a box (the next MeTube step will). | +| **R-761** | **[P3-LOW] The canonical example template tells a new app's author the logo is `-logo.webp`; the controller loads `-logo.svg`, then `.png`.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | app-catalog `29ac711`: the canonical template comment, `REUSE.md` and `NEW-APP-CHECKLIST.md` name `{slug}-logo.svg` then `.png` (controller `config.go` AppLogoURL/AppLogoPNGURL). Comment-only. | +| **R-391** | **Gate 11 (observations) is registered in three of the four runners; `app-catalog-felhom.eu` is the exception.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | app-catalog `29ac711`: `CLAUDE.md` states its `REPORT.md` carries no observations section by convention (the row's second option); the shared gate was not copied. | +| **R-291** | **CI's installability assertion is now BOUNDED by a retention number, and the narrowing is recorded here so it can be widened deliberately rather than discovered.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom-agent `d833163`: `scripts/retention-policy.json` names the source of its 10 — the R-267 newest-10 prune, established 2026-08-10 (R-287) — and drops the non-existent `registry-retention.md` reader; `check-published-versions.py` still reads 10 (checked). The min_agent-floor bound stays recorded in the file as the better bound. | +| **R-348** | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom-agent `d833163`: `internal/backup/store.go` says a restart blanks the reported backup list until the next run; only the hub's verdict (7-day look-back) is unaffected. Comment-only. | +| **R-263** | **C7 — „This is the ONLY writer of `StoragePath.BackupTarget`" is false, and nothing pins it.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom-controller `114ff27`: comment „the only writer that GRANTS"; `internal/settings/r263_backup_target_writers_test.go` scans every non-test file under internal/ and cmd/ (assignments and composite-literal keys). Red-proofs: ClearBackupTarget writing true; a `BackupTarget: true` literal in internal/web — both convict (`audits/burndown-2026-10-05/r263-red-proof.txt`). | +| **R-368** | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom-controller `114ff27`: the `IsDefault` comment says the deploy FORM pre-selects it and the deploy API applies no default (the row's second option; behaviour unchanged on purpose). | +| **R-418** | **`repo_gates.py`'s docstring listed ELEVEN gates while THIRTEEN were registered** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `repo_gates.py` docstring lists all 17 gates; `scripts/test_repo_gates_docstring.py` asserts list == GATES in order (red-proof: one line removed → FAIL), run on every push by the script-tests gate. | +| **R-345** | **`hub/Makefile` tags and pushes `:latest`, which the project's own rules forbid in two places.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `hub/Makefile` AND the real release script `scripts/build-hub.sh` (3 lines; the build dir links to it) no longer tag or push `felhom-hub:latest` — nothing pulls it (grep of all repos + homelab-manifests). `scripts/test_no_latest_push.py` (walk of hub/ + scripts/; red-proofs: the old Makefile and the old build-hub.sh each FAIL). Whether a stale `:latest` sits on the registry was not checked. | +| **R-416** | **`closed_register_gate.py` still has no within-register duplicate-id rule.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `closed_register_gate.py` RULE 4 refuses an id twice in CLOSED-ITEMS.md (OPEN duplicates were already `register_shape_gate.py` RULE 3); 0 duplicates existed, so it registered green. Decoy `closed-register/duplicate-closed-id` (red-proof: RULE 4 off → LIVE HOLE). | +| **R-261** | **C6 — `CountSelfBindTokens` exists so that callers can assert an invariant, and no production caller asserts it.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `hub/internal/store/selfbind.go` names `CountSelfBindTokens` a test accessor and the two tests that pin the auto-mint invariant. Comment-only; no hub release needed. | +| **R-262** | **C7 — a comment claims a cross-repo contract is mirrored „field-for-field" and „the key-set tests guard drift"; it is two fields short, AND THE FIXTURE THE TEST READS OMITS THE SAME TWO FIELDS.** (P3) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: the hub comment says hostRestoreTest is a deliberate SUBSET; `hub/internal/api/r262_restoretest_subset_test.go` pins the hub fields and the known-unmodelled agent fields and cross-checks the agent source beside it — which found a THIRD unmodelled field, `skipped` (agent v0.133.0, R-672; by design a skipped test reads as failed with its reason). Red-proofs: drop `skipped` from the list / stop decoding source_tier → FAIL. | +| **R-286** | **A control drawn from the same channel as the measurement cannot detect a defect in that channel — and this one passed while the measurement was wrong.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: workspace standing rule 3 (both `CLAUDE.md` copies, identical) adds: a control must come from a DIFFERENT channel than the measurement; a hub-state check copies `hub.db-wal` or asks the running pod. | +| **R-588** | **[P3-LOW] ISO release records live in two different places, so "was the gate run for this image?" cannot be answered by looking.** (P4) | CLOSED 2026-10-05 — FIXED (burn-down Part B) | felhom.eu this commit: `runbooks/iso-release-gate.md` names `documentation/tests/iso-release--/` as the one home; `tests/iso-release-1.28.0-2026-09-16/README.md` points at the 1.28.0 record inside the 2026-09-16 audit. | --- diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 47b2db52..e47a04b5 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -192,7 +192,7 @@ stopping line that lies. | **R-687** | App updates | P4 | **[P3-LOW] Part 7's live proof has four gaps a scratch box cannot close, and one observability gap.** (1) **W+5h reached with steps left** is proven by unit test only (`TestLeg_NoStepAtOrAfterW5h`) — the leg starts at W+105m and would need a 3-hour leg live; (2) **the off-site leg FAILING** before the update leg: 9202 has no off-site tier, so only the "no target" path ran live — failure and panic are `TestChainUpdateLeg_EveryPath`; (3) **a `files_may_change` step WITHOUT a whole copy**: both drill apps given the mark (wishlist, navidrome, romm) turned out whole on 9202 by the backup side's truth table (why, per app, is not logged — see the observability gap), so only "with a copy" ran live; (4) **the full-system gate waiting** cannot run on 9202 (no agent), and **did not occur on the demo boxes' real night either** (2026-09-25: both legs ended by 04:19, before the gate opened at 04:30, and no whole-box backup was due on either box) — unit + red-proof only (`TestD20_GateWaitsForTheLeg`). **Also cosmetic:** a leg with no steps reports `"steps": null` to the hub, not `[]`. **Observability:** when the leg TAKES a `files_may_change` step it does not log which whole copy allowed it (only the skip says why). `audits/night-2026-09-25/C/` **-- NARROWED 2026-09-25 (controller v0.273.0):** the cosmetic `"steps": null` → `[]` and the taken `files_may_change` step's missing log line are FIXED (red-proofed, `audits/night-2026-09-26/F/`). Items (1)–(4) stay; (4) did not occur on 2026-09-25 either (demo-felhom's whole-box backup ran at 07:29, three hours after its leg; demo-hp had none due). **-- 2026-09-28 (night 27/28):** (4) did not occur again — on demo-hp the leg ended 04:23:54 and the whole-guest backup began 04:37:06, after the gate opened at 04:30; demo-felhom's backup ran at 07:36 (`audits/evidence-golden-0276-2026-09-28/phaseD2-night-read.txt`). **-- 2026-09-30 (by day, demo-hp 9201): item (4) PROVEN LIVE.** The night chain pressed by hand, the window moved to W = now − 2h05m the moment the leg started, `quiesce.poll_interval` 1m: `[quiesce] full-system backup due and inside its window, but the automatic update leg is running … deferring` at 11:35:11 and 11:36:11 UTC while bookstack (55.1 s) and kimai (75.1 s) stepped; the leg's end line at 11:36:29; the backup quiesced at 11:37:11 (the first poll after), job done 11:47:19, the agent's `backup: completed` 9.98 GB. Config and window put back and read back (`audits/pg-last-six-2026-09-30/C/`). **Found, cosmetic, manual chain only:** the deferral names the moved window's W+5h (16:29) while the manual leg's own deadline was its start + the leg length (16:49). | **OPEN — P3, gaps (1)–(3) + the manual-chain deferral text; item (4) proven live 2026-09-30; owner: CC** **Re-ranked 2026-10-03: P3→P4: the gaps are covered by unit tests; left is live-proof completeness and one log text.** | — | — | CC | | **R-734** | App updates | P4 | **[P3-LOW] The harness marks immich `files_may_change` because immich rewrites six 13-byte `.immich` folder markers at every start.** MEASURED 2026-09-30 on the bench (v3.2.2 → v3.2.4): the bind-tree hash of `appdata/immich` changed; the only changed files were `{encoded-video,library,backups,profile,thumbs,upload}/.immich`, rewritten at each start — no household file. The mark is honest by the harness's rule and the ladder writer copies it (never edited by hand), so immich's v3.2.4 night step needs a fresh WHOLE copy (decision 13); on a box without one the night leg skips it and a person presses. The 2026-09-23 immich entry did not carry it (`files_changed []`). **Needs:** a decision whether app-owned marker files are excluded from the file hash (a per-template ignore list, or a size/name rule), or the mark stays. **-- 2026-09-30 (evening):** the harness now NAMES the files behind the mark (`files_changed_detail`, catalog `5b1972b`); on immich's step `0b82…` re-proof it named exactly the six `.immich` markers again. calibre-web's step (v4.0.6 → v4.0.8) carries the mark too, and the named files are its LIBRARY DATABASE: `media/books/metadata.db`, `metadata.db-shm`, `metadata.db-wal` changed; the book file did not (bench names re-run, `A/calibre-names/`). That is household data (the template's backup class `mandatory` holds DB and books as one unit), so the mark is right there and the night leg takes that step only with a fresh whole copy. On 9202 the same step changed no file in that folder (per-file hashes before/after) — not explained. `audits/more-night-apps-2026-09-30/` | **READY — rank P3-LOW; owner: CC (harness); the rule change needs a word** **Re-ranked 2026-10-03: P3→P4: the effect is an update that waits for a press; no data risk.** | — | — | CC + operator | -## Backup & restore — 52 rows (P2 8, P3 22, P4 22) +## Backup & restore — 51 rows (P2 8, P3 21, P4 22) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -210,7 +210,6 @@ stopping line that lies. | **R-240** | Backup & restore | P3 | **A backup that covered nothing calls itself „Sikeres".** On a configured box with no app selected for off-site backup, a run reports status `ok` with the warning „Sikeres — nincs mentésre jelölt alkalmazás" — *successful* immediately beside *nothing is selected*. Measured as T4 on the final walk, 2026-08-07; flagged once before (2026-08-06) and deliberately not touched then, because the task that noticed it forbade changing that path. **It is the same rhetorical shape the project has spent a fortnight removing** — R-203's *a warning beside a success is read as a success*, R-234's *„✓ Rendben" over an app that was skipped*, R-225's *unknown rendered as zero* — one notch weaker each time, and this is the weakest and last of them. The state itself is honest and must stay `ok`: an unconfigured box reporting `incomplete` forever is its own defect, pinned by a test. **The defect is the word „Sikeres", not the verdict.** Wording such as „Nincs mentésre jelölt alkalmazás — ez a futás semmit nem mentett" says the same thing without congratulating the customer on it. | **READY** — owner Viktor | — | — | operator | | **R-251** | Backup & restore | P3 | **The recovery listing renders one row per restic TAG, so the customer is shown an "app" they never installed and their data counted twice.** Measured on the fifth walk, 2026-08-07, on the screen the customer reaches after entering R. The snapshot carries tags `felhom-offbox,calibre-web`; the listing renders **two rows** — `calibre-web · 2026-08-07 14:57 · 12.8 MB` and `felhom-offbox · 2026-08-07 14:57 · 12.8 MB`. `felhom-offbox` is the tier's own marker tag, not an application. **The screen's whole job is to let the customer check that what is in the store is what they expect** (*"Nézd át, hogy tényleg azt találod-e itt, amire számítasz"*), and it shows them a stranger's name beside their own data and a total that is double the truth. **Cosmetic, not a data defect** — the restore page correctly offers only `calibre-web`. **Fix:** filter the marker tag out of the listing, or key the rows on the app tag. | **READY** — owner Viktor | — | — | operator | | **R-257** | Backup & restore | P3 | **C2 — „Az offsite tároló nincs elárvult állapotban." puts an English loanword and an internal state name in front of a Hungarian household customer, and names no route.** `web/offbox_handlers.go:270` (the Go error it mirrors is `backup/offbox.go:343`). „Offsite" is untranslated; „elárvult állapot" is the codebase's own `OffboxOrphaned()` predicate surfacing verbatim. A customer who pressed a button and got this cannot tell whether something failed, whether they did something wrong, or what to do instead. **This is a refusal that is CORRECT and fail-closed and still a dead end** — the same shape R-241 recorded for `--recover-offsite-install`. **Fix shape, not a decision:** say what the customer tried to do, why it does not apply right now, and where to look — or, since this is a state they cannot reach deliberately, do not offer the action at all | **READY** — owner Viktor | — | — | operator | -| **R-262** | Backup & restore | P3 | **C7 — a comment claims a cross-repo contract is mirrored „field-for-field" and „the key-set tests guard drift"; it is two fields short, AND THE FIXTURE THE TEST READS OMITS THE SAME TWO FIELDS.** `hub/internal/api/handler.go:682-687` covers `hostBackup` **and** `hostRestoreTest`. **It is TRUE of `hostBackup`** (verified field-for-field against `agent/internal/hub/Backup`). **It is FALSE of `hostRestoreTest`:** the agent emits `mount_parity` and `mount_inventory` (`hub/report.go:432-433`, populated in production from `reconcile/restoretest.go:277-283` via `backup/runner.go:517`), and the hub has no field for either — **0 occurrences in the entire hub repo** outside the CHANGELOG. **The guard is blind in exactly the place the drift is:** `TestHostReport_GoldenContract` reads `testdata/host-report.golden.json`, the two copies of which are byte-identical as required — and **neither contains `mount_parity` or `mount_inventory` at all**, so the key sets agree on a shape that is not the shape the agent sends. A test that cannot fail on the drift it names is the R-97b lesson (*prove the consequence, not the mechanism*) landing on a contract test. **Consequence, stated precisely:** the verdict is not lost (a parity mismatch fails the test before `Pass` is set), but the hub cannot distinguish a full-fidelity restore-test pass from a boot-only one, for any agent, ever. **Fix shape, not a decision:** add the two fields and put them in the fixture — or narrow the comment to name `hostBackup` only and say plainly that `hostRestoreTest` is a subset. **Attached observation:** the same fixture carries `cpu_temp_c` and `loadavg`, which no hub struct decodes — a fixture carrying keys the receiver cannot read is the same shape one level down | **READY** — owner Viktor | — | — | operator | | **R-314** | Backup & restore | P3 | **`StopAbandon` has no web route — a customer who telephones is served by a command line.** `--abandon-stop` exists on the controller binary (`cmd/controller/main.go:86`) and refuses rather than silently no-opping, which is right. But there is no handler: the only in-product way to cancel a countdown is the customer finding their recovery code (Scenario G cancels it at the moment the code proves they still have it). **An operator who is telephoned instead has to reach a shell on the customer's machine.** Used this session on the operator's ruling, container stopped first so the running controller could not overwrite `settings.json` from memory — a sequencing subtlety that is itself an argument for a route | **READY (S) — NEW 2026-08-12, RANK 3** | R-241, R-307 | An operator-authenticated POST that calls the same `StopAbandon`, so the telephone path and the code path converge | CC | | **R-362** | Backup & restore | P3 | **A data drive detached mid-restore is reported as „permission denied".** Observed 2026-08-21 23:15: the guest-visible bind was unmounted 4 s into a scratch restore; the restore correctly failed and wrote nothing to the wrong place, but said „A visszaállítás sikertelen: restore dir: mkdir /mnt/felhom-drives/hdd_1/backups: permission denied". The controller has a drive-state concept (`IsDisconnected`, used by both backup legs) and the restore path never consults it. **A correct refusal that misdescribes why sends the reader at a permissions problem that does not exist.** Creditable in the same test: the agent re-bound the drive 5 s later, unaided. | **OPEN — MEDIUM** | — | Consult drive state when a restore path operation fails on ENOENT/EACCES and name the drive. | CC | | **R-401** | Backup & restore | P3 | **Revisit the off-site integrity depth when a real store is LARGE — the default rests on ONE measurement, on ONE 134 MB store.** Controller v0.228.0 (R-399) made `--read-data-subset=100%` the default for every box. The whole justification is a single data point: `demo-hp`, 2026-08-30, 140 829 678 B / 2 651 blobs / 67 snapshots, structure 35.0 s vs 100% **39.2 s** — four seconds. Re-proven live 2026-08-31 at 38.7 s. **It does not extrapolate, and the reason is structural: the structure check's cost tracks the INDEX, a read-data run's tracks the DATA.** A 50 GB store is ~370x the data and this curve says nothing about it. **Nothing was invented from that one point** — no rotation schedule, no size threshold, no bandwidth budget — because four production designs in this project were specced against unvalidated mechanisms and all four were wrong. **`readDataSubsetRe` already accepts `n/m`**, so a rotating schedule (`1/7` on a different seventh each week) needs no parser work when the time comes; the missing input is a measurement on a large store, not code. **THE TRIGGER IS AN EVENT, NOT A DATE:** the slow-check WARN from v0.228.0 firing on any box (`integritySlowNoticeThreshold`, 5 min) — that line names the duration, the depth and this row. **WHAT HAPPENS IF NOBODY ACTS:** every box re-reads its entire store every week, however large it grows, and the first person to notice is a customer whose upload is saturated. **Whoever acts must also revisit `integrityCheckTimeout` (30 min)**, which is now the number a large store meets first. | **OPEN — WATCHING** | — | When the WARN fires: measure the curve on that store, then choose between a rotation (`n/m`), a size-conditional default, or leaving it. Do NOT choose from this row's numbers — they are the small-store case. | CC | @@ -249,7 +248,7 @@ stopping line that lies. | **R-832** | Backup & restore | P4 | **ep0's copy in a place outside both Hetzner and the operator's home (roadmap).** Today DooPlex (the operator's home) holds it (decision 71). A Hetzner Storage Box would share a provider with ep0 and with every household's file backups, and cannot run PBS, so the copy could not be verified or restored from directly. | **DEFERRED — later, if the product grows** | — | — | operator | | **R-878** | Backup & restore | P4 | **A catch-up (R-871) runs the database-dump leg in the DAY, and that leg stops an app with a volume for its copy — the household may notice the stop, and a large volume makes it longer.** MEASURED 2026-10-05 on demo-felhom: the catch-up at 08:25:02 stopped opengist, copied 182.5 KB, started it again — about 1 s, then a few seconds of `health: starting`; the night does exactly the same, unseen. Nothing measured for a large volume. Fix direction (if it matters): skip the volume copy of a running app in a DAYTIME catch-up and leave it to the next night, or warn. `audits/catchup-2026-10-05/partA/live-demo-felhom.txt` | **READY — owner: CC** | — | measure a large volume first | CC | -## Storage & devices — 12 rows (P3 7, P4 5) +## Storage & devices — 11 rows (P3 7, P4 4) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -263,7 +262,6 @@ stopping line that lies. | **R-25** | Storage & devices | P4 | **Device-node TOCTOU hardening (drive init).** Graduate the controller v0.141.0 Observation: the `format → resolveEnrollUUID(path) → AssignDisk(uuid)` sequence has a narrow /dev-re-enumeration window (agent-guarded on the destructive format via anti-retarget durable-id; benign fs-UUID mount). Bind resolve+assign to the format's durable-id so the mount can't target a moved node. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | From the v0.141.0 F6 commit's security-review finding (`felhom-controller` REPORT). Low real risk (single-operator, agent-guarded), but cheap to close | CC | | **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC | | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | -| **R-368** | Storage & devices | P4 | **The storage default DOES apply at deploy time — the earlier claim that it never does was wrong, and the residual defect is smaller and different.** R-352 and `SPEC-app-data-placement-2026-08-21.md` §2.2 stated *"the deploy route never reads it"*, from `grep -nE 'GetDefaultStoragePath|primaryHDDPath|IsDefault' internal/stacks/deploy.go internal/stacks/manager.go` → nothing. **That grep searched Go files only and never the templates.** `internal/web/templates/deploy.html:612` reads `.IsDefault` directly off each `DeployStoragePath` (which embeds `settings.StoragePath`, `web/handlers.go:89-99`) and **pre-selects the default drive for a new deploy**: `{{else if and .IsDefault (not .NotAllowed)}}selected{{end}}`. So `// new apps use this by default` (`settings.go:453`) is **IMPRECISE ABOUT THE MECHANISM, NOT FALSE** — nobody calls `GetDefaultStoragePath()` on that route, but the value is honoured. The customer-facing label promises exactly this and no more: **„Legyen alapértelmezett új telepítéseknél"** (`storage.html:469`). **THE RESIDUAL, and it is the whole finding:** the default lives in the TEMPLATE, not in the server. `POST /api/stacks//deploy` accepts `values` verbatim; omit `HDD_PATH` and `withPathVars` (`stacks/deploy.go:584`) receives `""` and no default is applied. **That is why the invariant has no test — there is nothing server-side to test.** | **OPEN — LOW** | corrects R-352(2); supersedes SPEC §2.2 | Either move the default into the server so the API and the form agree and a test can pin it, or reword the comment to say the template owns it. Do not "fix" the behaviour: it is correct on the path customers use. | CC | | **R-568** | Storage & devices | P4 | **[P3-LOW] The dashboard's drive-health rows swap order between visits — the same two disks, listed in a different order a minute apart.** MEASURED 2026-09-17 on demo-hp 9201 during slice 1 release C's live proof: `/dashboard` fetched on 0.249.0 listed „KXG50PNV1T02 NVMe TOSHIBA 1024GB” then „SanDisk X600 M.2 2280 SATA 128GB”; fetched on 0.250.0 a minute later, the reverse (`audits/i18n-slice1-2026-09-17/C/live/hu-before-vs-after.txt`). `diskHealthRows` (`disk_health.go` L135–141) keeps the agent's response order and does not sort; the agent's order is therefore not stable. Cosmetic, but a household that reads „the second disk” finds a different one. **Fix shape:** sort the rows controller-side by a durable key (device path or serial), with a test that feeds two orders and expects one. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: cosmetic.** | — | — | CC | ## Security & access — 31 rows (P2 2, P3 26, P4 3) @@ -322,7 +320,7 @@ stopping line that lies. | **R-853** | Box system & updates | P3 | **After a boot the box's versions and crash facts reach the hub up to ~15 minutes late.** MEASURED 2026-10-04 on demo-hp (crash-guard test): the agent's first report after a boot has no `system.facts` — the facts read needs a RUNNING customer guest (`firstGuest`), the guest starts ~1–2 min after the agent, and the failed read is cached for 10 minutes; so the HOST half (the crash guard, the kernel) is lost too. The crash events arrived 15 min after the boot (17:17 → 17:32 CEST); nothing was lost (the guard keeps 7 days). Fix direction: the facts mode reads the host without a guest (guest fields `unknown`), and a failed read is not cached. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — owner: CC** | — | — | CC | | **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC | -## Monitoring & notifications — 26 rows (P2 3, P3 15, P4 8) +## Monitoring & notifications — 25 rows (P2 3, P3 15, P4 7) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -346,14 +344,13 @@ stopping line that lies. | **R-266** | Monitoring & notifications | P4 | **A failed root `statfs` still reaches the hub as a 0-of-0 disk, and the hub cannot tell that from an empty one.** Split out of R-259 on 2026-08-08 so that fixing the CUSTOMER-facing half could not be mistaken for fixing the wire. `report/builder.go:93-95` copies `sysInfo.DiskTotalGB` / `DiskUsedGB` / `DiskPercent` into `r.Storage[0]` (`Mount: "/"`), and those are exactly the zeros a failed `statfs` leaves behind — the controller now KNOWS the measurement failed (`SystemInfo.DiskKnown`, controller v0.210.0) and the report still does not carry it. **Deliberately not fixed here, for a reason that is now structural rather than a preference:** adding a field to that report is a change to a declared wire, which since G-1 means the receiving side must model it in the same session (`scripts/wire_contract_gate.py` refuses otherwise) — a two-repo change with a hub bump, and this session deliberately touched no hub code. **RANKED LOW, and the reason is that the consequence is bounded:** the hub bands host storage on `disk_percent`, so a failed read presents as 0% used — the *quiet* direction. It cannot raise a false "nearly full" alarm; it can only fail to raise a true one, and only while the root filesystem is unreadable, which is a state with louder symptoms of its own. **Fix shape when it is taken:** carry `disk_known` on the storage entry and have the hub's fill checker skip an unknown reading rather than band it — never treat absent as 0 | **READY** — owner Viktor | — | — | operator | | **R-285** | Monitoring & notifications | P4 | **A planned, supervised reinstall pages the operator as if the machine had died — there is no notion of expected downtime anywhere.** During the 2026-08-09 rehearsal the hub sent, all `status: sent` to the operator channel: `host_stale` 08:58 UTC, `node_stale` 09:00, **`host_down` 09:28 (error)**, **`node_down` 09:30 (error)**, `host_leaf_changed` 09:31, `host_recovered` 09:31, `node_recovered` 09:34, `offsite_delivery_stuck` 09:34 — eight operator mails for work that was deliberate, attended and announced. **This is the OPPOSITE gap from the one R-281 filed:** the alarms are not missing, they are indiscriminate. `host_stale` at 30 min and `host_down` at 60 min (`monitor/host_staleness.go:22-23`, `downAfter = 2 * threshold`) cannot distinguish a wiped-on-purpose box from a dead one, and `host_leaf_changed` firing on a reinstall is correct-but-expected. **Note the interaction with the mute used on 2026-08-09 evening:** blocking a customer silences everything, so today the only two settings are *page me for planned work* and *tell me nothing at all*. **What is owed is a middle:** a maintenance window, or an operator-set expected-downtime flag, that suppresses staleness and leaf-change while leaving genuine faults audible | **READY (M) — NEW 2026-08-09** | — | The evidence is the operator's mailbox plus `events`/`notification_log` for 2026-08-09 | CC | | **R-337** | Monitoring & notifications | P4 | **`/backup/status` lagged a completed backup by minutes on one box and not the other — and it RESOLVED ITSELF, which is why this is WATCHING and not a defect.** During the R-336 recovery on 2026-08-18, `demo-hp`'s snapshot landed on ep0 at **03:58:43Z** (complete manifest; the host's own task index says `OK`) — yet `GET /backup/status` was **still serving the superseded 03:27:00Z failure at ~04:03Z**, four-plus minutes later. `demo-felhom` showed its new result within ~40 s of completion. **The lag cleared on its own:** demo-hp's 04:07:35Z host report carries `felhom-pbs success=true, 4.29 GB`, and the hub is green for both boxes. **The first draft of this row claimed the success was "still reported as failed" — that was written before the next report arrived and it was wrong; the corrected claim is a several-minute skew between the two boxes, not a stuck value.** It is recorded because a status field that can trail its own artifact by minutes will, during an incident, be read as a second failure — this session nearly did — and because the asymmetry between the two boxes is unexplained | **WATCHING — NEW 2026-08-18** | another observation, ideally during an incident rather than constructed | **Do not open a fix on this as written.** First establish the intended refresh path for `/backup/status` after an out-of-schedule run; only if the skew is not simply collection cadence is there anything to pin. If it is cadence, close this row and say so | CC | -| **R-348** | Monitoring & notifications | P4 | **Every agent restart blanks the reported backup list for up to ~18 hours, and the comment that covers it says "unaffected".** Observed 2026-08-20 while deploying R-344: the first host reports after `demo-hp`'s agent restart carry **`0 backups`** (11:15:50 and 11:30:52 CEST, two consecutive), while the box's own `pvesm list` shows archives present on **both** tiers. `internal/backup/store.go`'s `Store` is in-memory and `byTarget` is repopulated only when a backup **runs** — daily for the local tier, weekly for offsite — so the field reads 0 until the next run. `restore_tests` did **not** blank, because that half has a durable on-disk companion (`RestoreTestState`, R-189). **It blinds no alarm, and that was CHECKED rather than assumed.** `hub/internal/monitor/deadline.go` scans back over stored reports with a 7-day `backupEvidenceLookback` whose own comment names this exact case — *"when the LATEST report carries none... and against an agent that stayed restarted for days"* — and `pbs_snapshots` stayed populated at 2 regardless. So this is an observability wart, **not** a safety hole, and it is filed at that severity deliberately. **What is actually wrong is the comment.** The `Store` doc says *"Backups are unaffected — their freshness has a ground truth on the storage (R-84)"*. That is true of the **consequence** and false of the **field**, and it sits three lines below a paragraph explaining that the very same sentence about restore-tests *"used to be here and it is now FALSE"* — so the file already carries one correction of this shape and invites the next reader to trust the surviving half. | **READY (XS) — NEW 2026-08-20** | — | Say what is measured: the field IS lost on restart and repopulates only when a backup runs; the freshness VERDICT is unaffected because the hub looks back 7 days. **Name `backupEvidenceLookback` in the comment** so the cross-repo dependency is visible from the agent side — today the agent's claim of safety rests on a hub constant it does not mention. Per the workspace rule, a comment asserting an invariant needs a test pinning it: the pin belongs on the HUB side, asserting the verdict survives a report carrying `backups: []`. | CC | | **R-371** | Monitoring & notifications | P4 | **The off-site tier is the only backup tier that announces nothing on success.** Written down 2026-08-05 in `audits/CAMPAIGN-11-recovery-journey-2026-08-05.md:508-513` and explicitly *"recorded, not filed"*: the off-site run emits **no hub event at all**, while both lesser tiers do (`db_dump_completed`, `crossdrive_completed`). Failures are covered by `backup_run_failures` and staleness by the hub's 8-day tier deadline, which is why it was judged a wrinkle. **Still true 2026-08-22** — the 2026-08-21 drill's own event dump shows `db_dump_completed` and six `crossdrive_completed` rows and no off-site success event. **Age when filed: 17 days.** | **OPEN — LOW** | — | Either emit one, or record deliberately that the highest-value tier is silent on success and say why. | CC | | **R-856** | Monitoring & notifications | P4 | **A crash restart reaches the household twice: the hub's "restarted after an unexpected stop" line AND the controller's app mails.** 2026-10-04 crash-guard test on demo-hp: after the third crash and the power-on, the controller sent `app_start_failed` (operator) and `app_stopped_unhealthy` (operator AND the household's address) for apps that were still coming up. Each is true on its own; the app ladder has no "the host just crashed" suppression like its boot grace for an ordinary restart (`08` §5). A design question for the operator, not a defect yet. `audits/os-docker-crash-2026-10-04/partC/c4-hub-events.txt` | **READY — operator decision** | — | — | operator | | **R-872** | Monitoring & notifications | P2 | **A box that is off every night never raises a missed-backup alarm: the 05:00 deadline check skips every customer whose node is `down`, so missing database dumps, second copies and off-site copies stay silent indefinitely; the only nightly signal is `node_down`.** MEASURED 2026-10-05 05:00 Budapest, hub log: `Deadline check: … 0 backup missed … 1 skipped (down)` — the skipped one is Tester 2, off since 18:06 UTC (`hub/internal/monitor/deadline.go` ~360: `if st == "down" \|\| st == StateDisabled { skipped++; continue }`). R-195 / R-321's shape again — a skip keyed off the wrong fact: "down now" was meant to avoid a double alarm, but a box down at every deadline is never checked at all. Fix direction (after R-871): count the days since the last success regardless of the node state, and alarm on N missed nights. `audits/night-fixes-2026-10-05/partF/FINDINGS.md` | **NARROWED 2026-10-05 — FIXED hub v0.134.0, proven by tests (3 red-proofs, `audits/catchup-2026-10-05/partC/`): a down box is judged on 48 h (dump) / 72 h (whole-guest) lines (`08` §6.4, decision 115). LEFT: the first live 05:00 run — DATED CHECK 2026-10-06 (DUE-CHECKS): the hub log line `Deadline check: Tester-2 is DOWN — judged on the longer lines … dump missed=1 backup missed=1` (if Tester 2 is still off at 05:00), and the two events in `events`. Holds → close; does not → a new row.** | R-871 | — | CC | | **R-886** | Monitoring & notifications | P3 | **DooPlex's Alertmanager cannot write its own state since the Longhorn restart of 2026-10-05 13:20Z** — every 15 min `Running maintenance failed … open /alertmanager/nflog.…: permission denied` (and the same for `silences`), 8 times by 14:21Z; the pod was recreated 13:20:48Z by that restart. Mail still goes out (`alertmanager_notifications_total{integration="email"}` 5 → 6, `failed_total` 0, 14:23Z), but a silence set now and the record of what was already sent do not survive the next pod restart — so a restart can re-send every active alarm or drop a silence. Likely collateral of the restart (volume ownership on re-attach), not measured. | **OPEN** | — | Compare the volume's file owner with the pod's `securityContext` (`fsGroup`/`runAsUser`); fix in homelab-manifests; prove with a silence that survives a pod restart | operator | | **R-884** | Monitoring & notifications | P4 | **ArgoCD app `monitoring` shows `Deployment/prometheus` OutOfSync** (seen 2026-10-05 while syncing the R-173 alarm rules; only the rules ConfigMap was synced, so the Deployment drift is untouched and its cause unknown). A full sync would change the running Prometheus in an unknown way. | **OPEN** | — | `argocd app diff monitoring` (or the CR's resource diff) to see what differs, then decide git or live | operator | -## Hub & operator — 25 rows (P2 1, P3 9, P4 15) +## Hub & operator — 24 rows (P2 1, P3 9, P4 14) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -368,7 +365,6 @@ stopping line that lies. | **R-882** | Hub & operator | P3 | **Longhorn on DooPlex could not grow a volume online: its `instance-manager` (116 days up) called a host process that no longer existed** — `nsenter: cannot open /host/proc/196610/ns/mnt` on every expansion retry, and an offline growth was blocked by the expansion's own attachment ticket (found 2026-10-05 growing `hub-data` to 2 Gi). A restart of the instance-manager (operator-approved) fixed it: 77/77 volumes back `attached/healthy` in 110 s. **Why the cached PID went stale was not established** (likely a containerd/k3s or iscsid restart after the instance-manager started), so it will recur after the next such restart and stay invisible until a volume needs to grow. `audits/hub-db-offsite-2026-10-05/partA/step1-*.txt` | **OPEN** | — | Find which host process the PID was and whether Longhorn 1.10.x re-resolves it; until then, before growing any volume, check the instance-manager's age against the last k3s/containerd restart | operator | | **R-883** | Hub & operator | P3 | **8 DooPlex workloads run an image by a moving tag (`:latest` or none), so any pod restart is a silent upgrade.** Measured 2026-10-05: the Longhorn restart restarted zipline on `ghcr.io/diced/zipline:latest` (pull Always), which pulled 4.8.0; 4.8.0 refused its database (`cannot safely migrate from prisma to drizzle: expected migration 20260508022000 … was not applied`) and crash-looped. Fixed for zipline by pinning `4.7.0` (homelab-manifests `90f60e4`, `4c8ec7a`; 4.7.0 applied the four missing migrations; a dump from before is kept out-of-band). The other 7 were counted, not named or changed. | **OPEN** | — | List the 7 (`kubectl get deploy,sts -A` images without a fixed tag), pin each to the running version, and let Renovate move them | operator | | **R-92** | Hub & operator | P4 | Hub PBS-DR gauge is 0.1 GB-granular — small deltas unverifiable | READY (XS) | — | Widen precision when retention becomes customer-visible | CC | -| **R-261** | Hub & operator | P4 | **C6 — `CountSelfBindTokens` exists so that callers can assert an invariant, and no production caller asserts it.** `hub/internal/store/selfbind.go:106-111`. Its doc comment: *"it exists so callers can assert the 'after this runs, the only live link is one we just issued — or none' invariant that the auto-mint at customer-create / RESET-completion depends on."* **Census: only its own declaration in production; the two callers are `selfbind_automint_test.go:29` and `customer_delete_test.go:510`.** Tests are not callers (the campaign's rule), so the invariant the auto-mint *depends on* is checked in the test suite and never at the moment it matters. **This is the smallest of the eight rows and is filed at its true size, because the rest of the C6 sweep found INERT dead accessors rather than defects:** `OffboxOrphanedRenamedTo` and `OffboxEscrowState` have no caller but their data reaches the card another way (the template reads the settings field directly, `backups_remote.html:80`) — **R-228 is genuinely closed, and the sweep's first reading that it had regressed was wrong.** **The more consequential C6 result is a method result and is in the report, not here:** `golang.org/x/tools/cmd/deadcode` re-finds **neither** known instance, and a planted probe measured why — it reports an unreachable exported FUNCTION and not an unreachable exported METHOD on a widely-used type, and both known instances are methods | **READY** — owner Viktor | — | — | operator | | **R-264** | Hub & operator | P4 | **Twenty-one facts the boxes report that the hub can now decode nowhere, each allowlisted with a reason rather than silently skipped — and for these the reason is "no consumer today, and one is arguably owed".** Split out of R-260 on 2026-08-08 so that closing the CLASS (gated) and fixing its sharpest instance (`operator_key_configured`) could not be mistaken for having decided what the hub should do with the rest. **The list, grouped by what a consumer would be for.** **(a) Guest-network health — `guest_net` and its seven children** (`checked_at`, `has_route`, `dhclient_alive`, `heal_succeeded`, `heals_last_hour`, `last_heal_at`, `damped`). The R-54 watchdog reports per-guest network state and self-heal counts every cycle and the hub — the component that emails the operator — models none of it. There is a live incident in this project's own record where a killed `dhclient` took a tunnel down for 1 h 15 m (`audits/INCIDENT-guest-dhclient-killed-2026-07-20.md`); a recurring-heal signal is exactly what would have surfaced it. **This is the strongest candidate of the twenty-one.** **(b) `selfupdate_pending` + `selfupdate_pending_version`** — an agent that has flipped its binary and never committed reports pending on every heartbeat so that "the operator sees WHY the version isn't advancing", and no operator can see it. **(c) `mgmt_plane.healed_recently`** — bounded: the hub DOES alarm on the `privsep_healed_at` timestamp beside it, so the recurring-clobber signal is not lost, only this flag. **(d) `restore_tests.mount_parity` + `mount_inventory`** — R-262's subject; the verdict is not lost (a mismatch fails before `Pass` is set) but the hub cannot tell a full-fidelity pass from a boot-only one. **(e) `pbs_dr.applied_at`.** **(f) Controller-side: `config_hash`, `reporting_disabled`, `stacks`, `storage.migrated_to`, `backup.last_db_dump`, `backup.last_integrity_check`** — the last two are backup-integrity timestamps, which is the "presence is not success" neighbourhood. **For each the question is the same and is NOT answered here: is it wanted? If the hub should act on it, model it and name what consults it. If it should not, the honest end is that the emitter stops sending it** — a fact emitted forever and consumed nowhere is a future false green waiting for someone to write a check against it. **Deliberately not decided in the G-1 session**, whose scope was the gate plus the operator-access instance; unilaterally removing emitters would also break the byte-identical cross-repo host-report golden and is a coordinated two-repo change. **⚠ DISPOSITIONS RECORDED 2026-08-13 — and the first thing to say is that they were NOT in this register before today.** The rulings were made on 2026-08-12; this row still read **READY — owner Viktor** and the gate's twenty entries still all said *"arguably owed"*, so a session told to *"re-read the dispositions from the register"* would have found none. They are written down now, which is the point of writing them down. **THE COUNT WAS ALSO WRONG:** this row says *twenty-one*; the gate's allowlist held **twenty**, measured. Twenty is the number the dispositions below account for, exactly. **(1) BUILD A READER — four groups, fourteen facts.** (a) guest-network health, (b) the staged-update-pending pair, (d) the restore-test depth pair, (f-part) the two backup-integrity timestamps. **(2) NO READER WANTED — five facts, now recorded as `not consumed, DELIBERATELY` with the ruling and its date** in `scripts/wire_contract_gate.py`, each with its own reason rather than a bare refusal: `mgmt_plane.healed_recently` (the hub already alarms on the timestamp beside it), `pbs_dr.applied_at` (`pbs_dr.state` is the verdict; the timestamp alone is the attempt-read-as-result trap), `config_hash` (the hub authors the config and knows its own generation), `stacks` (the app view is built from the purpose-built `app_telemetry` wire), `storage.migrated_to` (box-local bookkeeping with no hub-side intent to reconcile against). **The emitters are deliberately left alone** — removing one is a coordinated two-repo change and breaks the host-report golden; the honest end here is a recorded decision, not a deletion. The gate grew a THIRD entry kind to carry them, because an undecided fact and a decided one must not read alike. **(3) `reporting_disabled`, decided on its own merits: RECLASSIFIED `redundant`** — `health.status = "disabled"` travels in the same minimal report, is decoded into `reports.health_status`, and IS rendered. **The decision surfaced a real defect the flag would not have fixed: the staleness checker is age-only, so a deliberately-silent box still alarms → R-321.** **PROGRESS, 2026-08-13: the first reader is BUILT — guest-network health, R-319.** Its eight allowlist entries are **removed** (an allowlisted tag is skipped, so leaving them would have meant the new reader's fields were never checked); the gate's checked-tag count rose **182 → 190** and skipped fell **88 → 80**, which is the positive control that the wiring is real. **WHERE THE TWENTY NOW STAND: 8 read · 5 deliberately unread · 1 redundant · 6 still owed a reader** (`selfupdate_pending`, `selfupdate_pending_version`, `restore_tests.mount_parity`, `restore_tests.mount_inventory`, `backup.last_db_dump`, `backup.last_integrity_check`) — counts measured from the allowlist, not estimated. **Only ONE reader was built on purpose:** four at once is a design session pretending to be an implementation, and this one now tells us what the other three cost | **OPEN — 6 of 20 still owed a reader; dispositions recorded 2026-08-13, first reader shipped (R-319)** — owner Viktor | — | — | operator | | **R-292** | Hub & operator | P4 | **The artifact-save flash conflates three different facts, and a failing test found it rather than a reading.** `artifact_sha_invalid` reads *"the Gitea sha lookup failed (version missing / Gitea unreachable) or the manually-entered sha is invalid"* — three causes, one message, and the operator acts differently on each. It surfaced because scenario E of the new installability gate kept reporting `artifact_sha_invalid` where it expected `artifact_unverifiable`: `resolveArtifactSHA` ran first and swallowed the distinction. **Worked around in v0.102.0 by ORDERING** — the installability probes now run before the sha resolution, so an unreachable registry is reported as unreachable — **but the underlying message is untouched and still conflates on its own paths** | **READY (XS) — NEW 2026-08-09** | — | Split it into "version not found", "registry unreachable" and "invalid sha" | CC | | **R-372** | Hub & operator | P4 | **A Tier-2 copy that has NEVER been produced because its source path is missing is not surfaced prominently to the operator.** Written down **2026-07-15**, in `audits/CAMPAIGN-6E-2026-07-15.md:128` (F-6E-1), whose disposition ends *"Optional product idea: surface 'tier-2 has never produced a copy (source missing)' more prominently in the operator UI — not filed."* The finding it sits on was correctly judged demo-data churn rather than a product defect (the code warns loudly and does not silently succeed), but the surfacing idea was never carried anywhere. **Age when filed: 38 days — the oldest gap this sweep recovered.** | **OPEN — LOW** | — | Decide whether "never produced a copy" deserves its own operator surface, distinct from "last copy failed". | CC | @@ -395,7 +391,7 @@ stopping line that lies. | **R-793** | Business & legal | P4 | **[P3-LOW] Enterprise / BUSL code ships inside four open images — Cal.com and Docmost (EE folders, off without a key), Outline (BUSL-1.1: no commercial "Document Service"), meilisearch v1.36 in Wanderer (EE modules).** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Each is fine as the catalog runs them: no EE key, the household's own Outline is not a Document Service, Wanderer uses plain search. **Watch:** never turn on an EE feature, never switch Karakeep's/Wanderer's meilisearch to the `-enterprise` image, and re-read on each major. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a watch item; nothing is wrong as the catalog runs them.** | — | — | CC | | **R-794** | Business & legal | P4 | **[P3-LOW] redis 7.4 (RSALv2 / SSPL, not OSI) runs as a private cache in seven apps: dawarich, docmost, immich, nextcloud, outline, paperless-ngx, romm.** READ 2026-10-02 (`audits/licences-2026-10-02/TABLE.md`). Read as permitted (a private cache only its app uses is not Redis offered as a service — inferred). Valkey (BSD-3) or redis 8 (AGPL option) removes the question. **Needs:** a ladder step per app to valkey or redis 8, through the harness — no hurry. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: the row itself says no hurry; usage read as permitted.** | — | — | CC | -## Process & tooling — 75 rows (P3 4, P4 71) +## Process & tooling — 65 rows (P3 4, P4 61) | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| @@ -413,25 +409,18 @@ stopping line that lies. | **R-209a** | Process & tooling | P4 | **The SSD2 move has NOT survived a reboot, so by this project's own standard it is not fully validated** | **WATCHING — NEW 2026-08-05** | the next DooPlex reboot | **Operator ruled explicitly: do NOT reboot DooPlex.** Uptime verified unbroken (7 weeks 6 days, since 2026-06-10). **The distinction is stated rather than glossed: the MECHANISM is proven** — the guard is wired into both units and containerd refuses to start when a required mount's device is absent — **but the CONSEQUENCE is not**: that a real boot mounts `/mnt/ssd_2` before containerd starts, in this host's actual ordering. Mount-ordering reasoning is precisely the class this project has been burned by (`RequiresMountsFor` RE-MOUNTS rather than refusing — the ep0 lesson), and `CLAUDE.md` prefers a consequence assertion over a mechanism one. **Two deliberate consequences: (1)** the rollback copy `/var/lib/containerd.pre-move-2026-08-05` (**34.3 GB on `/`**) **STAYS** until a reboot validates — which is why `/` sits at 54% and not lower; deleting it now would trade a cheap 34 GB for the only cheap way back. **(2)** validation is **automatic and needs no one to remember it**: `felhom-store-postboot-check.service` (oneshot, enabled, dry-run PASS at install) runs at **every** boot and writes `RESULT: PASS`/`FAIL` to `/var/log/felhom-store-postboot-check.log`, asserting positively that `/mnt/ssd_2` is mounted, that containerd's root is on it, that **`/var/lib/containerd` does NOT exist** (the empty-store trap), that ≥100 images are visible and that both dev containers run. **Next action: after the next reboot — planned or not — read that file; on PASS, `rm -rf /var/lib/containerd.pre-move-2026-08-05` returns ~34 GB to `/`** **P3's prune already removed the urgency: `/` went 86% → 53% used and SSD1's Longhorn disk went `Schedulable=False (DiskPressure)` → `Schedulable=True` (18.85% → 50.32% available).** The move was ruled "cap then move"; the cap is in and the pressure is gone, so this is now a deliberate choice rather than a rescue. **The numbers, measured (`Crucial-SSD-240G`, `/mnt/ssd_2/data/longhorn`, `storageMaximum` 235,148,750,848):** available today **214,958,080,000 (91.41%)**; 25% floor **58,787,187,712**. Moving the whole containerd tree at steady state (~31.5 GB images + ≤30 GB cache ≈ 65 GB) leaves **63.77%, i.e. +38.8 pp above the floor — comfortably safe as measured.** **But `storageScheduled` on SSD2 is 139,586,437,120 while `df` says only 20,094,939,136 is actually used** — Longhorn has overcommitted 6.9× — and if those volumes ever inflate to their scheduled size, the same disk lands at **13.00%, i.e. 12 pp BELOW the floor → `Schedulable=False`**, which is exactly the failure that just took SSD1 out. **Recommendation: do the move only together with setting `storageReserved` on SSD2 to cover the containerd tree (~80 GB); SSD2 reserving zero while HDD2 and HDD4 each reserve 500 GB is an anomaly in its own right.** Mechanism, if it goes ahead: **containerd's `root` in `/etc/containerd/config.toml`** (the key is present but commented out) — **not** Docker's `data-root`, which would move only 0.62 GB. Guard: `RequiresMountsFor=/mnt/ssd_2` on `containerd.service` **and** `docker.service`, remembering that **`RequiresMountsFor` RE-MOUNTS rather than refusing** ([[ep0-datastore-volume-move-2026-07-27]]) — so it must be tested with a genuinely absent device, and the move is not validated until it has survived a **reboot**. Full pre-analysis: `audits/SPIKE-dooplex-buildcache-2026-08-05.md` §P6 | operator + CC | | **R-210** | Process & tooling | P4 | **Which of the 345 local images may be deleted — 131 controller tags and 62 hub tags exist ONLY on this box and are not recoverable by `docker pull`** | **WAITING-ON-OPERATOR** — NEW 2026-08-05 | an operator ruling | **Nothing was deleted; this is a list, not an action.** The registry was queried directly: `felhom-controller` has **76** tags in Gitea vs **207** locally, `felhom-hub` **45** vs **107**. The **131 + 62 local-only tags are all OLD** — controller `0.39.0`–`0.135.0` plus `v0.35.0`–`v0.39.0`, hub `0.9.0`–`0.57.0` plus `v0.7.2`–`v0.13.0` — while everything from controller `0.136.0` and hub `0.58.0` upward IS in the registry and therefore re-pullable. **Size the prize honestly before spending a decision on it:** per-tag sizes sum to 139.29 GB, but that double-counts shared layers — `docker system df` puts the **real** dedup'd image footprint at **31.02 GB with 27.02 GB reclaimable**, i.e. an order of magnitude less than the build cache P3 already returned. `docker image prune -a` would remove 343 of 345 (only `redis:7-alpine` and `postgres:16-alpine` are held by running containers). **CC's view: not worth doing for the space** — it buys ~27 GB against 199 GB now free, and its only real benefit is dropping unrecoverable clutter | operator | | **R-230** | Process & tooling | P4 | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | — | — | operator | -| **R-263** | Process & tooling | P4 | **C7 — „This is the ONLY writer of `StoragePath.BackupTarget`" is false, and nothing pins it.** `settings/settings.go:1317-1319`, on `SetBackupTarget`. **`ClearBackupTarget` (`:1358-1363`) also writes the field**, 17 lines below, in the same file. **The GUARANTEE the comment protects is intact and that is why this is filed small:** the sentence continues *"registration must never set it (E-2 §3: a drive never acquires a role by appearing)"*, and `ClearBackupTarget` only ever writes `false`, so no path other than `SetBackupTarget` **grants** the role. **What is wrong is the claim as written, and the absence of anything holding it:** `backup_target_role_test.go` exercises the behaviour and asserts nothing about writer uniqueness, so if a third writer appeared tomorrow — one that granted — the comment would still read as settled and the suite would still be green. **This is the class's own definition:** an invariant asserted in prose with no test pinning it. **Fix shape:** one word (*"the only writer that GRANTS the role"*) plus a test that fails when a second granting writer appears. **Method honesty:** found in a sample of **60 of 2652** production invariant comments — the "is the only" form only, chosen because a uniqueness claim is the one form a grep can falsify. **~2592 production and all 1440 test invariant comments were NOT examined**, so C7 has the weakest coverage of the seven classes and the task's instruction to check the tests' own claims is **owed, not discharged** | **READY** — owner Viktor | — | — | operator | -| **R-286** | Process & tooling | P4 | **A control drawn from the same channel as the measurement cannot detect a defect in that channel — and this one passed while the measurement was wrong.** The P7 check asked *"did the hub record anything?"* against a stale snapshot, got "no", and then validated itself with *"is the hub recording ANY events today, for anyone?"* — **against the same stale snapshot**. It answered "2 events all day", which was internally consistent and entirely false. The standing rule (*an absent log line is not evidence*) was followed in form: a positive control WAS run. **It was the wrong kind of control**, and nothing in the rule as written says so. **The independent channel existed and was available the whole time: the operator's mailbox.** One glance at it would have shown eight alarms in the window. **The durable lesson, to be added where the standing rules live:** a control must come from a DIFFERENT channel than the measurement — same query, same snapshot, same API, same clock all fail this. **Concrete follow-through owed:** (a) add this to the standing rules in `runbooks/workspace-CLAUDE.md`; (b) any hub-state check in a runbook must copy `-wal` or query the pod directly, never `cat hub.db` alone — the trap `operations/nodes.md` already documents | **READY (S) — NEW 2026-08-09** | — | Parent: R-281 (withdrawn) | CC | | **R-288** | Process & tooling | P4 | **The capability map is too long to be read, and that is why it stops being true.** `architecture/00-capability-map.md` is **134 642 bytes / 19 456 words across 99 table rows in only 159 lines** — because the rows ARE the length. Measured, longest first: the unaided-recovery-journey row is **3 024 words**, the offsite-password-recovery row **1 087**, the unattended-restore-proof row **971**, the app/guest-network-failure row **904**. That single longest row is a novella of nested corrections, each appended rather than resolved. Its own verification stamp reads **2026-07-16 against evidence corpus @ felhom.eu tip `4b18cc5`** (line 23) — three weeks stale, which is the measurable consequence: nobody re-reads a row they cannot finish. **This is the project's memory, so restructuring it is surgery and wants daylight** — filed, deliberately not attempted in the 2026-08-09 session. **What the shape should probably be:** one line of status per capability plus a dated evidence pointer, with the argument moved to the audit it came from | **READY (M) — NEW 2026-08-09** | — | Do not fold this into another session; it needs its own. **SECOND CONCRETE COST, 2026-08-10 — and it is a different failure mode from the first.** An *execution record* — 33 package deletions on the operator's rule — was undiscoverable for two days because it lives **inside the row about the Configuration page being slow**. Two sessions searched for it: one reported "no register row records a package prune", the other exhausted the Gitea logs, the activity feed and the schema before concluding it might be unestablishable. It was in `OPEN-ITEMS.md` the whole time. **The first cost (2026-08-09) was two records that looked contradictory and were not; this one is a record that could not be found at all.** Illegibility now has two measured costs and they are different in kind: prose rows make claims ambiguous, and rows-about-other-things make facts unfindable. **The rule this earns is in CONTEXT.md: a record that lives inside a row about something else has not been recorded** | Viktor | | **R-290** | Process & tooling | P4 | **Most capability-map rows that back a green dot cite no evidence document at all — measured, 20 of 28 probed.** The page's *Walked* means *"done end to end on real hardware, evidence on file"*. Extracting the evidence column for the 28 rows behind the page's claims found a `tests/` or `audits/` path in **8**; the other 20 carry prose only. **Consequence, applied this session:** of 32 claims the page drew as Walked, **12 were downgraded to Built** because no walk document exists for them — `install.installer-by-tag`, `use.lifecycle`, `drives.enrol`, `drives.migrate`, `backup.tier1`, `backup.whole-machine`, `backup.restore-proof`, `fault.selfheal`, `fault.operator-email`, `fail.drive-filling`, `fail.lost-recovery-code`, `fail.hub-down`. **This is not a claim that those twelve are false** — several are near-certainly fine — it is a claim that nothing on file distinguishes them from an opinion, which is exactly what the status word promises. **The gate now enforces it going forward:** `scripts/check_stands.py` fails on `status: walked` with no `evidence:` source. **What is owed:** either a walk document per row, or an honest demotion in the map itself (the map is the source; the dataset only follows it) | **READY (M) — NEW 2026-08-09** | R-288 | The dataset was corrected; **the capability map itself still says PROVEN-LIVE for these rows** and is the thing to fix | Viktor | -| **R-291** | Process & tooling | P4 | **CI's installability assertion is now BOUNDED by a retention number, and the narrowing is recorded here so it can be widened deliberately rather than discovered.** `check-published-versions.py` demanded that **every** `v` tag still be downloadable while the registry demonstrably does not retain every version — two sensible rules that cannot both hold, which is why CI went red at a commit whose own run had been green the day before, and would have gone red again at the next publish. **The fix couples them:** `felhom-agent/scripts/retention-policy.json` is THE number (`generic_versions_kept: 10`) and the check reads it. **WHAT CI NO LONGER COVERS, stated plainly: a released version older than the retention window is no longer asserted downloadable.** Its git TAG and its config tree are still asserted — only the binary's presence is dropped — and the check **prints the dropped versions on every run**, so the narrowing cannot go quiet. Controls run: widened to 11 the evicted version re-enters and convicts (exit 1); the policy file removed gives INCONCLUSIVE (exit 2), never silently unbounded. **The number is an OBSERVED state, not a located ruling** (R-287) and the file says so. **The better bound, recorded rather than built:** the hub's vouched `min_agent` floor — nothing can install an agent below it, so a sub-floor version being un-downloadable costs nothing real; it needs the gate to read the hub, which is network it does not have today | **READY (S) — NEW 2026-08-09** | R-287 | ~~Widen or replace the number when the deleter is established~~ — **CONDITION RELEASED 2026-08-10: the deleter IS established (R-287), so the operator is no longer blocked on establishing what was already written down.** The number can now be confirmed or replaced on its merits. The better bound remains the vouched `min_agent` floor | CC | | **R-315** | Process & tooling | P4 | **The wire-contract gate's positive control FAILS on the new wire: it checks name-presence, not decodability.** R-311 declared `hub -> agent (GET /escrow/retained)` as a fourth ROOT, and the gate's tag count rose 174 → 182, so the fields ARE inspected. But renaming the agent-side `superseded_at` json tag to `superseded_at_RENAMED` **still passed** — because the string `superseded_at` also occurs as a map key in the agent's local-API response, and the check is a repo-wide name search. The gate documents this ("name-reachability is not use"), so it is a known limit rather than a regression — but it means **declaring this wire bought documentation, not enforcement**, and a report that claimed coverage would have been wrong. The mutation was asserted to have applied before the run | **READY (M) — NEW 2026-08-12, RANK 3** | R-311 | Make the check resolve the RECEIVER'S mirror type and compare field-by-field, or state per-root which kind of check it got. **A gate whose positive control fails is an instrument nobody has calibrated** | CC | | **R-325** | Process & tooling | P4 | **The shared copy vocabulary is imported by ONE of its two consumers, and drift-checked into the other.** `customer_copy_vocab.py` is the single list; `hub_copy_gate.py` imports it. **`felhom-controller/controller/scripts/retrieval_promise_gate.py` still carries its own `STEMS` literal**, because the session that created the shared module was under a hard end-state requirement to leave `felhom-controller` untouched — its target box was being re-deployed the same evening. **Two copies of a word list is not a theoretical risk in this project: it is the R-299 defect exactly**, where a guard asserted one inflection of a Hungarian verb and the plural walked past it. **So the gap is instrumented rather than left open: `hub_copy_gate.py` READS the controller gate's `STEMS` and FAILS if the two disagree** — single-source semantics tonight without a cross-repo edit. **Watched failing:** removing one stem from the shared list produced *"the shared vocabulary is no longer shared"* with both lists printed, and restoring it returned the gate to green. An ABSENT sibling clone is **INCONCLUSIVE (exit 2), never a pass** — the G-1 lesson. **This is a scaffold, not the destination** | **READY (S) — NEW 2026-08-13, RANK 3** | R-299, R-324 | Make `retrieval_promise_gate.py` import `felhom.eu/scripts/customer_copy_vocab.py` and delete its own literal — a felhom-controller change of a few lines, needing no bake (a gate is not shipped code). Then the drift check becomes redundant and should be removed with it, rather than left as a second mechanism nobody re-reads | CC | | **R-327** | Process & tooling | P4 | **The standing picture still describes a defect that has been fixed twice over.** Found by the first run of `unproven.py` (R-326), which is the argument for having built it. `where-felhom-stands.yaml`'s `claim.code-naming` is `status: partial` and its title reads *"The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not show"* — **both halves of which are now false.** The box side shipped 2026-08-10 (R-295), the hub half and the page-naming fix on 2026-08-13 (R-295 hub, new `reenroll` mail kind), and the third near-homograph on 2026-08-13 (R-323). **NOT MOVED BY THIS SESSION, deliberately and by the dataset's own rule:** *"A status may not be RAISED here — if the evidence supports a stronger status than the capability map records, the MAP changes first and this file follows it."* Raising it here would be the exact inversion the file's header forbids, and the map edit is a separate judgement about what "walked" means for a naming change that no customer has yet met | **READY (S) — NEW 2026-08-13, RANK 4** | R-295, R-323, R-326 | Decide the capability-map status for the naming arc, then let the dataset follow it. **Note the honest difficulty: no customer has typed „Tulajdonosi jelmondat” yet**, so `walked` would be an over-claim; `built` is probably right, and the title needs rewriting either way because it describes a defect rather than a capability | operator + CC | -| **R-345** | Process & tooling | P4 | **`hub/Makefile` tags and pushes `:latest`, which the project's own rules forbid in two places.** Lines 21-22 of `docker-push`: `docker tag $(IMAGE):$(VERSION) $(IMAGE):latest` then `docker push $(IMAGE):latest`. `.claude/rules/hub.md:35` says *"Pin explicit versions, never `:latest`"* and `.claude/rules/manifests.md:15` repeats it. Verified present on `848368ec3`. Small, and the deployed manifests do pin a version, so nothing is currently broken by it — but a documented command that performs the prohibited action is a trap for whoever next reads the Makefile as the reference for how to publish, and a floating `:latest` on the registry is exactly the thing an emergency `kubectl set image` reaches for. Noticed while running the 2026-08-20 connections spike; unfiled until now. | **READY (XS) — NEW 2026-08-20** | — | Delete the two lines, or keep them behind an explicit opt-in target that says in a comment why it exists. Check whether a stale `:latest` tag already sits on `gitea.dooplex.hu/admin/felhom-hub` before deciding — an existing floating tag is the more dangerous half. | CC | | **R-346** | Process & tooling | P4 | **`ActiveEnterTimestamp` answers a different question than the one a slope measurement asks, and on ep0 right now it is wrong by 5 h 56 m.** Found while taking R-341's first dated check. ep0's `proxmox-backup-proxy` has `MainPID=551655` started **2026-08-18 09:51:04Z** (`ps -o lstart=`), but `systemctl show -p ActiveEnterTimestamp` reads **03:54:54Z** and `NRestarts` reads **0** — because the 4.2.5-1 upgrade **re-exec'd** the daemon rather than restarting the unit, so systemd never observed a stop. Anyone anchoring "when did this proxy generation start" on `ActiveEnterTimestamp` would divide 388 descriptors by 52.1 h instead of 46.2 h and report **178/day instead of 201.6/day — ~12% low** — while every field consulted looks healthy and consistent. **This is the workspace rule's own case, in a new place:** ask of a timestamp *what exactly must have happened for this to be set?* Here the answer is "the unit entered active", which is not "this process started". `NRestarts=0` is the tell, and it reads like reassurance. | **READY (XS) — NEW 2026-08-20** | — | R-341's command already uses `ps -o lstart= -p $MainPID` and is correct; the risk is a future reader "improving" it to a systemd property. Add the reason as a comment beside that command in the R-341 row (done), and check whether any other slope or uptime check in the repo or in `scripts/felhom-tenantsync.sh` anchors on a systemd timestamp where it means a process start. | CC | | **R-364** | Process & tooling | P4 | **Accented-text search is an instrument that silently transforms its input, and discipline alone has failed at least three times.** (1) 2026-07-20, `ssh → pct exec → bash -c`, nearly a wrong "banner cleared" claim (`felhom-controller/.claude/rules/ui-hungarian.md:19-22`). (2) 2026-08-13, `kubectl exec … sh -c grep` returned **0 for three strings that were present**, one step from a wrongly-reported failed hub deploy. (3) 2026-08-21, `tar -tf` rendered `őszibarack.md` as `\305\221szibarack.md`; recording the fixture's name bytes from that listing would have been wrong. **NOTE: that is two inside two weeks plus the founding case a month earlier — a third inside the two-week window is not on record.** | **OPEN — LOW** | — | **PROPOSED, NOT BUILT:** a helper that refuses to report a zero for any pattern containing a byte ≥ 0x80 unless a negative control also returns zero AND an ASCII anchor known to be present returns non-zero. Three probes, one helper, no judgement at the call site — because judgement is what failed. | CC | | **R-374** | Process & tooling | P4 | **Three C1 refusal cases were judged borderline, left unfiled, and never named — so nobody can re-open the judgement.** `audits/CAMPAIGN-12-class-sweep-2026-08-08.md:123`: *"'Names a route' is a judgement, not a predicate — two readers could disagree on the borderline cases, and three of the 19 were called borderline and left unfiled."* The disclosure is honest and is exactly the right thing to write; **what is missing is WHICH three.** An unnamed borderline case cannot be re-judged by a second reader, which is the only remedy a judgement call has. **Age when filed: 14 days.** | **OPEN — LOW** | — | Name the three in that document, or file them as one row listing them. No code. | CC | | **R-377** | Process & tooling | P4 | **`CONTEXT.md`'s standing rulings are 188 KB in one section with no sub-headings, and that is why nobody reads them.** Measured 2026-08-22: `CONTEXT.md` is 217 KB, of which **187,913 bytes — 86% — is a single `## Standing rulings` section** carrying 39 `S-` ids and 153 bullets under **one** heading. **This session deliberately did NOT compress or split it**, and the reason is the ruling itself: `PROMPT-TEMPLATE.md` §3.4 and this file's own contract say the decision log is *dated, never edited afterwards*, and it is the only place that answers *"has this been proposed before, and why did we say no?"* — compressing it destroys exactly that. With no per-ruling delimiter, any mechanical split risks cutting a live ruling from its reason, which is the failure this whole arc is correcting. **So the problem is navigational, not volumetric, and the fix is structural: give each ruling a sub-heading with its `S-` id and date.** Then it can be linked, cited and found without a single word being edited. **13 mentions of SUPERSEDED already sit inside that blob** and cannot be separated from live text safely today. | **OPEN — LOW** | R-369 | Add per-ruling sub-headings only. Do not compress, do not reorder, do not edit any ruling's text. | CC | -| **R-391** | Process & tooling | P4 | **Gate 11 (observations) is registered in three of the four runners; `app-catalog-felhom.eu` is the exception.** The controller and agent runners already carried a shared-gate mechanism (`SHARED_REUSE`, `SHARED_INSTRUCTIONS` pointing into `felhom.eu/scripts/`), so registering there was one constant and one `GATES` line each. **`catalog_gates.py` has no such mechanism:** its `run_gate` joins every entry against its OWN `scripts/` directory, so it cannot invoke a sibling repo's script at all; and its loop appends `--all` to every gate unconditionally, which the observations gate would read as a path. Registering there therefore needs `run_gate`'s contract widened AND the argument handling changed — a refactor of a runner whose shape is deliberately different (per-app scoping, network/runtime gates excluded from `--fast`), in a repo this task marked out of scope. **The exposure today is nil** — `app-catalog-felhom.eu/REPORT.md` has no observations section, and the gate passes quietly on that — but a future catalog session could write one and nothing would read it. **Filed rather than left as a sentence in a report, which is the exact failure R-389 records.** | **OPEN — LOW** | — | Either give `catalog_gates.py` the `SHARED_*` absolute-path mechanism the other two runners already have and stop appending `--all` to gates that do not take it, or state in that repo's CLAUDE.md that its REPORT.md carries no observations section by convention. **Do not copy the gate script** — the shared checker lives in ONE place (`felhom.eu/scripts/`) and copying it is the drift the shared pattern exists to prevent. | CC | | **R-392** | Process & tooling | P4 | **No architecture document covers the two-AI workflow.** `documentation/architecture/` holds eight documents and **all eight cover the product** — topology, host agent, control-plane authorization, hub, off-site connectivity, backup, controller modules, capability map. Nothing records how the Claude.ai / Claude Code split works, what each side owns, how skills and `.claude/rules/` are scoped, or why. **The absence was found by trying to fill the template field, not by a survey:** the task that added the five process skills (2026-08-25) had to name an owning architecture document and could not, and the template requires that be recorded rather than passed over. The exposure today is low — the split is stable and both sides work — but it lives entirely in the operator's head and in chat, which is precisely the shape of a commitment nothing enforces. | **OPEN — LOW** | — | Write one architecture document for the agent-tooling layer: which AI owns which artifact class (`TASK-*.md`, `RUNBOOK-*.md`, validation), how skills are scoped and installed, what belongs in a `CLAUDE.md` versus a skill versus a rules file, and the reasoning for each boundary. **The rules themselves already exist** in `skills/felhom-doc-authoring/SKILL.md`; what is missing is the map of who owns what. Do not restate the doc-authoring rules there — point at that skill. | CC | | **R-393** | Process & tooling | P4 | **A decision-log skill for unattended runs was considered and deliberately deferred.** Filed 2026-08-25 by the session that added the five process skills, so the deferral is a decision on the record rather than a thing that was dropped. **The gap it would close:** an overnight or unattended run makes dozens of decisions and the operator can only reconstruct them by reading the whole transcript, which is exactly what nobody does. The proposal is an appended row per decision — what was chosen, why, the evidence pointer, and the result — so a long run is reconstructable in a page. **Why it was NOT built with the other five:** the other five are text files that need nothing but the existing installer glob. This one needs a helper script to append rows and a storage convention for where the log lives and when it is rotated, which makes it an implementation task with its own acceptance criteria, not a skill file. | **OPEN — LOW** | — | Decide the storage convention FIRST — most likely a per-session file beside the session's evidence directory, never `REPORT.md`, which is overwritten every session (the R-341 shape). Then the skill, then the helper. **Check it does not duplicate `felhom-handoff`**, which already owns the end-of-session note; a decision log is the during-the-run half and the two must point at each other rather than overlap. | CC | | **R-394** | Process & tooling | P4 | **`felhom-build-deploy/SKILL.md` is 179 lines, over the 150-line limit its own repo now enforces.** Found 2026-08-25 by `scripts/check_skills.py` on its first run — the over-length was discovered BY the new checker, on the day the limit was written down, which is the checker working as intended. **It is not edited and not trimmed here:** the task that introduced the limit explicitly scoped the four pre-existing skills out, and trimming a build-and-deploy skill without exercising its commands is how a wrong command ships to a live host. **It is a named single-entry exception in `GRANDFATHERED` in `scripts/check_skills.py`, printed as a WARN on every run**, so it cannot fade; a NEW skill over the limit is convicted normally, and growing the set requires editing that file in a commit with a row to name. **The rationale for the limit** — attention thins across the excess, so the lines that matter are not the ones that survive — is in `skills/felhom-doc-authoring/SKILL.md` §5. | **OPEN — LOW** | — | Trim `felhom-build-deploy/SKILL.md` under 150 lines in a session that can VERIFY the commands it keeps, then delete its `GRANDFATHERED` entry in the same commit. The likely trim is the per-artifact command blocks moving behind a pointer to the runbooks, keeping the gotchas inline — but that is a judgement for a session with a build to run, not a line-count exercise. | CC | -| **R-416** | Process & tooling | P4 | **`closed_register_gate.py` still has no within-register duplicate-id rule.** R-406 closed by renumbering the only collision (R-133 → R-415), so the register is clean today and nothing stops the next one. The rule was deliberately NOT added in the same commit: with its only real subject removed, the red-proof would have had to be a planted fixture rather than the live defect, and this project's own standard is that a guard ships with a proof against something real. **Now that the register is clean it can be added safely** — a fresh duplicate would be the first thing it ever sees. | **OPEN — LOW** | R-406, R-405 | Add a third rule to `closed_register_gate.py`: no `R-` id may appear twice within either register. Ship it with a planted red-proof, and note that suffixed ids (R-88a/R-88b, R-209/R-209a) are distinct and must NOT be convicted. | CC | -| **R-418** | Process & tooling | P4 | **`repo_gates.py`'s docstring listed ELEVEN gates while THIRTEEN were registered** — `one-register` (R-369) and `closed-register` (R-405) ran on every push from 2026-08-24 to 2026-09-01 while being documented nowhere. Found while adding the `exemptible` field. This is the drift `felhom-controller/.claude/rules/gates.md` already warns about in its own runner (*"it has already drifted once — it said seven while nine were registered"*), recurring in the sibling repo that the warning does not load for. FIXED in the same commit: all thirteen listed, plus a line saying the table is the list and this is a pointer to it. **Nothing enforces the correspondence** — a gate added tomorrow drifts again, and the fix is a test that compares the docstring list against `len(GATES)`. | **OPEN — the enumeration is fixed, the mechanism is not** | — | — | CC | | **R-420** | Process & tooling | P4 | **`controller_gates.py` could not express a NON-BLOCKING gate before 2026-09-01** — every registered gate's non-zero exit failed the run, so the only way to add a check was to give it the power to refuse a push. That is the wrong trade for a notice that must fire at the moment a release is committed, when the golden legitimately cannot exist yet. **The capability was added rather than the notice compromised** (a fifth `blocking` field, False for exactly one gate; the felhom.eu runner already had the shape from its `--scope` work). Recorded because the ABSENCE was invisible: nobody had wanted a non-blocking gate before, so nothing said it was impossible. `felhom.eu/scripts/repo_gates.py` still has no `blocking` field — it has `exemptible`, which is a different idea (scope-dependent, not permanent). If a permanently-advisory gate is ever wanted there, it needs the same addition. | **OPEN — noted, not needed yet** | — | — | CC | | **R-421** | Process & tooling | P4 | **THE CLASS: an instrument that matches a LABEL rather than the fact it names — five instances, every one found by accident.** R-410 (a `mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was ABSENT), R-94 (a test comparing a constant to itself). **The gates are the machinery that enforces everything else in this project, and they were the one part nothing had ever checked.** The 2026-09-01 decoy sweep read all 29 scripts and fooled **16**. Ten were fixed the same day; four remain with rows (R-422..R-425); six could not be given a plausible decoy and are named. **The shapes, so the next one is cheap to recognise:** (1) name-for-fact — it matches a path or directory NAME while the fact lives inside the file; (2) substring-for-field — it matches a token anywhere in a body instead of in the field that carries it; (3) declaration-for-reachability — it checks a thing is declared, not that it RESOLVES; (4) constant-for-measurement — it compares a value against itself. **The single largest cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level), so every one was green and correct today and would have gone blind the moment anyone added a subdirectory. `decoy_coverage_gate.py` now refuses a new gate that ships without a decoy. | **OPEN — the class row; it stays open as the place the next instance is recorded** | — | — | CC | | **R-422** | Process & tooling | P4 | **`reuse_refs_check.py` only checks citations whose extension is one of `go py html css yml yaml sh`.** A cited `.md` path that does not exist is invisible — MEASURED 2026-09-01: `documentation/architecture/99-does-not-exist.md` added to `REUSE.md` passed, while the `.go` control was correctly convicted. REUSE.md and the CLAUDE.md files cite `.md` paths routinely, so this is the common case, not an exotic one. Fix: widen `PATH_RE`, then walk the false positives it produces across all four repos — that pass is the work, not the regex. The decoy is kept in `scripts/test_gate_decoys.py` asserting TODAY's behaviour, so the day this is fixed the test fails and is updated deliberately. | **OPEN** | — | — | CC | @@ -453,7 +442,6 @@ stopping line that lies. | **R-574** | Process & tooling | P4 | **[P3-LOW] `web/handler_debug.go` mixes page copy with JSON payload, so neither half could be converted safely.** FOUND 2026-09-18 by localisation slice 2 release A (R-557): the file holds 39 Hungarian literals and the inventory classifies them by STATEMENT, not by data flow (`I18N-INVENTORY-2026-09-17.md` §4), so which are section headings the debug page renders and which are values inside a diagnostic dump the operator copies out is not established. Converting a dump value would change what an operator pastes into a report; leaving a heading Hungarian leaves a half-English page. **Fix shape:** walk the file once and label every literal page-copy or payload in the same table slice 2 release A used, then convert only the page-copy half. Belongs to slice 2 release B or C. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: debug page is operator-facing.** | — | — | CC | | **R-576** | Process & tooling | P4 | **[P3-LOW] `i18n_go_parity.py` cannot see a call site that LOST text; it only checks that the text a key carries is real.** FOUND 2026-09-18 by localisation slice 2 release B, and found the hard way: the bulk converter silently dropped the continuation of a multi-line concatenation (`fmt.Errorf("a: "+ "b: %s", x)` kept only `"a: "`), damaging **7** producers — and the gate stayed GREEN throughout, because every surviving fragment WAS a byte-equal base-commit literal. Its question ("is this text real?") was answered yes while the CALL had lost half its sentence and its arguments. Two behaviour tests caught it (`TestR356_ScenarioC_UndeployedAppIsStillRefused`, `TestR379_ScenarioA_RollbackSucceeds_AppComesBack`), because they assert the sentence a customer READS. **Fix shape:** the gate learns a second question — for every `util.MsgError("key", …)` call site, the count of its arguments must equal the count of printf verbs in the key's Hungarian value, and no key-naming literal may be adjacent to a `+`. Both are cheap and would have convicted all 7. **The general lesson, worth keeping whatever is built: a structural gate over the TEXT cannot see a defect in the CALL.** | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: gate improvement; tooling only.** | — | — | CC | | **R-579** | Process & tooling | P4 | **[P3-LOW] Five page shells loaded `style.css` with NO cache-buster, so a browser holding an older copy kept being served CSS that did not know about the newest UI.** FOUND 2026-09-18 by the operator's screenshot of the v0.254.0 language globe: it rendered as a bare, unstyled `
` — a stray triangle and two plain words outside the card — because `login.html`, `claim.html`, `recovery.html`, `launcher_shared.html` and `launcher_share_password.html` requested `/static/style.css` with no `?v=`, while `layout.html` has used `?v={{.Version}}` since v0.166.0. **`.Version` was also absent from three of those five data maps.** FIXED AND CLOSED in the same session (controller v0.255.0): the parameter on all five, and `Version` set once in `executeTemplateLang` so a new shell cannot miss it; `TestGlobeOnAnonymousShells` now refuses an absent or EMPTY `?v=`. **The general form, which is the part worth keeping: a template that loads a versioned asset WITHOUT its version is invisible to every test that reads markup — the markup is correct and the browser fetches the wrong file.** A gate over "every stylesheet/script link in a template carries `?v=`" would catch the class; not built, because 8 first-boot-wizard templates would fail it and R-554 deletes them. | **DEFERRED** (2026-10-03 triage: the row's verdict was finished, but it names open work no other row carries — the gate "every stylesheet/script link carries `?v=`" was deliberately not built until R-554 deletes the old wizard; nothing else owns it) — **CLOSED 2026-09-18 - controller v0.255.0** | — | — | CC | -| **R-588** | Process & tooling | P4 | **[P3-LOW] ISO release records live in two different places, so "was the gate run for this image?" cannot be answered by looking.** FOUND 2026-09-18: the release records for 1.27.0 and 1.27.1 are directories under `documentation/tests/iso-release--/`, and there is none for **1.28.0** — which led me to conclude its gate had not been run. It had: the record is `documentation/audits/evidence-backup-promise-2026-09-16/phaseD-iso-gate.txt`, inside an audit about something else. **The gate's own rule is "a criterion with no recorded observation is a criterion that was not run"**, and that rule is unenforceable while the records have no single home — the question it answers has to be settled by a full-text search for a checksum, which is what it took here. **Fix shape:** one home, `documentation/tests/iso-release--/`, and a line in `iso-release-gate.md`'s Result-recording section naming it; optionally a check that every published ISO version has a record directory. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: documentation hygiene.** | — | — | CC | | **R-591** | Process & tooling | P4 | **[P3-LOW] `Stack.Copy()` is a deep copy with one shallow field, and the field is new.** FOUND 2026-09-20 while adding the catalog's English overlay: `controller/internal/stacks/manager.go` `Copy()` deep-copies `Meta.DeployFields` (with nested `Options`), `Meta.OptionalConfig` (with nested `Fields`), `Meta.Integrations`, `Meta.HealthCheck` and `Meta.InitialCreds` — and does NOT copy the new `Meta.I18n` map, which the struct assignment leaves shared between the original and the "copy". **It is safe TODAY and that is exactly the shape worth filing:** `Metadata.For` reads the overlay and never writes to it (pinned by `TestForDoesNotMutateTheReceiver`), so nothing can observe the sharing yet. The hole is in the CONTRACT — a function whose whole purpose is "a snapshot the caller may mutate" now has a field that is not one, and the next person to write through an overlay will find a bug with no failing test in front of it. **Fix shape:** deep-copy `I18n` in `Copy()` and pin it with a test that mutates the copy's overlay and asserts the original is unchanged. Alternatively state in `Copy()`'s comment that `I18n` is deliberately shared and immutable, and pin THAT with a test. Either is fine; silence is not. | **READY - rank P3-LOW; owner: CC (controller)** **Re-ranked 2026-10-03: P3->P4: latent contract gap, safe today.** | — | — | CC | | **R-594** | Process & tooling | P4 | **[P3-LOW] The catalog copy gate can CONVICT a retrieval promise but has no way to REGISTER a true one.** FOUND 2026-09-20 translating batch 3 (R-560 slice 5). Vaultwarden's invite step and its sign-up setting both ended „…can open an account", and the English retrieval-promise pattern reads `can … open` as the claim that sealed backups can be opened. **The conviction was a FALSE POSITIVE** — opening an account is not opening a backup — and the two sentences were reworded to „can sign up", which is also the better copy, so nothing is blocked today. **The gap is structural.** The shared vocabulary this gate copies (`scripts/customer_copy_vocab.py`) states the design explicitly: these stems are NOT banned, because each carries a claim that is sometimes TRUE, and *"an occurrence must be REGISTERED with a reason in the consuming gate's allowlist"*. The hub gate has `ALLOWLIST_EN`; `app-catalog-felhom.eu/scripts/check-copy-i18n.py` has none, so the only ways past it are to reword or to bypass the gate — and a catalog app whose English genuinely says a file can be restored (a backup app, a versioned document store) has no honest third option. **Fix shape:** an `ALLOWLIST_EN` of `(app, path, reason)` in the gate, a decoy proving a REGISTERED occurrence passes and an unregistered one still convicts, and a check that every entry still matches something (a stale allowlist entry was R-299's shape, and the retrieval gate has gone red on stale entries before). | **READY - rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3->P4: gate tooling; nothing blocked today.** | — | — | CC | | **R-603** | Process & tooling | P4 | **[P3-LOW] An English string containing an apostrophe silently never matches on a rendered page, and a `strings.Contains` assertion reads exactly like a missing sentence.** FOUND 2026-09-21 while writing the R-598 render tests. `backup.target.absent` was first written as *"The system backup's drive cannot be reached…"*; `html/template` escapes `'` to `'`, so the page carried the sentence and every assertion for it failed. **The failure mode is the expensive part:** the test said *"the English absent-drive copy never reached the page"*, which is indistinguishable from the handler not being wired — and the obvious next move is to go and re-fix the handler. Reworded to avoid the possessive, and all 23 new English values were then swept for `' " < > &` (zero). **The Hungarian bundle has never hit this** because Hungarian copy uses „quotes" and few apostrophes; **English copy will hit it again.** **Fix shape:** either a bundle gate that refuses an HTML-escapable character in a value destined for a page (and an allow-list for the ones that legitimately need one), or a test helper that compares against `html.EscapeString(want)` so the assertion cannot be fooled. The gate is the better shape — the helper only protects tests that remember to use it. | **READY — rank P3-LOW; owner: CC (controller)** **Re-ranked 2026-10-03: P3->P4: test tooling.** | — | — | CC | @@ -464,11 +452,9 @@ stopping line that lies. | **R-731** | Process & tooling | P4 | **[P3-LOW] The catalog-currency comparison cannot see an upstream that changes its TAG SHAPE, and two newest tags need a release check.** MEASURED 2026-09-30 (`audits/catalog-currency-2026-09-30.md` §2 item 5): the same-shape rule read three apps as up to date that were not — gramps-web (`v25.6.0` → upstream dropped the `v`, at `26.9.1`), jellyfin (`10.11.11` → two-part `12.1`), kimai (`apache-2.57.0` → plain `2.67.0`; the plain tag's digest equals `apache`'s, so kimai moved on 2026-09-30). A control pass over every shape caught them. Not checked: whether `mariadb:13.0` and `gitea/gitea:28.0.0` (pushed 2026-09-30 00:15 UTC) are general releases. **Needs:** the currency script's shape-switch control made standing (it is in the audit's tools today), and the two release checks before either is walked. **-- 2026-09-30: the two release checks answered** (`audits/immich-first-start-2026-09-30/D/D2-release-checks.txt`): **gitea v28.0.0** is a general release (GitHub: not prerelease, not draft, 2026-09-29) — upstream renumbered 1.27.x → 28; **mariadb 13.0** is `Stable` but a short-term `Rolling` line (no EOL date), while 12.3 and 11.8 are the LTS lines — so a move of any of the four MariaDB apps to 13.0 would leave LTS. Not done: the shape-switch control made standing. | **NARROWED 2026-09-30 — only the standing shape-switch control is left; owner: CC** **Re-ranked 2026-10-03: P3→P4: only a standing check in a tooling script is left.** | — | — | CC | | **R-739** | Process & tooling | P4 | **[P3-LOW] The test bench cannot run wanderer at all: its web server calls the database at the PUBLIC name `https://.`, which the bench has no name or TLS for.** MEASURED 2026-09-30 (more-night-apps brief, Part B): on bench 9401 the catalog template (v0.20.0) came up `wanderer-db` healthy, `wanderer-search` healthy, `wanderer` **unhealthy** for 12 min, every page 500 („Error 0: Something went wrong"); `PUBLIC_POCKETBASE_URL=https://hike-db.gate.invalid`. So the harness can only ever answer `inconclusive` at FROM for wanderer, never about an update. Its step `v0.20.0 → v0.21.0` (web + db) and meilisearch `v1.36 → v1.54` stay untested; **the brief's question — does the search index survive meilisearch's move or is it rebuilt — is NOT measured.** **Needs:** a bench venue that gives the stack the DB name (an `extra_hosts` + plain-http override in the bench's render only, stated in the verdict), or wanderer proven on the box venue alone with the operator's word. `audits/more-night-apps-2026-09-30/B/wanderer-probe.txt` **-- 2026-09-30 late: the bench CAN run wanderer now** — `upgrade-test.py` `BENCH_ENV_OVERRIDES` points `PUBLIC_POCKETBASE_URL` at the database container on the bench only (the only address the web image reads, measured in v0.20.0 and v0.21.0), recorded in every verdict: web 200, sign-up (`PUT /api/v1/user`) 200, login 200. **The meilisearch question, answered:** v1.54.2 REFUSES v1.36's database („Your database version (1.36.0) is incompatible”); with `MEILI_UPGRADE_DB=true` it upgrades in place and all three indexes (actors, lists, trails) are there — so wanderer's step needs that switch in the template first. Not done: a fixture (creating a list answered PocketBase's „Failed to create record” — likely a rule for unverified users), whether a household's record survives in the index, the step itself. `audits/night-rulings-2026-09-30/` | **NARROWED — the bench runs it; the step needs a fixture and `MEILI_UPGRADE_DB`; owner: CC** **Re-ranked 2026-10-03: P3→P4: test bench coverage; the template switch and fixture are tooling work.** | — | — | CC | | **R-759** | Process & tooling | P4 | **[P3-LOW] wger's onboarding record (the checklist pilot) keeps rows open that no other row owns.** 2026-10-01, `app-catalog-felhom.eu/onboarding/wger.md`: **2.5** no backup → remove → restore → read back of wger exists — the box has no per-app backup press outside an Update (R-648) and wger has no newer step to carry one; **3.7** changing the password and adding a family member not measured, and the template has no `add_people` text; **6.3** no forced-fail undo for wger; **8.2** the app page not read on 9202 this session; **9.1** the runtime volume-persistence gate not re-run (last CLEAN 2026-08-02). The other open rows have their own: 1.5 (R-755), 1.7/2.8 (R-762), 3.4 (R-763), 7.1 (R-764). wger is exempt from the onboarding gate (published before the checklist), so nothing blocks; this row is what keeps the record honest. **Needs:** the five measured on 9202 — 2.5 and 6.3 ride wger's next ladder step (the update's backing-up phase is the per-app backup). `audits/new-app-checklist-2026-10-01/` | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: record-keeping for a hidden app; nothing blocks.** | — | — | CC | -| **R-761** | Process & tooling | P4 | **[P3-LOW] The canonical example template tells a new app's author the logo is `-logo.webp`; the controller loads `-logo.svg`, then `.png`.** READ 2026-10-01 (checklist row 8.3): `templates/paperless-ngx/.felhom.yml` lines 20–24 (the comment block REUSE.md §2 says to copy) name `{assets.base_url}/assets/{slug}-logo.webp`; `felhom-controller` `internal/config/config.go` `AppLogoURL`/`AppLogoPNGURL` ask for `-logo.svg` and `-logo.png`; `https://felhom.eu/assets/wger-logo.webp` answers 404, `wger-logo.png` 200 (negative control `nosuchapp-logo.png` 404). A new app's author following the comment publishes a logo the box never loads. **Needs:** the comment corrected (a comment-only template change, in a session allowed to touch templates). The checklist row 8.3 already names the right files. `audits/new-app-checklist-2026-10-01/B/B1-upstream-reads.txt` | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: a misleading comment for app authors; no household meets it.** | — | — | CC | | **R-781** | Process & tooling | P4 | **[P3-LOW] The catalog's `scripts/test_gate_decoys.py` fails 4 of its own "genuine" onboarding cases — on the untouched tree.** Measured 2026-10-01 (same 4 FAIL lines before and after this session's checklist edit): the harness copies the catalog into a temp tree with a fake sibling `felhom.eu`, and the REAL published records (radicale, karakeep, dawarich) name evidence that the fake sibling lacks, so a complete record reads incomplete. The gate itself (`onboarding`) is green on the real tree. **Needs:** the genuine cases build their own records, or the fake sibling mirrors the cited evidence. | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: a test-harness self-check; the real gate is green.** | — | — | CC | | **R-786** | Process & tooling | P4 | **[P3-LOW] SparkyFitness's onboarding record has six open rows** (`app-catalog-felhom.eu/onboarding/sparkyfitness.md`): 0.5 runtime internet (food search providers), 0.7 the phone app's sign-in route through traefik, 1.6 the env names the server reads, 1.7 the entrypoint read, 5.4 a second memory watch at another limit, 8.3 no logo/screenshots on felhom.eu (404). Everything else measured this session (bench + 9202). **Needs:** each row measured, or n/a with a reason. | **READY — rank P3-LOW; owner: CC (catalog)** **Re-ranked 2026-10-03: P3→P4: onboarding record completeness; no household meets it directly.** | — | — | CC | | **R-797** | Process & tooling | P4 | **[P3-LOW] `check-family-gate.py` rule 3 (a family_gate template needs a baked golden ≥ 0.287.0) is checked only where the felhom.eu sibling exists — CI's single clone cannot.** MEASURED 2026-10-02 (`audits/family-gate-2026-10-02/C-metube-bench/C3b-family-gate-with-sibling.txt`): without the sibling the gate printed NOT CHECKED and its summary read OK. The summary line now says "rule 3 … NOT CHECKED here". The pre-push hook (with the sibling) is where it bites. **Needs:** nothing more unless CI gets the sibling. | **WATCHING — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a CI coverage gap; the pre-push hook still checks it.** | — | — | CC | -| **R-799** | Process & tooling | P4 | **[P3-LOW] The MeTube fixture's `POST /add` leaves out `download_type`, which upstream's validator lists as required.** MEASURED 2026-10-02 (`audits/family-gate-2026-10-02/C-metube-bench/`): both 2026.09.28 and .29 answered 200 and downloaded anyway. Fragile if a later tag enforces it. **Needs:** add `download_type: video` to the fixture with the next MeTube step. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a test fixture detail.** | — | — | CC | | **R-805** | Process & tooling | P4 | **[P3-LOW] The volume-persistence gate judges an empty NAMED volume (R-788) but not an empty BIND mount.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): Grimmory's `/app/data` bind held 0 files and read CLEAN (its state is in MariaDB); komga, paperless-ngx, radarr and sonarr also carry empty binds that are not judged. **Needs:** decide whether an empty bind after the exercise is UNDETERMINED like a named volume, with a decoy. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a test-gate rule; no household meets it.** | — | — | CC | | **R-806** | Process & tooling | P4 | **[P3-LOW] The gate's GET exercise speaks plain http to an HTTPS backend (crafty-controller :8443), and gramps-web did not answer on :5000.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): both reached data only through their fixture seed. **Needs:** read the `loadbalancer.server.scheme` label (upgrade_boxport already does); find why gramps-web's first GET got no answer. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: a test-gate defect; no household meets it.** | — | — | CC | | **R-807** | Process & tooling | P4 | **[P3-LOW] 13 apps stay UNDETERMINED because their seeds write only to the database, so their upload / media / cache / redis volumes stay empty.** MEASURED 2026-10-02 (`audits/persistence-sweep-2026-10-02/A/sweep/TABLE.md`): claper, crafty-controller, dawarich (redis), docmost, gramps-web, immich (ML cache), outline (+redis), sparkyfitness, tandoor, vikunja, wger, wishlist, zipline — plus plex (no seed route: a plex.tv claim token) and wanderer (no fixture; unhealthy under the gate). None is BROKEN: nothing was written outside a preserved folder. **Needs:** per app, a seed that uploads one file (or the volume named n/a with a reason: a redis cache, an ML model cache). | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: test coverage; nothing was found broken.** | — | — | CC | diff --git a/documentation/runbooks/iso-release-gate.md b/documentation/runbooks/iso-release-gate.md index f20ec51f..d7bcc72a 100644 --- a/documentation/runbooks/iso-release-gate.md +++ b/documentation/runbooks/iso-release-gate.md @@ -320,3 +320,8 @@ record. Record each criterion as PASS/FAIL **with the observed value and what was scanned for**, in the release report. A criterion with no recorded observation is a criterion that was not run. + +**The ONE home for a release's record is `documentation/tests/iso-release--/`** (R-588, +2026-10-05) — one directory per published ISO, so „was the gate run for this image?" is answered by looking, +not by a full-text search for a checksum. A record made elsewhere (inside an audit) gets a directory here +whose `README.md` points at it. Today: 1.27.0, 1.27.1, 1.28.0 (pointer), 1.29.0. diff --git a/documentation/runbooks/workspace-CLAUDE.md b/documentation/runbooks/workspace-CLAUDE.md index de249a49..48c4f18d 100644 --- a/documentation/runbooks/workspace-CLAUDE.md +++ b/documentation/runbooks/workspace-CLAUDE.md @@ -54,7 +54,13 @@ roles. **A file being open in the editor is NOT an instruction. If no task is st attempts. 3. **An absent log line is not evidence of correct behaviour.** Verify with a POSITIVE observable — something that MUST appear when the system is healthy. An empty log is equally consistent with - "working" and "stopped entirely". + "working" and "stopped entirely". **And the control must come from a DIFFERENT channel than the measurement** (R-286): + same query, same snapshot, same API or same clock all share the defect they are meant to catch — a + stale hub snapshot once "confirmed" itself (2 events all day) while the operator's mailbox held eight + alarms. A hub-state check copies `hub.db-wal` too, or asks the running pod. **And the control must come from a DIFFERENT channel than the measurement** (R-286): + same query, same snapshot, same API or same clock all share the defect they are meant to catch — a + stale hub snapshot once "confirmed" itself (2 events all day) while the operator's mailbox held eight + alarms. A hub-state check copies `hub.db-wal` too, or asks the running pod. 4. **A recommendation that is not followed gets one line saying why.** Silence reads as agreement and the disagreement is lost. 5. **Evidence is copied off the machine at the end of the phase that produced it — before any revert, diff --git a/documentation/tests/iso-release-1.28.0-2026-09-16/README.md b/documentation/tests/iso-release-1.28.0-2026-09-16/README.md new file mode 100644 index 00000000..d49f8c49 --- /dev/null +++ b/documentation/tests/iso-release-1.28.0-2026-09-16/README.md @@ -0,0 +1,9 @@ +# ISO release 1.28.0 — 2026-09-16 (pointer) + +The gate record for this image was made inside an audit about something else, before the one-home rule +(`runbooks/iso-release-gate.md` „Result recording", R-588): + +- `documentation/audits/evidence-backup-promise-2026-09-16/phaseD-iso-gate.txt` — run 2026-09-16T15:07:32Z against + `felhom-installer-1.28.0-pve9.2-1.iso`, sha256 `a4cd9b6ddcb55bae3700ab307084d2b699330cc20710688d5816318f04f6d635`. + +Nothing was re-run for this pointer; it only makes the record findable by looking. diff --git a/hub/CHANGELOG.md b/hub/CHANGELOG.md index fd917c17..f7aa866d 100644 --- a/hub/CHANGELOG.md +++ b/hub/CHANGELOG.md @@ -1,3 +1,14 @@ +## unreleased — comments, a pinned subset and the build without `:latest`; no image change (burn-down 2026-10-05: R-261, R-262, R-345) + +The next hub release carries these lines into its own entry. Nothing here changes the running hub. + +- **R-262:** `hostRestoreTest` is a deliberate SUBSET of the agent's `RestoreTest` — `mount_parity`, `mount_inventory` + and `skipped` are not modelled (a skipped test reads as failed with its reason, by the agent's design). + `internal/api/r262_restoretest_subset_test.go` pins both lists and cross-checks the agent source when it sits beside + this repo; it found `skipped`, which the comment had not named. +- **R-261:** `CountSelfBindTokens` is documented as the test accessor it is. +- **R-345:** `Makefile` `docker-push` no longer tags or pushes `:latest` (nor does `scripts/build-hub.sh`). + ## v0.136.0 — the hub writes a nightly, checked copy of its own database (R-173, decision A) (2026-10-05) **Operator action on deploy: none.** The hub's volume grows from 1 GiB to 2 GiB and joins Longhorn's nightly backup diff --git a/hub/Makefile b/hub/Makefile index 865ce439..b1b6f4b2 100644 --- a/hub/Makefile +++ b/hub/Makefile @@ -18,8 +18,7 @@ docker: docker-push: docker docker push $(IMAGE):$(VERSION) - docker tag $(IMAGE):$(VERSION) $(IMAGE):latest - docker push $(IMAGE):latest + # never :latest (R-345) — the manifest pins a version; a moving tag is how a restart upgrades by surprise clean: rm -rf bin/ diff --git a/hub/internal/api/handler.go b/hub/internal/api/handler.go index c8dc75e9..13002f7f 100644 --- a/hub/internal/api/handler.go +++ b/hub/internal/api/handler.go @@ -724,8 +724,9 @@ type hostPBSSnapshot struct { VerifyUPID string `json:"verify_upid,omitempty"` } -// hostBackup / hostRestoreTest mirror the agent's hub.Backup / hub.RestoreTest wire -// contract field-for-field (slice 6, doc 03 §8). DUPLICATED contract — the golden stays +// hostBackup mirrors the agent's hub.Backup wire contract field-for-field; hostRestoreTest is a deliberate +// SUBSET of hub.RestoreTest — the agent's mount_parity, mount_inventory and skipped are NOT modelled here (R-262, +// pinned by TestR262_RestoreTestFieldsAreAKnownSubset) (slice 6, doc 03 §8). DUPLICATED contract — the golden stays // byte-identical with felhom-agent's copy and the key-set tests guard drift. The hub // persists these via report_json (no new columns this slice) and surfaces a FAILED // restore-test prominently (the loudest DR signal). The rich backup policy is slice 10. diff --git a/hub/internal/api/r262_restoretest_subset_test.go b/hub/internal/api/r262_restoretest_subset_test.go new file mode 100644 index 00000000..f20d271b --- /dev/null +++ b/hub/internal/api/r262_restoretest_subset_test.go @@ -0,0 +1,58 @@ +package api + +import ( + "os" + "reflect" + "regexp" + "sort" + "strings" + "testing" +) + +// R-262: hostRestoreTest is a deliberate SUBSET of the agent's hub.RestoreTest. This pins WHICH subset: the hub decodes +// exactly `hubRestoreTestFields`, and the agent fields it leaves out are exactly `knownUnmodelled`. When the agent's +// source is beside this repo (the workspace), its RestoreTest json tags must equal the union — so an agent field added +// later fails here until someone decides to model it or to list it as unmodelled. RED-PROOF: drop a field from the +// struct, or delete an entry from knownUnmodelled → this test fails. +func TestR262_RestoreTestFieldsAreAKnownSubset(t *testing.T) { + // skipped (agent v0.133.0, R-672): a skipped test arrives as Pass=false with Error "skipped: …", which the hub + // reads as a failed test with that reason — the agent's own comment calls that the honest reading. + knownUnmodelled := []string{"mount_inventory", "mount_parity", "skipped"} + var got []string + rt := reflect.TypeOf(hostRestoreTest{}) + for i := 0; i < rt.NumField(); i++ { + got = append(got, strings.Split(rt.Field(i).Tag.Get("json"), ",")[0]) + } + sort.Strings(got) + want := []string{"duration_seconds", "error", "pass", "scratch_vmid", "source_archive", "source_tier", + "tested_at", "verified", "warnings", "warnings_recognized"} + if !reflect.DeepEqual(got, want) { + t.Fatalf("hostRestoreTest json fields = %v, want %v", got, want) + } + for _, f := range knownUnmodelled { + for _, g := range got { + if g == f { + t.Fatalf("%s is listed as unmodelled but hostRestoreTest decodes it — update knownUnmodelled", f) + } + } + } + src, err := os.ReadFile("../../../../felhom-agent/internal/hub/report.go") + if err != nil { + t.Logf("agent source not beside this repo (%v) — the cross-repo half is not checked here", err) + return + } + body := regexp.MustCompile(`(?s)type RestoreTest struct \{(.*?)\n\}`).FindSubmatch(src) + if body == nil { + t.Fatal("type RestoreTest struct not found in felhom-agent/internal/hub/report.go") + } + var agent []string + for _, m := range regexp.MustCompile("json:\"([a-z_]+)").FindAllSubmatch(body[1], -1) { + agent = append(agent, string(m[1])) + } + sort.Strings(agent) + union := append(append([]string{}, want...), knownUnmodelled...) + sort.Strings(union) + if !reflect.DeepEqual(agent, union) { + t.Fatalf("agent RestoreTest fields %v != hub fields + knownUnmodelled %v — model the new field or list it", agent, union) + } +} diff --git a/hub/internal/store/selfbind.go b/hub/internal/store/selfbind.go index 255fa713..16bc17aa 100644 --- a/hub/internal/store/selfbind.go +++ b/hub/internal/store/selfbind.go @@ -104,10 +104,11 @@ func (s *Store) DeleteSelfBindTokens(customerID string) error { } // CountSelfBindTokens reports how many capability tokens exist for a customer (v0.67.0). Minting is -// single-active (delete-then-insert), so this is 0 or 1 in practice; it exists so callers can assert -// the "after this runs, the only live link is one we just issued — or none" invariant that the -// auto-mint at customer-create / RESET-completion depends on. Read-only, no oracle risk: it is keyed -// by customer id, which the operator already knows. +// single-active (delete-then-insert), so this is 0 or 1 in practice. It is a TEST ACCESSOR (R-261): no +// production code calls it. The "after this runs, the only live link is one we just issued — or none" +// invariant the auto-mint at customer-create / RESET-completion depends on is pinned by +// web/selfbind_automint_test.go and web/customer_delete_test.go, which call it. Read-only, no oracle risk: +// it is keyed by customer id, which the operator already knows. func (s *Store) CountSelfBindTokens(customerID string) (int, error) { var n int err := s.db.QueryRow(`SELECT COUNT(*) FROM selfbind_tokens WHERE customer_id = ?`, customerID).Scan(&n) diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index a9b8751f..25aa736b 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,3 +1,11 @@ +## build + gates — no `:latest`; the gate list pinned; duplicate closed ids refused (2026-10-05, burn-down) + +- **R-345:** `build-hub.sh` no longer tags or pushes `felhom-hub:latest` (`--push`, `--multiarch`, local); nothing + pulls it. `test_no_latest_push.py` walks `hub/` and `scripts/` build files (red-proofs: the old Makefile, the old + build script). +- **R-418:** `repo_gates.py`'s docstring lists all 17 gates; `test_repo_gates_docstring.py` keeps list == `GATES`. +- **R-416:** `closed_register_gate.py` RULE 4 — an id twice in `CLOSED-ITEMS.md`; decoy in `test_gate_decoys.py`. + ## gates — `script-tests`: every Python test suite under scripts/ runs on every push (R-885) (2026-10-05) - `scripts/script_tests_gate.py` (registered in `repo_gates.py`, fast): walks `scripts/` for `test_*.py` and runs each; diff --git a/scripts/build-hub.sh b/scripts/build-hub.sh index 4f57272f..3c0e10f2 100755 --- a/scripts/build-hub.sh +++ b/scripts/build-hub.sh @@ -147,12 +147,11 @@ case "${ACTION}" in info "Building for current platform + pushing..." docker build "${BUILD_ARGS[@]}" \ -t "${IMAGE}:${VERSION}" \ - -t "${IMAGE}:latest" \ . info "Pushing..." docker push "${IMAGE}:${VERSION}" - docker push "${IMAGE}:latest" + # never :latest (R-345): the manifest pins a version, and nothing pulls hub:latest (checked 2026-10-05) ;; --multiarch) @@ -169,7 +168,6 @@ case "${ACTION}" in docker buildx build "${BUILD_ARGS[@]}" \ --platform linux/amd64,linux/arm64 \ -t "${IMAGE}:${VERSION}" \ - -t "${IMAGE}:latest" \ --push \ . ;; @@ -178,7 +176,6 @@ case "${ACTION}" in info "Building for current platform (local only)..." docker build "${BUILD_ARGS[@]}" \ -t "${IMAGE}:${VERSION}" \ - -t "${IMAGE}:latest" \ . ;; esac diff --git a/scripts/closed_register_gate.py b/scripts/closed_register_gate.py index 58e2bc8a..bba07bd6 100644 --- a/scripts/closed_register_gate.py +++ b/scripts/closed_register_gate.py @@ -50,9 +50,9 @@ WHAT THIS GATE CANNOT SEE — the residual holes, named rather than implied: printed as a WARNING. 3. **A closed-sounding verdict that is not true escapes.** `PARTLY CLOSED` leads with no open word. This gate checks where a row FILED, never whether the verdict is honest. - 4. **A duplicate id WITHIN one register escapes.** `OPEN-ITEMS.md` carries two unrelated findings - both numbered R-133 (filed as R-406). Adding that rule would fail the gate on a pre-existing - defect, and a registered-but-failing gate refuses every push, so it was deliberately left out. + 4. ~~A duplicate id WITHIN one register escapes.~~ **Closed 2026-10-05 (R-416):** in `OPEN-ITEMS.md` + `register_shape_gate.py` RULE 3 refuses it; in `CLOSED-ITEMS.md` this gate's RULE 4 does (no duplicates + existed when it was added, so it was registered green). 5. Nothing here reads audits, spikes or inventories. A finding that never reaches either register is invisible to this gate, as it is to `one_register_gate.py`. @@ -153,6 +153,16 @@ def main(): for rid in sorted(closed_ids & set(open_ids), key=lambda r: (int(re.sub(r"\D", "", r)), r)): convicted.append((open_ids[rid], rid, "", "has a row in BOTH registers")) + # RULE 4 — a duplicate id WITHIN CLOSED-ITEMS.md (2026-10-05, R-416). The open register's duplicates are + # register_shape_gate.py's RULE 3; this file had no such rule, so two closed rows under one id (the R-133/R-406 + # shape) would make `git show` of "the row that closed R-n" ambiguous. Suffixed ids (R-88a, R-88b) are distinct. + first_seen = {} + for n, rid, _, _ in rows(CLOSED): + if rid in first_seen: + convicted.append((n, rid, "", "CLOSED-ITEMS.md has this id twice (first at line %d)" % first_seen[rid])) + else: + first_seen[rid] = n + # RULE 3 — a finished row left in the OPEN register (2026-10-03) finished_in_open = [] for n, rid, columns, cells, _ in register_table.rows(OPEN): diff --git a/scripts/repo_gates.py b/scripts/repo_gates.py index dcef373b..32200d74 100644 --- a/scripts/repo_gates.py +++ b/scripts/repo_gates.py @@ -24,6 +24,8 @@ Gates, in order (all must pass; **non-zero exit on any failure**): 13. observations a REPORT.md observation with no register row behind it (R-389) 14. register-shape a register row whose state cell was eaten, a duplicated id, or a blank line splitting the table — the register mis-stating how many findings exist (R-627) + 15. script-tests every Python test suite under scripts/ (found by a walk), by exit code (R-885) + 16. decoy-coverage every registered gate in all four repos has a decoy, or a named exemption (R-421) **THE `GATES` TABLE BELOW IS THE LIST; THIS IS A POINTER TO IT.** It drifted once already — it read eleven while thirteen were registered, from 2026-08-24 until 2026-09-01, so diff --git a/scripts/test_gate_decoys.py b/scripts/test_gate_decoys.py index 74d09e3f..266e8ba3 100644 --- a/scripts/test_gate_decoys.py +++ b/scripts/test_gate_decoys.py @@ -25,6 +25,7 @@ Run from the repo root: python3 scripts/test_gate_decoys.py Exit 0 all decoys rejected · 1 a decoy passed (a live hole). """ import io +import re import json import os import shutil @@ -217,6 +218,20 @@ decoy("closed-register/unreadable-row", "closed_register_gate.py", append_to(os.path.join(ROOT, "documentation", "backlog", "CLOSED-ITEMS.md"), u"\n| **R-905** | A row with no state cell at all. |\n")) +# --- closed-register RULE 4 (2026-10-05, R-416): the same id twice in CLOSED-ITEMS.md --------------- +# The id is read from the file's FIRST closed row at run time, so the decoy duplicates a real id rather than an +# invented one that could never collide. +def _first_closed_id(): + for line in io.open(os.path.join(ROOT, "documentation", "backlog", "CLOSED-ITEMS.md"), encoding="utf-8"): + m = re.match(r"^\| \*\*(R-\d+[a-z]?)\*\* \|", line) + if m: + return m.group(1) + + +decoy("closed-register/duplicate-closed-id", "closed_register_gate.py", + append_to(os.path.join(ROOT, "documentation", "backlog", "CLOSED-ITEMS.md"), + u"\n| **%s** | The same id again. | CLOSED 2026-10-05 | none |\n" % _first_closed_id())) + # --- closed-register RULE 3 (2026-10-03): a FINISHED row left in the OPEN register --------------- # On 2026-10-03 the open register held 113 rows whose leading verdict was finished — a quarter of the # file. The decoy is that exact shape: a row a session closed in place and never moved. The genuine diff --git a/scripts/test_no_latest_push.py b/scripts/test_no_latest_push.py new file mode 100644 index 00000000..f7ff442f --- /dev/null +++ b/scripts/test_no_latest_push.py @@ -0,0 +1,27 @@ +#!/usr/bin/env python3 +"""R-345: nothing in this repo's build tooling tags or pushes an image as `:latest`. +Scans hub/Makefile, scripts/build-hub.sh and every Dockerfile/Makefile/*.sh under hub/ and scripts/ (a walk). +Run: python3 scripts/test_no_latest_push.py""" +import os +import re +import sys + +ROOT = os.path.dirname(os.path.dirname(os.path.abspath(__file__))) +BAD = re.compile(r"^(?![ \t]*#)[^\n]*?(\bdocker\s+(tag|push)\b[^#\n]*:latest\b|-t\s+\S*:latest\b)", re.M) +hits, scanned = [], 0 +for top in ("hub", "scripts"): + for dp, dns, fns in os.walk(os.path.join(ROOT, top)): + dns[:] = [d for d in dns if d not in (".git", "__pycache__", "node_modules")] + for f in fns: + if f in ("Makefile", "Dockerfile") or f.endswith(".sh"): + p = os.path.join(dp, f) + scanned += 1 + for m in BAD.finditer(open(p, encoding="utf-8", errors="replace").read()): + hits.append("%s: %s" % (os.path.relpath(p, ROOT), m.group(0).strip())) +if scanned < 5: + print("FAIL: scanned only %d build files — the scope is wrong" % scanned) + sys.exit(1) +if hits: + print("FAIL: a :latest tag/push in build tooling (R-345):\n " + "\n ".join(hits)) + sys.exit(1) +print("OK: %d build files, no docker tag/push of :latest" % scanned) diff --git a/scripts/test_repo_gates_docstring.py b/scripts/test_repo_gates_docstring.py new file mode 100644 index 00000000..ebb962a7 --- /dev/null +++ b/scripts/test_repo_gates_docstring.py @@ -0,0 +1,21 @@ +#!/usr/bin/env python3 +"""R-418: repo_gates.py's docstring list of gates names exactly the gates in GATES, in the same order. +It drifted twice (eleven listed while thirteen ran, 2026-08-24..09-01; then fifteen while seventeen ran). Run: +python3 scripts/test_repo_gates_docstring.py""" +import os +import re +import sys + +sys.path.insert(0, os.path.dirname(os.path.abspath(__file__))) +import repo_gates # noqa: E402 + +listed = re.findall(r"^ {1,3}\d+b?\. +([a-z][a-z0-9-]+) ", repo_gates.__doc__, re.M) +registered = [g[0] for g in repo_gates.GATES] +if listed != registered: + print("FAIL: the docstring lists %d gate(s), GATES registers %d" % (len(listed), len(registered))) + print(" only listed: %s" % sorted(set(listed) - set(registered))) + print(" only registered: %s" % sorted(set(registered) - set(listed))) + if set(listed) == set(registered): + print(" (same set, different order)") + sys.exit(1) +print("OK: the docstring lists the %d registered gates, in order" % len(registered))