diff --git a/documentation/architecture/05-hub-architecture.md b/documentation/architecture/05-hub-architecture.md index 730395b5..8f771562 100644 --- a/documentation/architecture/05-hub-architecture.md +++ b/documentation/architecture/05-hub-architecture.md @@ -419,7 +419,8 @@ Hungarian „öt szó" is correct and unchanged. ### 16.1 Form protection (CSRF) on both login paths (R-135) -The operator logs in two ways: a browser session (`hub_session` cookie + a per-session token on every form) and HTTP +The operator logs in two ways: a browser session (`__Host-hub_session` cookie since v0.138.0, R-136 — always Secure, `Path=/`, no Domain, so a sibling +subdomain cannot plant one; + a per-session token on every form) and HTTP Basic for scripts. Until v0.135.0 a state-changing request with NO cookie skipped the token check — on the reasoning that it must be a script. It need not be: a browser caches Basic credentials per origin and resends them on a cross-site form POST (SameSite does not govern the Authorization header). Now a request without a session passes @@ -439,9 +440,18 @@ Legacy rows are sealed in place at start-up (`SealLegacyRecoverySecrets`, measur refused; a wrong key → a reveal is a 500 with nothing in the body or the log. Both retrieval paths (the operator page and the global-key API) open through `GetHostRecoveryCredential`, so the break-glass route still works with the UI down. -**What a copy of `hub.db` still holds readable** (R-879): each box's hub API key, each household's owner passphrase -(`customer_configs.retrieval_password`) and API key, and the PBS-DR token values (`host_pbs_secrets.value`, kept after -use). **And the key is the other half:** a backup of the database restores a hub that can open the sealed columns only +**Since v0.138.0 (R-879) four more columns are sealed the same way:** each box's hub API key (`hosts.api_key`), each +household's controller API key and owner passphrase (`customer_configs.api_key`, `retrieval_password`) and the PBS-DR +token values (`host_pbs_secrets.value`). Hashing or deleting was not possible — every one of these values is served again +(re-enroll returns the key, the config ships the controller key, the operator page shows the passphrase, a PBS token can +be re-staged). The two API keys also carry an UNKEYED SHA-256 twin (`api_key_hash`, backfilled at every start), and a box +is looked up by that hash — **so a box authenticates even when the sealing key is missing or wrong** (`TestR879_BoxAuthSurvivesFailedSealing`). +Legacy rows are sealed at start-up (`SealLegacyBoxSecrets`, idempotent, non-fatal). A value that does not open marks the +record unreadable: its serve paths answer 500, a save refuses it (never blanks it), a PBS token is not consumed. **Roll-back +to a hub before v0.138.0** needs the columns opened first: `felhom-hub -unseal-box-secrets` (all or nothing; exit 1 = +nothing changed) — run in the running pod, then put the old image back. `guests.api_key` is an unused, always-empty +column and is not sealed. *Decided by CC unattended — operator may reverse (`09` §3 decision 132).* +**And the key is the other half:** a backup of the database restores a hub that can open the sealed columns only with the same `OFFSITE_SECRET_KEY`; today that key exists only on DooPlex (the k8s Secret, and the GPG secrets export on the same machine). The off-site plan for the database and its key: `runbooks/RUNBOOK-hub-db-offsite-backup.md` (R-173 — option A decided 2026-10-05, `09` decision 125; §16.3). diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index 571214d3..28b1dd9d 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -913,6 +913,19 @@ its length, and both fixes cost something the household would notice — operato (a 75 GiB drill box) and changes a behaviour operators rely on, with no measured floor behind 120. **Chosen (a)**; reversible by renaming back. Whether a REAL floor exists is unmeasured. Installer 1.32.0 (`installer-v1.32.0`, not cut). +132. **A hub with no sealing key, asked to write a NEW box secret (R-879).** Options: (a) refuse — enrolment, customer + creation, passphrase regeneration and a PBS mint fail until the key is back, existing boxes unaffected; (b) write it + in plaintext — silently re-creates the readable secret the row removes. **Chosen (a)**: `05` §16.2 already says + „no key → a save is refused" for every sealed column, and the production hub runs with the key. The box lookup is + an UNKEYED hash so that authentication never depends on the key. Hub v0.138.0. +133. **What removing a household's own off-site target does to its repository password (R-729, R-545).** Options: (a) + shred it always — a hub-held package then hits the R-241 mint refusal, and history on the NAS loses its on-box key; + (b) refuse the whole press while the hub holds a package — refused on practically every box, R-729 unsolved; (c) + clear the target, SSH key and known-host line always; keep the password whenever anything could depend on it (a + hub package, escrowed, a successful run, snapshots), delete it otherwise. **Chosen (c)**: R-241's rule (never drop + a key a package protects) and „never the repository"; keeping a key is reversible, deleting it is not. Cost: on + most boxes the 0600 password file stays with no target. `07`. Controller (next release). + ### 2026-10-05 (~21:00) — three rulings, the reviewer's picks given to the operator (recorded before the work; the burn-down night) 128. **R-126 — a `.fab` export onto a network drive.** **Refuse an export WITHOUT a password to any network drive; with diff --git a/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md b/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md index 8e15aeed..2cba79f3 100644 --- a/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md +++ b/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md @@ -32,3 +32,15 @@ cherry-picks onto `main`, writes CHANGELOG, closes rows, pushes and watches CI. | R-179 | fixed on main (NAS units; data-drive units left, GL6-F2) | 25 | `c11d4fbf` | | R-310 | fixed on main + runbook line | 15 | `b43705ac` | | R-274 | fixed on main (disclosure + pinning tests; check shipped with R-297) | 15 | `f790734e` | +| R-879 | fixed (hub v0.138.0) + roll-back command; decision 132; close after deploy | 75 | `5f060e3d`, `91e4ace9` | +| R-136 | fixed (hub v0.138.0); close after deploy | 20 | `e221476f` | +| R-283 | fixed (hub v0.138.0): box's own claim state shown; un-claim NOT taken (operator's call) | 30 | `daeed175` | +| R-540 | moved to D (multi-box design + money) | 10 | — | +| R-435 | moved to D (per-app counts, two repos, unmeasured threshold) | 10 | — | +| R-315 | fixed and closed (gate tooling is live on push) | 25 | `900e6d86` | +| R-325 | felhom.eu half fixed; controller half → helper ctrl-c | 15 | `7ee00d59` | +| R-422 | fixed and closed | 25 | `bf179847` | +| R-426 | partial (20 → 16 exemptions); row stays open | 15 | `68df84bc` | +| R-502 | moved to A (a Docker-running gate on DooPlex needs an operator word) | 10 | — | +| R-349 | hub half fixed (hub v0.138.0) | 20 | `a5b4d29a` | +| R-276 | hub half fixed (hub v0.138.0) | 15 | `0a460bcd` | diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 64d863c1..da91d8ff 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,15 @@ --- +## 2026-10-05 (night) — the burn-down night: gate fixes + +The full text of every row below: `git show 86def579:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-315** | **The wire-contract gate's positive control FAILS on the new wire: it checks name-presence, not decodability.** (P4) | CLOSED 2026-10-05 — FIXED (the burn-down night): mirrored roots are checked field by field | felhom.eu `900e6d86`: `scripts/wire_contract_gate.py` checks the receiver's mirror type for the hub→agent escrow/retained and hub→controller escrow-ACK roots and prints each root's check; decoys `wire-mirror/renamed` (convicts — the measured `superseded_at` rename) and `wire-mirror/genuine` (passes) in `scripts/test_gate_decoys.py`; red-proof: deleting the MIRRORS entry lets the renamed decoy pass. The two report wires stay name-reachability (printed as such). | +| **R-422** | **`reuse_refs_check.py` only checks citations whose extension is one of `go py html css yml yaml sh`.** (P4) | CLOSED 2026-10-05 — FIXED (the burn-down night): .md citations are checked | felhom.eu `bf179847`: `scripts/reuse_refs_check.py` checks `.md` citations; audits/ and documentation/tests are indexed for .md only; walk across all four repos' REUSE.md: 0 failures after one fix; tests `scripts/test_reuse_refs_check.py` (3) + the KNOWN-HOLE decoy now expects a conviction; red-proofs recorded in the night log. | + ## 2026-10-05 (night) — the burn-down night: stale rows The full text of every row below: `git show c8da8e04:documentation/backlog/OPEN-ITEMS.md`. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 799ba97e..fbe67094 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -120,7 +120,7 @@ stopping line that lies. | **R-179** | Install & onboarding | P3 | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** The teardown's residue-diff provenance (`day0-install.md` Part E: *"a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers"*) is from **v1.9.1**, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps `/etc/systemd/system/mnt-felhom\x2ddrives-.mount` and `.automount` after a full uninstall | **READY (S) — NEW 2026-08-03** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `c11d4fbf`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). Left: an enrolled DATA drive's mount unit still survives an uninstall (removing it unmounts customer drives — GL6-F2, its own question). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | **Observed on demo-hp 2026-08-03** after `--uninstall --vmid 9201`: `mnt-felhom\x2ddrives-Felhom\x2dShare.automount` **loaded failed failed**, its `.mount` `loaded inactive dead`, and `mnt-felhom\x2ddrives.mount` still `active mounted` — the uninstall's own output had warned `/mnt/felhom-drives/Felhom-Share is busy — NOT forcing` and `/mnt/felhom-drives root bind left mounted`, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, `daemon-reload`, unmount the autofs then the parent. **NEGATIVE CONTROL, same day:** demo-felhom's uninstall left **nothing** (`ls /etc/systemd/system | grep -i felhom` → only the unrelated `felhom-bootstrap.service`; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** `felhom-bootstrap.service` is NOT residue — it is the ISO first-boot unit, `disabled`+`inactive`, exactly-once and already fired | CC | | **R-180** | Install & onboarding | P3 | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** `felhom-host-install.sh` validates the archive storage EXISTS (`pvesm status --storage`, `:1583`) and that the golden volid RESOLVES on it (`:1661`), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default `local local-lvm felhom-pbs` (`--acl-storages`, which `runbooks/day0-install.md` tells the operator **not** to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step | **READY (S) — NEW 2026-08-03** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `7f3944ef`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | **Hit live on demo-hp 2026-08-03** during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on `felhom-backup` (the enrolled NVMe, where the box's vzdumps live) and `--archive-storage felhom-backup` passed. Pre-flight passed; steps 1–7 ran; step 8 returned `reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)`. **The cost is the ORDER, not the error** — by the time it fires, step 2 has minted the PVE token, step 4b has **rotated root@pam and vaulted it** (so the old console password is already dead), and step 5 has installed the agent. Recovery was `--resume` after moving the golden to `local`, which worked cleanly. **This is statically checkable in pre-flight**: `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a one-line assertion over two variables both known at `:1583`. Same class as R-29 — the checkable thing that nothing checks | CC | | **R-250** | Install & onboarding | P3 | **A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time.** Found 2026-08-07 creating the fifth walk's venue. `POST /configs/new` with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — **fail-closed by design** (`offsite.go:111-121`, *"don't serve a descriptor the controller can't verify"*), with `defaultScanBackoff` = 2+4+8+16+30 ≈ **60 s**, sized by its own comment to *"the observed DNS propagation lag"*. **Measured, both halves:** the first create exhausted the ladder — five `no such host`, then `dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable` (the name had just begun resolving, **AAAA-first, into a pod with no IPv6 route**) — and the hub logged `[ERROR] offsite provision for walk5`. An **identical second POST succeeded** ~70 s later on its own final rung (`shared already provisioned … (subaccount 285351)`). **Total settle ≈ 100 s against a 60 s budget.** The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, `SSH-2.0-OpenSSH_9.6p1`. **Two distinct things are wrong and should not be merged:** (1) the budget is sized against DNS *existence*, but what actually bit is the **AAAA-before-A** window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) **the operator is told nothing actionable** — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — `one_time_secrets` 1, one sub-account, no double-provision) **but that safety is invisible to the person deciding whether pressing again will double-charge them.** **Severity LOW-MEDIUM:** self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | **READY** — owner Viktor | — | — | operator | -| **R-283** | Install & onboarding | P3 | **After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page.** `customer_claims` for demo-hp still read `claimed_at 2026-07-21 16:29:25`, `generation 2`, `issued_at 2026-08-03` while the freshly provisioned guest — whose `settings.json` is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ **R-282**), and any previously issued code fails with *„Hibás vagy lejárt kód"* — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of **R-214/R-235** (an already-paired box still told to pair itself) | **READY (S) — NEW 2026-08-09** | — | Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page | CC | +| **R-283** | Install & onboarding | P3 | **After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page.** `customer_claims` for demo-hp still read `claimed_at 2026-07-21 16:29:25`, `generation 2`, `issued_at 2026-08-03` while the freshly provisioned guest — whose `settings.json` is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ **R-282**), and any previously issued code fails with *„Hibás vagy lejárt kód"* — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of **R-214/R-235** (an already-paired box still told to pair itself) | **READY (S) — NEW 2026-08-09** **2026-10-05 (burn-down night): FIXED in hub v0.138.0** (felhom.eu `daeed175`): the customer page shows the box's own claim state and warns on a mismatch. NOT taken: letting a report clear the hub's claim (reverses the set-only rule — the operator's call). Close after the deploy. | — | Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page | CC | | **R-306** | Install & onboarding | P3 | **`--preflight-only` says "no state written" and writes state — with an answer that can be wrong.** `_state_put` short-circuits on `DRY_RUN` only (`felhom-host-install.sh:418`), so a preflight-only run creates `/var/lib/felhom-install/state.json`. Observed live: after a run whose banner read `PRE-FLIGHT PASS (mode=byo) — no state written, no install step executed`, the file existed containing `{"completed": [], "dnsmasq_preexisting": "yes"}`. Both the banner and the flag's own comment at line 226 assert the opposite. **The harm is not the file, it is the value**: the runbook recommends preflight-only first, then the same command without the flag, so on a box carrying a Felhom leftover the wrong ownership answer is baked in before the real install begins | **READY (S) — NEW 2026-08-12, RANK 3** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `17af3b65`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | R-300, R-305 | Either make `_state_put` a no-op under `PREFLIGHT_ONLY` (and record ownership at install instead), or correct both claims. A comment asserting an invariant needs a test pinning it | CC | | **R-516** | Install & onboarding | P3 | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** | — | — | CC | | **R-554** | Install & onboarding | P3 | **[P3-LOW] Delete the first-boot setup wizard — obsolete by design, still reachable.** OPERATOR DECISION 2026-09-17 (localisation starter, decision 4: „out of scope, obsolete"). `02-controller-module-map.md` L56 calls `internal/setup/` obsolete; `cmd/controller/main.go` L322 still enters it when `setup.NeedsSetup(cfg)` — `customer.id` empty after bootstrap ingestion, or a `.needs-setup` marker (`internal/setup/setup.go` L17-25). Ingestion leaves `customer.id` empty on a missing/invalid `bootstrap.json`, a failed hub pull, or a failed merge/write/reload (`internal/bootstrap/bootstrap.go` L109-163) — so a box whose first boot cannot reach the hub shows a household an 8-page wizard (95 Hungarian strings, its own template set and CSRF). **Fix shape:** decide what such a box shows instead (a single „cannot reach Felhom yet, retrying" page — needs no decision beyond copy), then delete `internal/setup/` and `runSetupMode`; red-proof that a failed ingestion renders the waiting page, not a 404. **Check first** whether any drill/golden path still relies on `.needs-setup`. | **READY - rank P3-LOW; owner: CC** | — | — | CC | @@ -189,7 +189,7 @@ stopping line that lies. | **R-401** | Backup & restore | P3 | **Revisit the off-site integrity depth when a real store is LARGE — the default rests on ONE measurement, on ONE 134 MB store.** Controller v0.228.0 (R-399) made `--read-data-subset=100%` the default for every box. The whole justification is a single data point: `demo-hp`, 2026-08-30, 140 829 678 B / 2 651 blobs / 67 snapshots, structure 35.0 s vs 100% **39.2 s** — four seconds. Re-proven live 2026-08-31 at 38.7 s. **It does not extrapolate, and the reason is structural: the structure check's cost tracks the INDEX, a read-data run's tracks the DATA.** A 50 GB store is ~370x the data and this curve says nothing about it. **Nothing was invented from that one point** — no rotation schedule, no size threshold, no bandwidth budget — because four production designs in this project were specced against unvalidated mechanisms and all four were wrong. **`readDataSubsetRe` already accepts `n/m`**, so a rotating schedule (`1/7` on a different seventh each week) needs no parser work when the time comes; the missing input is a measurement on a large store, not code. **THE TRIGGER IS AN EVENT, NOT A DATE:** the slow-check WARN from v0.228.0 firing on any box (`integritySlowNoticeThreshold`, 5 min) — that line names the duration, the depth and this row. **WHAT HAPPENS IF NOBODY ACTS:** every box re-reads its entire store every week, however large it grows, and the first person to notice is a customer whose upload is saturated. **Whoever acts must also revisit `integrityCheckTimeout` (30 min)**, which is now the number a large store meets first. | **OPEN — WATCHING** | — | When the WARN fires: measure the curve on that store, then choose between a rotation (`n/m`), a size-conditional default, or leaving it. Do NOT choose from this row's numbers — they are the small-store case. | CC | | **R-409** | Backup & restore | P3 | **Nothing in the product can vouch for the bytes of a restored recovery unit — the only hash record covers 0.002 % of it.** MEASURED on demo-hp 2026-08-31 against kimai's restored unit: `manifest.json`'s `checksums` object carries sha256 for `.felhom.yml` (2 235 B), `app.yaml` (488 B) and `docker-compose.yml` (2 195 B) — **4 918 bytes of a 213 231 242-byte unit**. The database dump (48 217 B) and the two named-volume tars (160 331 776 B + 52 845 056 B) — the recoverable data, 99.998 % of the bytes — have no recorded hash anywhere. **And nothing else supplies one:** restic 0.14.0's `restore --verify` is a size-and-mtime reconciliation (a size-and-mtime-preserving one-byte corruption of a 160 MB tar passed clean, red-proofed), and `restic ls --json` file nodes in 0.14.0 carry name, size, mode, uid/gid and three timestamps and **no content hash**. **So "the restore produced correct files" is currently unanswerable by any automated means.** **What is NOT claimed here:** `restic check --read-data-subset=100%` already proves the STORE's packs, and the config files that ARE hashed are the ones a wrong-content failure would be hardest to spot in. | **OPEN — MEDIUM** | R-87, R-361 | Cheapest fix, and it is already half-built: extend the capture's `checksums` to cover `db_dumps` and `volume_dumps` — R-361 already computes a canonical dump sha256 to prove itself, so the value exists at capture time. Then a restore-test has a real reference and R-87's narrow version becomes a content check rather than a completeness one. Evidence: `audits/SPIKE-restic-restore-test-2026-08-31.md` §Q2, §Q3. | CC | | **R-433** | Backup & restore | P3 | **A sub-account cannot reach ANY Storage Box snapshot, by any name — so clause (b) of the 2026-09-01 R-95 re-scope ("the rest is recoverable file by file") is NOT SUPPORTED.** MEASURED 2026-09-01 on `demo-hp` over the credential the box already holds, read-only, no delete verb issued. **The sweep:** a batched `stat -c %n` over Hetzner's port-23 restricted shell (500–600 paths per round trip, stdout carrying only paths that exist) tried **777,600** names of the vendor form `YYYY-MM-DDTHH-MM-SS` across nine full days at second granularity, and 126 alternative shapes — **zero resolved.** **The control is what makes the zero mean anything:** the identical 600-name batch with one real path appended returned it in 6 of 6 batches. **The structural cause:** `/home` (the customer data, `u629488-sub3`) is **st_dev 0,82**; `/.zfs/snapshot` is **st_dev 0,276**; `/home/.zfs` does not exist. A snapshot under `/.zfs/snapshot` belongs to a different dataset than the one holding `felhom-repo`. **What still stands:** clause (a) — the box can delete its live repository but cannot WRITE into `/.zfs/snapshot` — is unchanged and re-confirmed. **What is now open again:** the only routes to the older copy are a panel rollback of the WHOLE Storage Box (deletes newer snapshots, hits every customer on it) and the provider API (fenced by §11-D, and `hub/internal/hetznerapi/hetznerapi.go` has **no snapshot method at all**, so it needs new code regardless). **NOT ESTABLISHED, and it is the question that decides whether per-file recovery exists for anyone:** whether the MAIN account can see the snapshots. No main-account credential exists in this project. **The ranking is Viktor's; I am not re-ranking R-95 — but the argument that moved it down is the argument this drill removed.** `audits/evidence-drill-r95-recovery-2026-09-01/` **BLOCKED-ON-PROVIDER 2026-09-01.** The one question that can move this is drafted and ready to send: **Question 1** of `documentation/runbooks/provider-questions-2026-09-01.md` — *can the MAIN account retrieve individual files from a snapshot, without a whole-box restore?* **Neither answer leaves this row where it is:** "yes" makes per-file recovery real but permanently operator-only, and closes this; "no, full restore only" means the snapshots do not bound a single customer's exposure at all, because using them costs every other customer on the box their newer snapshots — **and R-95 becomes urgent.** Nothing here can progress without it, and it is not CC's to send (§11-D). Tracked by a dated check in the DUE-CHECKS block. **⚠ A TRAP ON THE WAY TO THIS ANSWER, found by the operator 2026-09-01 and armed against in the questions file: HETZNER HAS TWO SIMILARLY-NAMED PRODUCTS WHOSE DOCS SAY OPPOSITE THINGS.** **Storage SHARE** (a managed Nextcloud, NOT us) documents *"Currently, we only support restores for the full backup ZFS snapshot to a specific point in time"* (`docs.hetzner.com/storage/storage-share/faq/backup-snapshot/`). **Storage BOX** (ours) documents the opposite — *"You can download individual files or entire directories as usual"* (`docs.hetzner.com/storage/storage-box/snapshots/`). **A web search for the obvious phrasing surfaces the SHARE page, and it reads like a definitive NO.** Tell them apart by the giveaways: the Share page talks about *Nextcloud's data cache*, a *database dump* and the *konsoleH* interface, and never mentions Storage Box. **Why this is filed and not left as trivia: a support agent could answer Question 1 from the wrong page, and that answer would push this row to the top of the register for no reason.** Question 1 now names the product, quotes the Storage Box line, and says explicitly that we know what the Share FAQ says. **If an answer cites full-snapshot-only, check WHICH PRODUCT it is about before acting on it.** **Nothing in this repository ever leaned on the Share claim** — verified by grep at the time; the only vendor line we cite is the Storage Box one. **RE-DATED 2026-09-15 → 2026-09-22 (due-checks gate):** the operator mailbox read through the Gmail connector (`(from:hetzner OR subject:hetzner OR "Storage Box") after:2026/09/01`) holds no Hetzner reply — one match, our own `offsite_snapshots_dropped` alarm. Whether the two tickets were ever sent is not visible to CC; if they were not, sending them is the operator's act. **-- ANSWERED. Hetzner replied on ticket #2026090103040671; the operator supplied the thread on 2026-09-22 after this session searched for it and WRONGLY reported no answer (see the instrument note below and R-628).** **Q1 - can the MAIN account read individual files and directories out of a specific snapshot, without restoring the whole box?** Hetzner: *"With the main account it should be possible to download files and directories from the snapshot as it is stated in the documentation"*, citing `docs.hetzner.com/storage/storage-box/snapshots#access-to-snapshots`; and separately *"A restore of a snapshot will revert the Storage Box completely to the state it was in when the snapshot was taken."* **So the Storage BOX documentation governs, not the Storage Share FAQ** - which is exactly the trap `provider-questions-2026-09-01.md` warned the reader about, and the answer came back on the right side of it. **File-level snapshot access is a MAIN-ACCOUNT capability; the sub-account cannot do it**, which matches R-433's own measured finding (777,600 names swept from a sub-account, zero hits). **HEDGE, kept because it is in the reply: "should be possible" is not "is", and nobody has yet read a file out of a snapshot from the main account on `u629488`. That is now the measurement this row needs, and it is cheap.** **Q2 - is `--append-only` enforced by Hetzner, or taken from what the client sends?** Hetzner: *"You can use --append-only and it would look like this: `command="rclone serve restic --stdio --append-only path/to/repo" `"*, with `fluix.one/blog/hetzner-restic-append-only/`. **So it is enforced by US, by pinning a FORCED COMMAND to the SSH key in the Storage Box's `authorized_keys`** - not by Hetzner globally, and not by anything the client asks for. A key without that prefix still gets a deleting server. **This is the answer R-95 has been blocked on** and it says append-only IS expressible on a Storage Box, which R-95 records as impossible over SFTP - see R-95 and R-436. **NOT YET MEASURED, and that distinction is the whole of what this row is worth: this is a support answer, not a proof.** What is owed is a real forced-command key on `u629488`, a restic `forget --prune` through it that is REFUSED, and a backup through it that still succeeds. | **BLOCKED** — **BLOCKED-ON-PROVIDER — Question 1 of `runbooks/provider-questions-2026-09-01.md`** | — | — | CC + operator | -| **R-540** | Backup & restore | P3 | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** | — | — | CC | +| **R-540** | Backup & restore | P3 | **[P3-LOW] The hub knows exactly ONE off-site pool box, so there is no rule for what happens when it fills.** Read from source 2026-09-16 while making off-site the default: `HETZNER_POOL_BOX_ID` is a single value, and every shared customer becomes a sub-account on that box. With off-site now ON for every new customer (hub v0.116.0) the box fills faster, and the fill warning (80%/90% of the box, `monitor/offsite.go`) tells the operator it is filling but nothing says which box a new customer should land on. **Needs a selection rule** (least-full, or explicit per-customer), not a bigger box. No customer is at risk today: the pool box read 0.3% full (2.7 GB of 1 TB), Σ shared quota 150 GB, oversub 0.15x. | **READY — rank P3-LOW; owner: CC (hub)** **2026-10-05 (burn-down night): NEEDS A DESIGN (and money).** A selection rule needs a multi-box configuration and a second box. Next: an operator pick of the rule and its trigger (e.g. add box 2 at 70 %). | — | — | CC | | **R-545** | Backup & restore | P3 | **[P3-LOW] There is no product action that un-configures an off-site target — only one that re-starts an ORPHANED repository.** FOUND 2026-09-16 while exercising R-543's paused state on the scratch guest: `POST /backup/offbox/config` configures a target and can disable it (`enabled` unchecked), but nothing removes it. `POST /backup/offbox/reset` refuses unless `OffboxOrphaned()` is true („Az offsite tároló nincs elárvult állapotban."), and it means „start a new remote backup, set the old history aside" — not „forget this destination". So a household that sets up the wrong NAS, or a box being handed to someone else, keeps the host, user, path, ssh key and minted repo password on disk with no route to clear them; a disabled target still holds its secrets in `data/offbox/`. **Why P3 and not higher:** a disabled target runs nothing and the secrets are 0600 on the box's own disk, so nothing leaks and no copy is lost. **Fix shape:** a „Távoli cél törlése" action beside the config form that clears the target and shreds `data/offbox/`, REFUSING while the hub holds a sealed package for this box (the R-241 rule — dropping the key would orphan the history that package protects). Teardown for this session's proof had to clear it out-of-band for exactly this reason, which is the measurement. | **READY — rank P3-LOW; owner: CC** | — | — | CC | | **R-548** | Backup & restore | P3 | **[P3-LOW] The whole-guest backup’s LOCAL tier cannot fit on a small-system-disk box, and will retry on that tier for ever.** MEASURED 2026-09-17 (chaos night) on `tester-1-022354`: a whole-guest backup wrote a **~29 GB source** (`mp0` = `local-lvm:vm-9201-disk-1`, 70 G provisioned, 40.58 % used, `backup=1`) into `pve-root`, which on a 32 GB system disk is **14 GB total with ~4.9 GB free**. Two samples thirty seconds apart showed the archive growing ~497 MB while free space fell ~475 MB — **~16 MB/s, i.e. under four minutes to a full `/`** on the nested PVE. **The product’s behaviour is correct and legible throughout:** it failed the tier and said which one — `whole_guest_backup_failed` (error, operator-only): „Whole-guest backup FAILED on the local tier — retrying with backoff (next attempt in 15m0s)” — its status surface agreed (`target_id:"local"`, `success:false`, `size_bytes:0`), and the **off-site tier then ran from the same snapshot and succeeded in ~8½ minutes, encrypted to ep0, consuming no local disk at all**. So the data still left the house. What is filed is the loop: on a box shaped like this the local tier can **never** succeed, and it keeps retrying on a backoff for ever, burning I/O and risking `/` each time. **Honest caveat:** the 32 GB system disk is this drill’s fixture choice, so the row is conditional on disk size — but nothing in the product checks whether the local target could ever hold the source before trying. **Fix shape:** compare the source size against the target’s free space before starting the local tier and skip it with a clear reason, rather than discovering it at ~16 MB/s. Evidence: `audits/evidence-chaos-night-2026-09-17/round-6.txt`. **-- 2026-09-30, the class on a real box: demo-hp's whole-guest LOCAL tier was refused by the space preflight (R-685) 10 times since 2026-09-27** („local has 12.1 GiB free; the last archive of guest 9201 was 8.9 GiB, so a new one needs about 12.1 GiB") — its newest local archive was 2026-09-29 04:42, 30 h old at the day's read (the weekly PBS tier, last 2026-09-24, is its whole copy meanwhile). The refusal is correct and named; what is missing is that nothing gives the box the room back (the host's `local` holds the 8.9 GiB archive of the only guest plus templates on a 39 GB root). `audits/pg-last-six-2026-09-30/C/`. | **READY — rank P3-LOW; owner: CC** | — | — | CC | | **R-645** | Backup & restore | P3 | **[P3-LOW] Lifting an update hold by the operator CLI lets the recovery unit be re-captured with the FAILED new definition within seconds — the copy the hold sentence names is overwritten.** MEASURED 2026-09-23 on 9202 during the undo bake-off: docmost was held at 07:59:16Z after a failed 0.95.0 → 0.96.0 update; `--clear-restore-hold docmost` + the controller restart it requires ran at ~07:59:23Z, and at **07:59:26Z** the controller logged *Recovery unit captured for docmost* — the unit's `compose/docker-compose.yml` now named `docmost/docmost:0.96.0`, the version that had just failed. The hold sentence had pointed the household at that unit („saját meghajtó, … 09:55"). The hold is what keeps the nightly legs off a held app (`isHeld`, v0.238.1); once it is lifted by hand, the checksum-gated refresh sees a changed definition and captures it. **Who it hits:** an operator who lifts a hold to inspect or repair, before restoring. With the undo (R-637) a failed update no longer holds unless the undo also fails, so the path is rarer — it does not go away. Candidate shapes, none chosen: the CLI refuses to lift an UPDATE hold (only a restore lifts it); or the lift also puts the pin back; or the capture skips an app whose pin is not what it is running. Evidence: `audits/undo-bakeoff-2026-09-23/docmost-40-undoF.txt` (the invalid run) and README §"Three things". **-- 2026-09-23 (controller v0.263.0):** the undo never reads the unit, so this no longer affects the automatic undo; it still affects an operator who lifts a hold by hand before restoring. | **OPEN — P3; owner: CC** | — | — | CC | @@ -232,13 +232,13 @@ stopping line that lies. | **R-861** | Security & access | P2 | **The agent's sudoers lets the agent user reach root without the operator key, so "root-minimized" (`03` §3) overstates it and the root-owned trust files (decision 93, the bundle's R17) are defence in depth, not a boundary.** READ 2026-10-04 from `felhom-agent/configs/felhom-agent.sudoers` (not exploited): `FELHOM_GUESTHOOK` installs `/tmp/felhom-guest-hook-*.sh` as a hookscript Proxmox runs as root at guest start, and `pct reboot` is granted; `FELHOM_INTERMEDIARY` installs a script + a systemd unit that run as root at boot; `FELHOM_ESCROW` runs `/usr/local/bin/felhom-agent` as root, and `FELHOM_SELFUPDATE apply` accepts a sha the agent itself passes. A compromised agent PROCESS is therefore root on its host. Fix direction: each of the four becomes a root-owned wrapper that checks its own input (fixed content or a signature), like `felhom-os-apply`; delivered by the config bundle. `11` §5.4.2, `03` §11. | **NARROWED 2026-10-05 — FIXED agent v0.146.1 for every root path found (nine, not four), delivered to demo-hp, demo-felhom and Tester 1 by a step bundle (R-880); live on both demo boxes: `sudo -l` 93/93 (64 commands allowed, 29 attacks refused — 23 of them allowed before), capability probe 67/67, a staged unit over /etc/sudoers.d refused. Design `03` §3.1, decision 122. LEFT, each named there: (a) the controller-swap image ref is guest-scoped (a compromised agent can run a chosen pinned-registry image in the guest); (b) the felhom-op SSH key is hub-delivered, not signed (felhom-op's sudo is scoped, not root); (c) the escrow ceremony hands the agent R by design (the box's PBS key). Tester 2: not delivered (offline).** | — | the operator decides whether (a)–(c) are accepted or need work before the first paying customer | CC | | **R-126** | Security & access | P3 | **A `.fab` bundle — plaintext secrets, OPTIONAL password — can be exported ONTO a NAS.** `storageDriveList()` (`internal/web/handler_export.go`) does not filter network paths | **READY — RULED 2026-10-05 ~21:00 (`09` §3 decision 128): refuse an export without a password to any network drive; with a password it is allowed. Being built (burn-down night).** | — | Split out of R-108, which closed without it: this is an explicit customer-chosen **export destination**, not a browsing surface reaching a backup tree, so it was never part of D5's precondition (`07` §7.3 records that reasoning). Was recorded inside R-108's row as its "second effect, independent of D5"; promoted to its own row so it does not vanish with R-108's closure. Fix = filter network paths out of the export destination list, or require the bundle password when the destination is a share | CC | | **R-132** | Security & access | P3 | **`curl -w '%{redirect_url}'` reconstructs the request URL WITH its basic-auth credential** — so a `-u ":$HUB_PW"` call that never put the password in a URL still printed it **Merged 2026-10-05 from R-350 (duplicate):** (1) 2026-08-20 occurrence: POST /configuration/artifacts answers 303; leak lives only in the CC transcript under ~/.claude/projects/ on DooPlex, not in git/evidence (checked then). (2) `-v` and `--libcurl` also re-render the credential, not only %{redirect_url}; confirm redirects with %{http_code} + follow-up GET. (3) Rotation path: hub /configuration form (current_password/new_password/confirm_password); DB override wins over ConfigMap (break-glass); CC can rotate file-to-file without printing (operator-present-one-time-secrets) if asked. | **WAITING-ON-OPERATOR** — **ACTION: rotate `HUB_PW`** **Folded R-580 2026-10-03** (the same `curl -w %{redirect_url}` credential echo, seen again 2026-09-18). | — | Happened on 2026-07-31 while red-proofing the R-120 gate: the hub operator password was written to the session transcript by the write-out format, not by the request. `-u` is safe; the *reporting* was not. Rule: read the redirect from `-D -` and grep `^Location:`, never `%{redirect_url}`, on any authenticated call. Rotate the hub password (`/configuration` → Login password; ConfigMap `auth.password_hash` is the reset path) and update `~/.config/credentials` | Viktor | -| **R-136** | Security & access | P3 | **Rename `hub_session` → `__Host-hub_session`** — makes cookie tossing structurally impossible | READY (XS, one line) | — | Verified on the live production response that all three prefix preconditions already hold: `Path=/`, `Secure`, no `Domain`. **Caveat for the ticket:** browsers reject a `__Host-` cookie without `Secure`, and `isSecure` is conditional on `r.TLS`/`X-Forwarded-Proto`, so plain-HTTP *browser* access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: `r.Cookie` returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 | CC | +| **R-136** | Security & access | P3 | **Rename `hub_session` → `__Host-hub_session`** — makes cookie tossing structurally impossible | READY (XS, one line) **2026-10-05 (burn-down night): FIXED in hub v0.138.0** (felhom.eu `e221476f`). Close after the deploy. | — | Verified on the live production response that all three prefix preconditions already hold: `Path=/`, `Secure`, no `Domain`. **Caveat for the ticket:** browsers reject a `__Host-` cookie without `Secure`, and `isSecure` is conditional on `r.TLS`/`X-Forwarded-Proto`, so plain-HTTP *browser* access to the hub would stop working (non-browser access uses Basic auth, unaffected). Tested consequence: `r.Cookie` returns the FIRST match and never tries the others, so a tossed cookie wins outright. Same audit §4.1-4.2 | CC | | **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | | **R-138** | Security & access | P3 | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) | READY (S) | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | | **R-255** | Security & access | P3 | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator | | **R-270** | Security & access | P3 | **R-268's stated rotation recipe is incomplete: the controller never re-reads `bootstrap.json`'s `local_api`, so a rotation leaves the agent channel dead across restarts.** `bootstrap.ensureLocalAPI` returns early when `cfg.LocalAPI.Endpoint != ""` — by design it FILLS an absent block and never refreshes a present one — so the token the controller uses lives in its own `controller.yaml`, not in the mount. Proved live 2026-08-09: two controller restarts after a correct `bootstrap.json` rotation, still `HTTP 401`; the channel came up only once `local_api.token` was written into `controller.yaml`. The neighbouring `DetectEndpointDrift` compares the ENDPOINT and deliberately does not compare the token (*"a token mismatch is a different failure"*), so this shape is knowingly unmodelled. Parent question — which file is authoritative — is **R-78** | **READY (S) — NEW 2026-08-09** | — | Either teach the drift detector the token, or make the rotation path write both files. Correct the R-268 row's recipe either way | CC | | **R-275** | Security & access | P3 | **`--uninstall` leaves five 0600 `agent.json.*` credential backups, and the reinstall hands them to the new service account.** `/etc/felhom-agent/` survives with `agent.json.{campaign8-before,campaign9-before,campaign9-prev,pre-e-target-move,pre-prunegate.bak}`, each carrying a 64-char `hub.api_key` and a 59-char `proxmox.token`. The teardown claims to remove *"config (+ its .bak backups)"* and `scripts/CHANGELOG` F1 records *"uninstall now purges the agent config's `.bak*` siblings (one held a live hub api_key)"* — **that fix does not match the filenames in use, and it misses `agent.json.pre-prunegate.bak`, a file that literally ends in `.bak`.** **Exposure assessed, not assumed:** these are SUPERSEDED — the orphaned key hashes to `a5d2222a…`, the hub's current demo-hp key to `8c59d1b6…`, and the Proxmox token was deleted by the same uninstall. **But the reinstall recreates `felhom-agent` at uid 999, the same uid the deleted account had**, so three of the backups become the new account's files — verified readable as `felhom-agent`. A fresh install's service account inherits read access to the prior install's credentials; superseded today, live if the backups were recent (R-179's precedent). **Also left, undeclared:** `/etc/felhom/{.bootstrap-done,appliance-pairing-code}`, `felhom-bootstrap.service` + `/usr/local/sbin/felhom-bootstrap.sh`, the `vmbr9` stanza in `/etc/network/interfaces`, and `/etc/sudoers.d/felhom-agent.bak-pre-e2a` (21 KB — **INERT: sudo skips dotted filenames, verified with `sudo -l -U felhom-agent`; `visudo -c -f` parsing it OK is NOT evidence sudo loads it**) | **READY (S) — NEW 2026-08-09** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `85de3f9b`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Purge by directory, not by glob; and do not let a new service account reuse a uid that owns old secrets | CC | -| **R-276** | Security & access | P3 | **RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing.** After `--uninstall` on demo-hp, `wg-quick@wg-felhom` is **enabled and active**, `/etc/wireguard/wg-felhom.conf` present, handshake to `167.233.158.164:443` **52 s old**, counters 5.86 GiB in / 2.48 GiB sent. It appears in **neither** the WIPED nor the KEPT list, though the BYO install disclosure names it prominently on the way in (*"an OUTBOUND WireGuard tunnel to the Felhom hub"*). A host told to leave Felhom retains a live network path into Felhom infrastructure, its hub-side peer registration intact, and nobody is told | **READY (S) — NEW 2026-08-09** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `85de3f9b`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). Left besides the tag: the hub deletes a removed host's WireGuard peer and pushes the peer list to the off-site endpoint (a hub change). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Tear the tunnel down and deregister the peer, or list it under KEPT with the reason and the removal command | CC | +| **R-276** | Security & access | P3 | **RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing.** After `--uninstall` on demo-hp, `wg-quick@wg-felhom` is **enabled and active**, `/etc/wireguard/wg-felhom.conf` present, handshake to `167.233.158.164:443` **52 s old**, counters 5.86 GiB in / 2.48 GiB sent. It appears in **neither** the WIPED nor the KEPT list, though the BYO install disclosure names it prominently on the way in (*"an OUTBOUND WireGuard tunnel to the Felhom hub"*). A host told to leave Felhom retains a live network path into Felhom infrastructure, its hub-side peer registration intact, and nobody is told | **READY (S) — NEW 2026-08-09** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `85de3f9b`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). Left besides the tag: the hub deletes a removed host's WireGuard peer and pushes the peer list to the off-site endpoint (a hub change). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. **Hub half in hub v0.138.0** (felhom.eu `0a460bcd`): host delete pushes the peer list at once. | — | Tear the tunnel down and deregister the peer, or list it under KEPT with the reason and the removal command | CC | | **R-338** | Security & access | P3 | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it **Checked from source 2026-10-05 (burn-down round 2):** nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246). | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-616** | Security & access | P3 | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** | — | — | CC + operator | | **R-717** | Security & access | P3 | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). | **OPEN — P3; owner: CC** | — | — | CC | @@ -251,7 +251,7 @@ stopping line that lies. | **R-870** | Security & access | P3 | **Tester 1's two Cloudflare credentials — the zone API token (`infrastructure.cf_api_token`) and the tunnel token (`infrastructure.cf_tunnel_token`) of the hub's `customer_configs` row `tester-1` — were printed into the 2026-10-04 night session's transcript** (not into any file): a read-only query selected `substr(config_json,1,400)`, and both values sit in the first 400 characters. Tester 1 is CC's disposable test customer (`enkicsifelhom.hu`). **Not rotated, by the operator's ruling of 2026-10-05 06:49 (option B).** **Rotation, whenever chosen (3 steps):** in the Cloudflare dashboard create a new API token for the `enkicsifelhom.hu` zone with the same permissions and refresh the Tester 1 tunnel's token (Zero Trust → Networks → Tunnels → the tunnel → refresh token) → hub → Configs → `tester-1` → Edit → the two Cloudflare fields → Save, then confirm on the box that cloudflared reconnected (`docker ps` health `healthy`) → delete the old API token. Rule for sessions (as R-831): never select a whole config row — name the fields, and never `config_json` without `json_extract` of a non-secret field. | **WAITING-ON-OPERATOR — rotation is his call (ruled: not now)** **Not rotated by the operator's rulings (2026-10-04 „keep using the current one"; 2026-10-05 option B) — restated 2026-10-05 18:23; the steps stay here.** | — | rotate when chosen | operator | | **R-525** | Security & access | P4 | **[P3-LOW] FileBrowser has its own login; putting it behind the dashboard session (traefik forwardAuth or Quantum proxy auth) is a new mechanism nobody has measured.** Filed 2026-09-15 by the P1-fixes task (B.5). R-513 closed the default-password hole with a generated password; a household still has two logins. **What it needs:** a spike on a scratch guest — forwardAuth to the controller session, and what FileBrowser Quantum does with a trusted header. | **READY — rank P3-LOW; owner: CC (spike)** **Re-ranked 2026-10-03: P3->P4: comfort feature needing a new unmeasured mechanism; the default password hole is closed.** | — | — | CC | | **R-779** | Security & access | P4 | **[P3-LOW] Part A's "two outside addresses seen as two" is proven through the simulated tunnel only; on the REAL tunnel the second outside address (ep0, one request allowed) was refused by Cloudflare's edge with 403 and never reached the box.** Measured 2026-10-01 19:51 UTC (`audits/visitors-2026-10-01/A/L2-demo-hp-real-tunnel.txt`): no log line on demo-hp; demo-hp's box has no geo restriction in its settings, so a Cloudflare ZONE rule (country or bot, not read) refused a German datacenter address. DooPlex's own address on the real tunnel was seen as itself. **Needs:** one sign-in from a second Hungarian address (the operator's phone off wifi) while DooPlex is locked out — 2 minutes; and say which Cloudflare rule refused ep0. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (a phone), CC reads the logs** **Re-ranked 2026-10-03: P3→P4: a proof gap on the real tunnel; operator-only follow-up.** | — | — | CC + operator | -| **R-879** | Security & access | P3 | **A copy of hub.db still holds secrets readable without any key:** each box's hub API key (`hosts.api_key`), each household's owner passphrase and API key (`customer_configs.retrieval_password`, `api_key`), and the PBS-DR token values (`host_pbs_secrets.value`, kept after use). Found 2026-10-05 while answering "what does a hub database backup now contain" (R-133 sealed the console passwords; the off-site passwords were sealed by R-821). A stolen database copy lets an attacker report as any box and read every owner passphrase. `05` §16.2 | READY | — | Seal or hash each with the same seal (the box keys and the passphrases are compared, so a hash may fit; the PBS token is served once, so it could be deleted after use) | CC | +| **R-879** | Security & access | P3 | **A copy of hub.db still holds secrets readable without any key:** each box's hub API key (`hosts.api_key`), each household's owner passphrase and API key (`customer_configs.retrieval_password`, `api_key`), and the PBS-DR token values (`host_pbs_secrets.value`, kept after use). Found 2026-10-05 while answering "what does a hub database backup now contain" (R-133 sealed the console passwords; the off-site passwords were sealed by R-821). A stolen database copy lets an attacker report as any box and read every owner passphrase. `05` §16.2 | READY **2026-10-05 (burn-down night): FIXED in hub v0.138.0** (felhom.eu `5f060e3d` + roll-back `91e4ace9`; `05` §16.2; decision 132). Close after the deploy's live check (sealed values in a snapshot with -wal, every box still reporting). | — | Seal or hash each with the same seal (the box keys and the passphrases are compared, so a hash may fit; the PBS token is served once, so it could be deleted after use) | CC | ## Box system & updates — 12 rows (P2 1, P3 11) @@ -263,7 +263,7 @@ stopping line that lies. | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC | | **R-242** | Box system & updates | P3 | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. **⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit message, in `hub/CHANGELOG.md` and in `REPORT.md` — a BYPASS, not a waiver**, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and these need one. **The operator was asked and ruled bypass-now-bake-later on 2026-08-30**, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. **That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it.** **OWED: bake a golden carrying 0.225.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change — `golden_version` + `agent_version` + `min_agent`; MinAgent is 0.129.0 per both CHANGELOG headers). **Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.** **⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0.** The `felhom.eu` push carrying that release's documentation used `git push --no-verify` on the operator's standing ruling from earlier the same day, declared in the commit and in `REPORT.md`. **The day-0 ground still holds for all three and was re-checked rather than assumed:** R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; **R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore.** **The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence.** **Owed: ONE bake carrying 0.226.0 covers all three** (`RUNBOOK-manual-build.md` §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. **✅ PAID THE SAME DAY — golden `0.226.1` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30).** Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`. `golden_currency_gate.py` went **red → green** on the same command, which is its proof that it measures something real. **The three declared bypasses above are now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back — **657 197 592 B, sha256 `70ed8e93…baefe69`**, both identical to what the bake reported — and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.226.1`, i.e. the delivered artifact naming the controller it will start. **The three-field vouch was checked deliberately, not assumed:** `MinAgent 0.129.0` read from the golden's controller CHANGELOG header, `agent_version 0.130.0 ≥ min_agent 0.129.0` (so NOT the R-216 shape), and the result verified by **re-reading the manifest** rather than trusting the flash — golden option `0.226.1 SELECTED`, all four shas matching. **The floor is proven ACTING, not merely set:** `demo-felhom` self-updated within 30 s, logging `[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1)`. **⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN.** v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that `felhom.eu` push used `--no-verify` and declared it, and golden **0.227.1** was baked, published, round-trip verified, **VOUCHED** and the floor **RAISED to 0.227.1** within the hour. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`. **THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day.** Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. **The floor was proven ACTING both times:** `demo-felhom` self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new `offsite-integrity` job **by itself, on a box nobody deployed to**, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. **This row's OTHER half is still open and untouched: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. **2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day.** v0.230.0 shipped in the morning with the newest golden at 0.229.0, **which is the build R-403 says deletes a good copy**, so the gate was red across four commits (`dddcc80`, `6e550ae`, `130f7a6`, `32a4c35`). Golden **0.230.0** baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; `demo-felhom` moved itself off the defective build unattended (`controller-swap: new controller healthy`, 16:21:40 CEST). Evidence: `documentation/tests/golden-0.230.0-2026-08-31/`. **The gate did its job and its own weakness surfaced doing it — R-410.** | **READY — the vouch half only** — owner Viktor **⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - **a BYPASS, not a waiver**, on the same reasoning as the five entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **The day-0 ground was RE-CHECKED rather than reused:** R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and `MinAgent` is unchanged at 0.129.0. **The ground still expires the moment a release changes first-boot behaviour.** **OWED: bake a golden carrying 0.229.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change - `golden_version` + `agent_version` + `min_agent`, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. **Golden and fleet delivery are the operator's (this row).** **✅ PAID THE SAME DAY — golden `0.229.0` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31).** Evidence: `documentation/tests/golden-0.229.0-2026-08-31/`. `golden_currency_gate.py` went **red to green** on the same command, which is its proof that it measures something real. **The `--no-verify` bypass declared above is now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back - **656 864 331 B, sha256 `39aa886d…d7bdae87`**, both identical to what the bake reported - and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.229.0`. **A THIRD independent reader agreed before anything was vouched:** the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. **The three-field vouch was checked deliberately, not assumed** (`MinAgent 0.129.0` read from the golden's controller CHANGELOG header; `agent_version 0.130.0` >= `min_agent 0.129.0`, so NOT the R-216 shape; `agent_sha256` and `wrapper_sha256` carried through explicitly because the handler clears a field it is not sent), and verified by **re-reading the manifest** rather than trusting the flash. The **R-120 gate passed rather than being bypassed** - fleet newest 0.229.0, golden 0.229.0. **The floor is proven ACTING:** `demo-felhom` self-updated `0.228.0 -> 0.229.0` and logged `settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0)` - **nobody deployed to that box.** **Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour.** This row's OTHER half is still open and untouched: **nothing gates the VOUCH itself** - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. **⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused.** Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). **The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here.** The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - a BYPASS, not a waiver. **OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor.** `demo-hp` was updated by hand; `demo-felhom` is still on 0.229.0 and still carries the defect. **See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported.** **2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN.** R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. **Nothing gates the VOUCH.** A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's `hub_settings` table, there is no copy in git, and a hub-reading gate could not be `--fast` so it would run in neither the hook nor CI. **Do not read R-404's closure as closing this.** **2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468).** The docstring's *"honest fix is a recorded waiver in the register, never a habit of bypassing"* is now a mechanism: `golden_currency_gate.py` reads `documentation/tests/golden-waiver.yml` (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. **What stays open on THIS row is exactly one thing: nothing gates the VOUCH.** The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). | — | — | operator | | **R-274** | Box system & updates | P3 | **A local golden is adopted with NO version and NO checksum check, so a reinstall can silently come up releases behind.** `felhom-host-install.sh` step 7: `if [[ -n "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then log_skip "using local golden"; return 0; fi` — **the hub manifest's `golden.sha256`, whose entire purpose is to vouch from a different trust root than Gitea, is consulted only on the FETCH path.** A locally-present archive bypasses the vouch: no version compare, no digest, no warning. On demo-hp 2026-08-09 the preflight selected `local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst`, whose baked marker reads `felhom-controller:`**`0.192.0`**, against a vouched golden of **0.210.0** — 18 releases stale. **The sharp consequence:** 0.192.0 is **below 0.200.0, where R-193's off-site recovery SCREEN shipped**, so a customer reinstalled today returns on a controller that cannot run the recovery ceremony their data depends on; it is also born below the managed-update floor (0.200.0), and the updater's auto-target is the floor, never the newest. **It compounds with the teardown**, which deliberately keeps the old golden (*"golden vzdump left in place"*). This is the R-111/R-115/R-120 drift family one layer down: the R-120 gate guards what may be VOUCHED, nothing guards what an install TAKES. **OBSERVED 2026-08-09, AND THE RESULT NARROWS THE ROW — recorded because it partly refutes what was written above.** On the RESUME path step 7 **fetched the vouched 0.210.0 correctly** (`fetching golden v0.210.0 from Gitea`), because `--resume` skips preflight and preflight is where local auto-discovery sets `GOLDEN_VOLID` (the GL6-F4 comment says so). **So the fresh-install and resume paths disagree on golden selection, and the resume path is the safe one.** Discovery is `… | sort | tail -1`, i.e. the NEWEST local archive by filename — a sensible heuristic, **and still no comparison against the manifest's version or sha**. The defect therefore stands as: *a box whose newest local golden predates the vouched one installs stale, silently* — which is exactly the state demo-hp was in before this run (newest local 0.192.0 vs vouched 0.210.0). It is now masked on this box because the freshly fetched 0.210.0 is the newest — **correct by recency, not by verification**. There are now **three** goldens on `local` (07-21, 08-03, 08-09), because the teardown keeps them. **Still not observed: a FRESH (non-resume) install taking a stale local golden.** | **VERIFY** (2026-10-03 triage: Local golden is now checked against the hub manifest's version and sha before use (GOLDEN_CHECK_WHY, R-297) — felhom.eu/scripts/felhom-host-install.sh:2855-2905) — **READY (S) — NEW 2026-08-09, NARROWED same day** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `f790734e`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). The version/checksum check itself shipped with R-297; tonight added the disclosure line and the tests that pin the check. **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Compare the local golden's version/sha against the manifest and refuse or re-fetch on mismatch; say so in the BYO disclosure, which today lists only what the install CREATES, never what it REUSES | CC | -| **R-349** | Box system & updates | P3 | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** **2026-10-05 (burn-down night): agent half on `main`** (felhom-agent `64f704d`: the host report carries `agent_sha256` of the running binary; test + red-proof). Left: the hub compares it with the vouched sha and shows drift; then a release. | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC | +| **R-349** | Box system & updates | P3 | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** **2026-10-05 (burn-down night): agent half on `main`** (felhom-agent `64f704d`: the host report carries `agent_sha256` of the running binary; test + red-proof). Left: the hub compares it with the vouched sha and shows drift; then a release. **Hub half in hub v0.138.0** (felhom.eu `a5b4d29a`). Left: the agent release that sends the field. | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC | | **R-444** | Box system & updates | P3 | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** **2026-10-05 (burn-down night): NEEDS THE OPERATOR.** A scheduled trim needs a new root grant (`pct fstrim `, widening what R-861 narrowed) and a new scheduled action on every customer guest's disks. Next: an operator yes on the grant and a cadence, then one measured run on demo-hp. | — | — | CC | | **R-468** | Box system & updates | P3 | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | — | — | CC | | **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC | @@ -281,7 +281,7 @@ stopping line that lies. | **R-333** | Monitoring & notifications | P3 | **Two disk-health questions the deploy raised and did NOT act on.** **(a) The 55/60 °C bands are SPINNING-DISK bands applied to NVMe.** They were adopted unchanged from the operator's Prometheus config so the two systems cannot disagree — a deliberate, stated decision — but **measured on demo-hp 2026-08-14 the healthy Toshiba KXG50PNV1T02 NVMe idles at 53 °C, two degrees below Figyelmeztetés and seven below Hiba**, and NVMe routinely exceeds 60 °C under load with no fault whatever. As it stands a healthy customer NVMe under sustained write can be reported as **Hiba** — the single worst outcome this feature can produce. **(b) The agent runs bare `smartctl -a -j` with no `-n standby`** (`felhom-agent/internal/storage/hostops.go:368`), so every poll WAKES a spun-down drive; going 6h → hourly multiplies that by six. demo-hp is all-flash so the cadence measurement could not reveal it, and it was recorded rather than acted on per the task's own instruction. Mitigating datum from the fixture: the failing drive logged only **3375 load cycles in 60505 hours** (~one per 18h), i.e. that duty cycle barely spins down at all | **READY (S each) — NEW 2026-08-14** | — | (a) split the temperature bands by device class, or drop them for NVMe and rely on `critical_warning`; (b) add `-n standby` to the agent's smartctl invocation (an agent change, so fold it into R-330's session) | Viktor decides (a); CC does (b) | | **R-340** | Monitoring & notifications | P3 | **The new reachability check does not touch the surface that actually failed.** R-339 reports when the hub cannot READ ep0 — but the read it performs is the `usage` op, which is `proxmox-backup-manager` plus `df` over SSH, and therefore rides the **local API daemon**. The 2026-08-18 incident explicitly CLEARED that daemon: `proxmox-backup.service` was healthy throughout, and it was the **HTTPS proxy on 8007** that was wedged with a full accept queue. **So R-339's check would have returned green for all 9 h 37 m of that outage.** It closes the case where ep0 is unreachable *as a host*; it does not close the case that actually happened. This is not a defect in R-339 — it is the honest boundary of what it watches, recorded so a future reader does not mistake a green box gauge for a working off-site tier | **READY (M) — NEW 2026-08-18** | a tenantsync endpoint-script version bump (the op is added on ep0, so it needs the same version-gated rollout `ErrUsageUnsupported` already models) | Add a **health op** to `scripts/felhom-tenantsync.sh` that probes `https://127.0.0.1:8007/` on ep0 and reports the proxy's fd count and listen-queue depth, then surface it as a third signal. **Overlaps the connections spike (R-336's remaining half): both want the same observations from ep0, so whichever runs SECOND must reuse the first's evidence rather than re-measuring a protected machine twice** **REUSE, per this row's own instruction — the connections spike ran FIRST (2026-08-20) and already produced most of what the health op wants; do not re-measure a protected machine a third time.** Available in `audits/evidence-ep0-established-connections-2026-08-20/`: the proxy **fd count** and its type breakdown (`lsof` + `/proc//fd`), the **listen-queue depth** (`ss -lnt` — `Recv-Q 0`, `Send-Q 1024`), the **ESTAB/CLOSE-WAIT split**, the **per-peer** connection histogram, a **31-minute persistence diff** of full 4-tuples, and a **46.18 h** slope with Poisson bounds. What the health op would still add beyond these: a loopback `GET https://127.0.0.1:8007/` probe — the observation that distinguished "process problem" from "network problem" on 2026-08-18 and the one thing this spike did NOT take, because it is the surface R-339 cannot see. **And this spike sharpens what the op should report:** a rising **ESTAB** count is the live signal (CLOSE-WAIT was **0**, not merely flat), and per **R-344** the fd ceiling that matters may be the **agent's**, not only ep0's. | CC | | **R-388** | Monitoring & notifications | P3 | **PRODUCT DECISION (not a defect): the customer notification model is the wrong shape, and the settings page grows by one toggle per detector.** The operator's framing, recorded verbatim 2026-08-23: *"A customer should be notified only about things they can act on or are responsible for — the drive they unplugged, the storage they filled. **A failed backup is our incident, not theirs.** The intended shape is that we detect it, we tell them we noticed and are dealing with it, and they are not handed an error they cannot solve. The subscription should feel like being looked after, not like being on call."* Today's page is the opposite shape — one switch per detector, and it **grew from 12 to 15 in a single session** (one new alarm plus two compound toggles split into four). That growth is the argument, not an aside: a page that grows per detector keeps asking a household to make engineering decisions. | **OPEN — DIRECTION, operator's call** | a decision on scope; nothing here is a bug | Recorded as a dated **[DESIGN — DIRECTION]** entry at `documentation/architecture/08-alarm-ladder.md` §8, marked plainly as *not current behaviour*. **Deliberately NOT implemented in the session that recorded it.** `app_start_failed` defaulting OFF is consistent with the direction and reversible either way, but was ruled on its own merits and does not pre-judge the redesign. | Viktor | -| **R-435** | Monitoring & notifications | P3 | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** | — | — | CC | +| **R-435** | Monitoring & notifications | P3 | **The snapshot-drop detector is blind to a single-app deletion — the exact shape the box can produce.** `snapshotDropFraction = 0.5` and `snapshotDropFloor = 5` (`hub/internal/monitor/offsite.go`) require a fall of MORE than half the previous count. demo-hp's baseline is **69** across 9 apps, so ~35 snapshots must go before it speaks; **one app's tag is ~9 and is invisible.** `offbox.go:1388` runs `forget --prune` **grouped by host,tags** — a per-tag wipe is precisely the shape a faulty retention or a targeted deletion produces. **This is deliberate, not accidental:** the constant's own comment argues the insensitivity, and *"a detector that cries wolf is switched off within a fortnight"* is a lesson this project paid for. **So this row is NOT a demand to lower the threshold.** It is a demand that the blind spot be written where the operator reads it, because "an unexplained fall is noticed within a day" (R-431, and `STATUS.md`) is true only of falls above half. **Discovered by arithmetic while planning the drill's Phase 2/3 pairing, which could not have worked:** Phase 2 deletes one app and Phase 3 expects the alarm to fire. **LIMITATION NOW WRITTEN INTO THE ALARM'S OWN DOCUMENTATION, hub v0.111.1** — the comment above `snapshotDropFraction` in `hub/internal/monitor/offsite.go` now states what the detector does NOT see, with the demo-hp arithmetic and the `forget --prune` grouping that makes the blind spot sit on the most likely single-app failure. **It also says explicitly that the numbers must NOT be lowered to "fix" this** and that per-app detection needs a SECOND signal keyed on the per-tag count. **The row stays OPEN because documenting a blind spot is not covering it** — and because `STATUS.md` and R-431 both still say "noticed within a day", which is true only of falls above half. | **OPEN — documented in code v0.111.1; the coverage gap itself is unclosed** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Covering one-app deletions needs per-app snapshot counts from the controller and a per-tag threshold measured on real counts — two repos, a mechanism nobody has measured. | — | — | CC | | **R-521** | Monitoring & notifications | P3 | **[P3-LOW] One unplugged drive sends the operator five e-mails and the household none.** MEASURED 2026-09-14 (BIGNIGHT F4, VM 333): `storage_disconnected (error)` at 21:58:02 CEST plus `app_start_failed (warning)` for each of the four apps the drive carries at 21:58:15, each with its own operator mail (hub log: five `Operator email sent`). The customer's mailbox (`tester1@felhom.eu`, read through the connector) received nothing; the household learns of it only on the dashboard, which is honest and says what to do. The apps' stop is a consequence of the drive event, so the four warnings add no information. **Fix shape:** suppress `app_start_failed` for apps stopped by a `storage_disconnected` (the dead-app check already knows the reason — „Hiányzó tárhely"), and decide whether a household gets a mail for a lost drive. **F6, 40 min later, the opposite failure:** a second, separate drive loss (20:38:33Z) produced `storage_disconnected (error)` and four `app_start_failed`, and the hub logged `Operator email suppressed … cooldown` for all five — **no mail at all for the second unplug**; only `health_degraded (warning)` mailed. A per-key cooldown that outlives the recovery (`storage_reconnected` came between them) silences a new incident. **F7:** the system disk at 95 % produced only `health_degraded (warning)`, whose operator mail was **suppressed by the cooldown** left by F6's `health_degraded` 15 minutes earlier; no disk-specific event reached the hub at all — the operator was not told the disk was nearly full. | **READY — rank P3-LOW; owner: CC (controller) · operator (customer mail policy)** | — | — | CC + operator | | **R-522** | Monitoring & notifications | P3 | **[P3-LOW] While the box has no internet, the dashboard's „Cloudflare Tunnel" tile keeps saying „Fut", and no page tells the household the box is offline.** MEASURED 2026-09-14 (BIGNIGHT F8, VM 333): VM 333's traffic off the LAN and to the hub was dropped at demo-hp's bridge 21:08:36 → 21:26:07Z. Throughout, the LAN dashboard (probed every 26 s from demo-hp) answered 200 and, polled every 2 min, showed no banner and the tile „Cloudflare Tunnel — Biztonságos internetkapcsolat — a szerver portnyitás nélkül érhető el kívülről. · **Fut** · Védett"; meanwhile cloudflared logged ≈ 20 errors every 2 minutes, the public name answered 530, and the controller logged `[report] Push failed … context deadline exceeded` and `Job hub-report failed: hub push failed after 3 attempts`. The tile reports the container, not the connection. A household whose remote access is gone sees „Fut". **Fix shape:** the tile reads the tunnel's connection state (cloudflared's registered connections or the report push result) and says „Nincs internetkapcsolat" when either fails. | **READY — rank P3-LOW; owner: CC (controller)** | — | — | CC | | **R-585** | Monitoring & notifications | P3 | **[P3-LOW] Six event producers still send Hungarian only, so an English household can see one Hungarian line in some mails.** FOUND 2026-09-18 by localisation slice 3 Part B (R-558, controller v0.256.0), which converted 19 of the customer-facing producers and left these: `backup_failed`, `db_dump_failed`, `backup_integrity_ok`, `backup_integrity_failed`, `offbox_enlarge_blocked` and `local_api_endpoint_drift`. **Why they were left:** each receives its sentence already FINISHED from another package, so the key and its arguments no longer exist by the time the notifier sees it — converting them means changing their callers, not the notifier. The 15 operator-tier types are deliberately excluded and are NOT part of this row: the operator reads Hungarian. **Why it matters more than it looks: `offbox_enlarge_blocked` has no `customerMessages` entry on the hub**, so its raw sentence IS the household's mail rather than an extra line under a translated headline — for that one type an English household gets a wholly Hungarian mail, not a mostly-English one. **Fix shape:** push the key and its arguments down from each caller (the shape slice 2 release B already used for errors, `util.MsgError`), then add each to the `convertedProducers` table in `internal/notify/message_customer_test.go`, which is the list both language tests walk. | **READY - rank P3-LOW; owner: CC** | — | — | CC | @@ -334,18 +334,16 @@ stopping line that lies. | **R-230** | Process & tooling | P4 | **Three instruction/memory follow-ups deliberately left by the part-2 session (2026-08-06), each needing a decision rather than an implementation.** (a) **A ruling is owed on auto-written staleness.** The hand-written `CLAUDE.md` files are now clean of version literals and expired blocks — the gate enforces it — but `MEMORY.md`, which Claude writes and which is the LARGER half of what loads (8.4k tokens vs the root file's 6.6k), carries **21 lines with component version literals**, **5 with bare host addresses**, and an entry still reading *"demo boxes REMOTE till ~08-02"* — the same expired-TEMPORARY class the gate was built to kill, now surviving in the one file the gate's content rules do not cover. **Partly actioned 2026-08-06 (close-out), and the ruling is STILL OWED:** the **three statements that were actively false** were corrected — `R-193 decision open` (closed 2026-08-05), `demo boxes REMOTE till ~08-02` (the box answers on the home LAN), `OPEN R-25b` (shipped 2026-07-21) — and gate check 6 now **WARNs** on version literals, host addresses, expired statements and stale-open citations in the index. WARN, never FAIL: Claude writes that file between sessions, so a hard failure would refuse a human's push over a line no human typed, and the warning is read by the model that will next edit it. **The remaining 32 version literals and 4 host addresses were deliberately left** for that loop. What is still owed is the bulk-correction ruling. **Correcting the premise:** the earlier report's "three expired statements" were all FALSE POSITIVES — each matched an ISO date inside a markdown link target, i.e. a filename — while the one real expired claim carried no ISO date at all. (b) **CLOSED 2026-08-06 (close-out)** — the workspace-root `CLAUDE.md` **is now a relative symlink** to the versioned copy, so the divergence class is gone rather than policed. Check 5 learned two shapes: for a link it asserts the target resolves to a real file (**a dangling link is worse than a diverged copy — the instructions load NOTHING and there is no content left to notice is wrong**), for two files byte-identity as before, so a clone elsewhere is unaffected. **Proven, not assumed:** three fresh sessions logged `session_start` for the link path, and a fourth **with no tools at all** quoted standing rule 1 verbatim — the content reaches the model, not just the path. (c) **The spec-as-failing-test pilot**, approved in principle and not started (was R-229(d)). | **READY** — owner Viktor | — | — | operator | | **R-288** | Process & tooling | P4 | **The capability map is too long to be read, and that is why it stops being true.** `architecture/00-capability-map.md` is **134 642 bytes / 19 456 words across 99 table rows in only 159 lines** — because the rows ARE the length. Measured, longest first: the unaided-recovery-journey row is **3 024 words**, the offsite-password-recovery row **1 087**, the unattended-restore-proof row **971**, the app/guest-network-failure row **904**. That single longest row is a novella of nested corrections, each appended rather than resolved. Its own verification stamp reads **2026-07-16 against evidence corpus @ felhom.eu tip `4b18cc5`** (line 23) — three weeks stale, which is the measurable consequence: nobody re-reads a row they cannot finish. **This is the project's memory, so restructuring it is surgery and wants daylight** — filed, deliberately not attempted in the 2026-08-09 session. **What the shape should probably be:** one line of status per capability plus a dated evidence pointer, with the argument moved to the audit it came from | **READY (M) — NEW 2026-08-09** | — | Do not fold this into another session; it needs its own. **SECOND CONCRETE COST, 2026-08-10 — and it is a different failure mode from the first.** An *execution record* — 33 package deletions on the operator's rule — was undiscoverable for two days because it lives **inside the row about the Configuration page being slow**. Two sessions searched for it: one reported "no register row records a package prune", the other exhausted the Gitea logs, the activity feed and the schema before concluding it might be unestablishable. It was in `OPEN-ITEMS.md` the whole time. **The first cost (2026-08-09) was two records that looked contradictory and were not; this one is a record that could not be found at all.** Illegibility now has two measured costs and they are different in kind: prose rows make claims ambiguous, and rows-about-other-things make facts unfindable. **The rule this earns is in CONTEXT.md: a record that lives inside a row about something else has not been recorded** | Viktor | | **R-290** | Process & tooling | P4 | **Most capability-map rows that back a green dot cite no evidence document at all — measured, 20 of 28 probed.** The page's *Walked* means *"done end to end on real hardware, evidence on file"*. Extracting the evidence column for the 28 rows behind the page's claims found a `tests/` or `audits/` path in **8**; the other 20 carry prose only. **Consequence, applied this session:** of 32 claims the page drew as Walked, **12 were downgraded to Built** because no walk document exists for them — `install.installer-by-tag`, `use.lifecycle`, `drives.enrol`, `drives.migrate`, `backup.tier1`, `backup.whole-machine`, `backup.restore-proof`, `fault.selfheal`, `fault.operator-email`, `fail.drive-filling`, `fail.lost-recovery-code`, `fail.hub-down`. **This is not a claim that those twelve are false** — several are near-certainly fine — it is a claim that nothing on file distinguishes them from an opinion, which is exactly what the status word promises. **The gate now enforces it going forward:** `scripts/check_stands.py` fails on `status: walked` with no `evidence:` source. **What is owed:** either a walk document per row, or an honest demotion in the map itself (the map is the source; the dataset only follows it) | **READY (M) — NEW 2026-08-09** | R-288 | The dataset was corrected; **the capability map itself still says PROVEN-LIVE for these rows** and is the thing to fix | Viktor | -| **R-315** | Process & tooling | P4 | **The wire-contract gate's positive control FAILS on the new wire: it checks name-presence, not decodability.** R-311 declared `hub -> agent (GET /escrow/retained)` as a fourth ROOT, and the gate's tag count rose 174 → 182, so the fields ARE inspected. But renaming the agent-side `superseded_at` json tag to `superseded_at_RENAMED` **still passed** — because the string `superseded_at` also occurs as a map key in the agent's local-API response, and the check is a repo-wide name search. The gate documents this ("name-reachability is not use"), so it is a known limit rather than a regression — but it means **declaring this wire bought documentation, not enforcement**, and a report that claimed coverage would have been wrong. The mutation was asserted to have applied before the run | **READY (M) — NEW 2026-08-12, RANK 3** | R-311 | Make the check resolve the RECEIVER'S mirror type and compare field-by-field, or state per-root which kind of check it got. **A gate whose positive control fails is an instrument nobody has calibrated** | CC | -| **R-325** | Process & tooling | P4 | **The shared copy vocabulary is imported by ONE of its two consumers, and drift-checked into the other.** `customer_copy_vocab.py` is the single list; `hub_copy_gate.py` imports it. **`felhom-controller/controller/scripts/retrieval_promise_gate.py` still carries its own `STEMS` literal**, because the session that created the shared module was under a hard end-state requirement to leave `felhom-controller` untouched — its target box was being re-deployed the same evening. **Two copies of a word list is not a theoretical risk in this project: it is the R-299 defect exactly**, where a guard asserted one inflection of a Hungarian verb and the plural walked past it. **So the gap is instrumented rather than left open: `hub_copy_gate.py` READS the controller gate's `STEMS` and FAILS if the two disagree** — single-source semantics tonight without a cross-repo edit. **Watched failing:** removing one stem from the shared list produced *"the shared vocabulary is no longer shared"* with both lists printed, and restoring it returned the gate to green. An ABSENT sibling clone is **INCONCLUSIVE (exit 2), never a pass** — the G-1 lesson. **This is a scaffold, not the destination** | **READY (S) — NEW 2026-08-13, RANK 3** | R-299, R-324 | Make `retrieval_promise_gate.py` import `felhom.eu/scripts/customer_copy_vocab.py` and delete its own literal — a felhom-controller change of a few lines, needing no bake (a gate is not shipped code). Then the drift check becomes redundant and should be removed with it, rather than left as a second mechanism nobody re-reads | CC | +| **R-325** | Process & tooling | P4 | **The shared copy vocabulary is imported by ONE of its two consumers, and drift-checked into the other.** `customer_copy_vocab.py` is the single list; `hub_copy_gate.py` imports it. **`felhom-controller/controller/scripts/retrieval_promise_gate.py` still carries its own `STEMS` literal**, because the session that created the shared module was under a hard end-state requirement to leave `felhom-controller` untouched — its target box was being re-deployed the same evening. **Two copies of a word list is not a theoretical risk in this project: it is the R-299 defect exactly**, where a guard asserted one inflection of a Hungarian verb and the plural walked past it. **So the gap is instrumented rather than left open: `hub_copy_gate.py` READS the controller gate's `STEMS` and FAILS if the two disagree** — single-source semantics tonight without a cross-repo edit. **Watched failing:** removing one stem from the shared list produced *"the shared vocabulary is no longer shared"* with both lists printed, and restoring it returned the gate to green. An ABSENT sibling clone is **INCONCLUSIVE (exit 2), never a pass** — the G-1 lesson. **This is a scaffold, not the destination** | **READY (S) — NEW 2026-08-13, RANK 3** **2026-10-05 (burn-down night): felhom.eu half on `main`** (`7ee00d59`: `hub_copy_gate.py` accepts the controller gate importing the shared list). Left: the controller gate imports it; then delete the drift check. | R-299, R-324 | Make `retrieval_promise_gate.py` import `felhom.eu/scripts/customer_copy_vocab.py` and delete its own literal — a felhom-controller change of a few lines, needing no bake (a gate is not shipped code). Then the drift check becomes redundant and should be removed with it, rather than left as a second mechanism nobody re-reads | CC | | **R-327** | Process & tooling | P4 | **The standing picture still describes a defect that has been fixed twice over.** Found by the first run of `unproven.py` (R-326), which is the argument for having built it. `where-felhom-stands.yaml`'s `claim.code-naming` is `status: partial` and its title reads *"The same word is used for two different secrets across three surfaces; the email points at a page a rebuilt machine does not show"* — **both halves of which are now false.** The box side shipped 2026-08-10 (R-295), the hub half and the page-naming fix on 2026-08-13 (R-295 hub, new `reenroll` mail kind), and the third near-homograph on 2026-08-13 (R-323). **NOT MOVED BY THIS SESSION, deliberately and by the dataset's own rule:** *"A status may not be RAISED here — if the evidence supports a stronger status than the capability map records, the MAP changes first and this file follows it."* Raising it here would be the exact inversion the file's header forbids, and the map edit is a separate judgement about what "walked" means for a naming change that no customer has yet met | **READY (S) — NEW 2026-08-13, RANK 4** | R-295, R-323, R-326 | Decide the capability-map status for the naming arc, then let the dataset follow it. **Note the honest difficulty: no customer has typed „Tulajdonosi jelmondat” yet**, so `walked` would be an over-claim; `built` is probably right, and the title needs rewriting either way because it describes a defect rather than a capability | operator + CC | | **R-377** | Process & tooling | P4 | **`CONTEXT.md`'s standing rulings are 188 KB in one section with no sub-headings, and that is why nobody reads them.** Measured 2026-08-22: `CONTEXT.md` is 217 KB, of which **187,913 bytes — 86% — is a single `## Standing rulings` section** carrying 39 `S-` ids and 153 bullets under **one** heading. **This session deliberately did NOT compress or split it**, and the reason is the ruling itself: `PROMPT-TEMPLATE.md` §3.4 and this file's own contract say the decision log is *dated, never edited afterwards*, and it is the only place that answers *"has this been proposed before, and why did we say no?"* — compressing it destroys exactly that. With no per-ruling delimiter, any mechanical split risks cutting a live ruling from its reason, which is the failure this whole arc is correcting. **So the problem is navigational, not volumetric, and the fix is structural: give each ruling a sub-heading with its `S-` id and date.** Then it can be linked, cited and found without a single word being edited. **13 mentions of SUPERSEDED already sit inside that blob** and cannot be separated from live text safely today. | **OPEN — LOW** | R-369 | Add per-ruling sub-headings only. Do not compress, do not reorder, do not edit any ruling's text. | CC | | **R-392** | Process & tooling | P4 | **No architecture document covers the two-AI workflow.** `documentation/architecture/` holds eight documents and **all eight cover the product** — topology, host agent, control-plane authorization, hub, off-site connectivity, backup, controller modules, capability map. Nothing records how the Claude.ai / Claude Code split works, what each side owns, how skills and `.claude/rules/` are scoped, or why. **The absence was found by trying to fill the template field, not by a survey:** the task that added the five process skills (2026-08-25) had to name an owning architecture document and could not, and the template requires that be recorded rather than passed over. The exposure today is low — the split is stable and both sides work — but it lives entirely in the operator's head and in chat, which is precisely the shape of a commitment nothing enforces. | **OPEN — LOW** | — | Write one architecture document for the agent-tooling layer: which AI owns which artifact class (`TASK-*.md`, `RUNBOOK-*.md`, validation), how skills are scoped and installed, what belongs in a `CLAUDE.md` versus a skill versus a rules file, and the reasoning for each boundary. **The rules themselves already exist** in `skills/felhom-doc-authoring/SKILL.md`; what is missing is the map of who owns what. Do not restate the doc-authoring rules there — point at that skill. | CC | | **R-394** | Process & tooling | P4 | **`felhom-build-deploy/SKILL.md` is 179 lines, over the 150-line limit its own repo now enforces.** Found 2026-08-25 by `scripts/check_skills.py` on its first run — the over-length was discovered BY the new checker, on the day the limit was written down, which is the checker working as intended. **It is not edited and not trimmed here:** the task that introduced the limit explicitly scoped the four pre-existing skills out, and trimming a build-and-deploy skill without exercising its commands is how a wrong command ships to a live host. **It is a named single-entry exception in `GRANDFATHERED` in `scripts/check_skills.py`, printed as a WARN on every run**, so it cannot fade; a NEW skill over the limit is convicted normally, and growing the set requires editing that file in a commit with a row to name. **The rationale for the limit** — attention thins across the excess, so the lines that matter are not the ones that survive — is in `skills/felhom-doc-authoring/SKILL.md` §5. | **OPEN — LOW** | — | Trim `felhom-build-deploy/SKILL.md` under 150 lines in a session that can VERIFY the commands it keeps, then delete its `GRANDFATHERED` entry in the same commit. The likely trim is the per-artifact command blocks moving behind a pointer to the runbooks, keeping the gotchas inline — but that is a judgement for a session with a build to run, not a line-count exercise. | CC | | **R-421** | Process & tooling | P4 | **THE CLASS: an instrument that matches a LABEL rather than the fact it names — five instances, every one found by accident.** R-410 (a `mkdir` turned the release gate green), R-400 (seven debug controls answering nothing), R-378 (a status word inside a sentence), R-419 (a phrase inside prose, including prose saying the marker was ABSENT), R-94 (a test comparing a constant to itself). **The gates are the machinery that enforces everything else in this project, and they were the one part nothing had ever checked.** The 2026-09-01 decoy sweep read all 29 scripts and fooled **16**. Ten were fixed the same day; four remain with rows (R-422..R-425); six could not be given a plausible decoy and are named. **The shapes, so the next one is cheap to recognise:** (1) name-for-fact — it matches a path or directory NAME while the fact lives inside the file; (2) substring-for-field — it matches a token anywhere in a body instead of in the field that carries it; (3) declaration-for-reachability — it checks a thing is declared, not that it RESOLVES; (4) constant-for-measurement — it compares a value against itself. **The single largest cause was mundane:** eight gates set their SCOPE with `os.listdir` (one level), so every one was green and correct today and would have gone blind the moment anyone added a subdirectory. `decoy_coverage_gate.py` now refuses a new gate that ships without a decoy. | **OPEN — the class row; it stays open as the place the next instance is recorded** | — | — | CC | -| **R-422** | Process & tooling | P4 | **`reuse_refs_check.py` only checks citations whose extension is one of `go py html css yml yaml sh`.** A cited `.md` path that does not exist is invisible — MEASURED 2026-09-01: `documentation/architecture/99-does-not-exist.md` added to `REUSE.md` passed, while the `.go` control was correctly convicted. REUSE.md and the CLAUDE.md files cite `.md` paths routinely, so this is the common case, not an exotic one. Fix: widen `PATH_RE`, then walk the false positives it produces across all four repos — that pass is the work, not the regex. The decoy is kept in `scripts/test_gate_decoys.py` asserting TODAY's behaviour, so the day this is fixed the test fails and is updated deliberately. | **OPEN** | — | — | CC | -| **R-426** | Process & tooling | P4 | **The decoy-coverage exemption list — 20 registered gates that ship WITHOUT a decoy test, each named.** `scripts/decoy_coverage_gate.py`'s `EXEMPT` map is debt, and this row owns it so it lives in the register and not only in a Python literal. **Four kinds:** (a) genuinely covered in the 2026-09-01 sweep but not yet moved into a suite — `hub-copy`, `instructions`, `docker-v`, `image-pins`; (b) blocked by an open hole and therefore un-assertable as rejecting — `site` (R-423), `one-register` (R-424), `offbox-rename` (R-425); (c) shared scripts whose decoy lives in `felhom.eu` and is counted there — `reuse-refs`, `instructions`, `observations` in the controller and agent runners; (d) **no plausible decoy constructed yet** — `hostinstall`, `wire-contract`, `due-checks`, `published`, `image-resolvable`, `volume-persistence`. Group (d) is the honest unknown: six gates whose soundness is UNTESTED, not established. **The list is green today and shrinks; a NEW gate with no decoy fails immediately.** | **OPEN — 20 names; group (d) is six untested gates** | — | — | CC | +| **R-426** | Process & tooling | P4 | **The decoy-coverage exemption list — 20 registered gates that ship WITHOUT a decoy test, each named.** `scripts/decoy_coverage_gate.py`'s `EXEMPT` map is debt, and this row owns it so it lives in the register and not only in a Python literal. **Four kinds:** (a) genuinely covered in the 2026-09-01 sweep but not yet moved into a suite — `hub-copy`, `instructions`, `docker-v`, `image-pins`; (b) blocked by an open hole and therefore un-assertable as rejecting — `site` (R-423), `one-register` (R-424), `offbox-rename` (R-425); (c) shared scripts whose decoy lives in `felhom.eu` and is counted there — `reuse-refs`, `instructions`, `observations` in the controller and agent runners; (d) **no plausible decoy constructed yet** — `hostinstall`, `wire-contract`, `due-checks`, `published`, `image-resolvable`, `volume-persistence`. Group (d) is the honest unknown: six gates whose soundness is UNTESTED, not established. **The list is green today and shrinks; a NEW gate with no decoy fails immediately.** | **OPEN — 20 names; group (d) is six untested gates** **2026-10-05 (burn-down night): PARTIAL** (`68df84bc`): the instructions gate's decoy is in the suite; the exemption list is **16** (was 20; wire-contract has decoys now). The row stays open for the 16. | — | — | CC | | **R-488** | Process & tooling | P4 | **[P3-LOW] `go test ./internal/backup` takes 5½ minutes: 89 off-site tests wait on real clocks.** MEASURED 2026-09-13 (`-v` timings, run alone: 581 tests, 333 s in total, 89 of them ≥ 1 s — `TestOffbox*`, `TestOffbox3a*`, `TestOffboxRun*`, `TestR4xx*` reconstitute fixtures at 3–8 s each). The controller's per-commit gate is therefore ~6 minutes, most of it sleeping, and two concurrent runs of the package looked like a hang. **Fix shape:** the waits are `waitForHealthy`-style polls and retry back-offs with fixed durations; make them seams the fixtures shorten (the R-457 rule: one clock). Not a correctness defect. **Checked from source 2026-10-05 (burn-down round 2):** The fixed real-clock waits named in the fix shape are unchanged: felhom-controller@114ff27 controller/internal/backup/restore.go:208-227 waitForHealthy has hard-coded `interval := 5 * time.Second` and `time.Sleep(3 * time.Second) // initial settling time`, called from offbox_reconstitute.go:927, tier2_restore.go:443, restore.go:90, restore_unit.go:471. No commit since 2026-09-13 touching internal/backup mentions R-488/test speed. Runtime (333 s) was NOT re-measured here (read-only). | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: developer speed only.** | — | — | CC | | **R-492** | Process & tooling | P4 | **[P3-LOW] `cfg.Paths.HDDPath` is empty on every box and still has readers; delete it.** R-465 audited its six readers and found every one falling back; R-490 (v0.242.0) gave the last one, `systemInfo`, the same fallback. The global now carries no information on any box and its deletion was deferred twice. **Fix shape:** remove the field, its env binding and the readers' fallback branches; a build proves nothing reads it. Next controller release. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: dead-code cleanup; every reader already falls back correctly.** | — | — | CC | -| **R-502** | Process & tooling | P4 | **[P3-LOW] The bootstrap regression harness is run by NO gate and NO CI — and it had never exercised the pairing banner.** MEASURED 2026-09-14: `grep -rn bootstrap-modes scripts/*.py .gitea/workflows` → nothing; the harness is run by hand in `felhom-iso-assistant:trixie`. While adding the R-496 checks, its fake hub's `register` reply turned out to carry no `pairing_code`, so `print_pairing_banner` returned early in every run: the banner that a household reads first had no test at all until ISO v1.27.0. **Fix shape:** register the harness in `repo_gates.py` behind a container-available check that reports INCONCLUSIVE (never skip-as-pass) when docker is absent, with a decoy. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: test harness coverage; no household meets it directly.** | — | — | CC | +| **R-502** | Process & tooling | P4 | **[P3-LOW] The bootstrap regression harness is run by NO gate and NO CI — and it had never exercised the pairing banner.** MEASURED 2026-09-14: `grep -rn bootstrap-modes scripts/*.py .gitea/workflows` → nothing; the harness is run by hand in `felhom-iso-assistant:trixie`. While adding the R-496 checks, its fake hub's `register` reply turned out to carry no `pairing_code`, so `print_pairing_banner` returned early in every run: the banner that a household reads first had no test at all until ISO v1.27.0. **Fix shape:** register the harness in `repo_gates.py` behind a container-available check that reports INCONCLUSIVE (never skip-as-pass) when docker is absent, with a decoy. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: test harness coverage; no household meets it directly.** **2026-10-05 (burn-down night): NEEDS THE OPERATOR.** The harness runs `docker run`; registered as a gate it would run Docker on DooPlex on every full gate run. Next: an operator word that a non-fast container gate may run there (INCONCLUSIVE without docker; never in --fast — CI has no docker). | — | — | CC | | **R-507** | Process & tooling | P4 | **[P3-LOW] The proof-install harness cannot drive the graphical installer, so a release's graphical entry is proven only up to its password screen.** MEASURED 2026-09-14 on VM 332 (ISO 1.27.0): `qm sendkey 332 tab` did not move focus (both password copies landed in one field), `mouse_move 1237 772` + `mouse_button 1` did not move the cursor or press Next, while `alt-n` did advance a page. The TUI entry is fully drivable. The gate's "proof install on BOTH menu entries" was met for 1.26.1 (by a person) and not for 1.27.x. **Fix shape:** measure QEMU `input-send-event` with absolute coordinates, or a VNC client on DooPlex; until then a release's graphical proof is an operator click-through. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: release-test tooling; an operator click-through covers it.** | — | — | CC | | **R-569** | Process & tooling | P4 | **[P3-LOW] Four more API handlers pick their status code by matching ENGLISH words in an error — the same shape as R-553, one language over.** FOUND 2026-09-17 while fixing R-553: `controller/internal/api/router.go` matches `"protected"`, `"not found"`, `"not deployed"`, `"still running"`, `"not orphaned"` in `err.Error()` at the stop/start, remove and orphan-cleanup handlers (three separate blocks). These strings are internal English, so localisation does not move them — the risk is a reworded internal error, not a translation, which is why this is P3 and was NOT folded into R-553's release. **Fix shape:** the same `util.KindErrorf` sentinels in `internal/stacks` (`ErrProtectedStack`, `ErrStackNotFound`, `ErrStillRunning`, …), a `statusFor` helper per handler family, and one table test per family passing a reworded message. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: internal robustness; only a future rewording would break it.** | — | — | CC | | **R-574** | Process & tooling | P4 | **[P3-LOW] `web/handler_debug.go` mixes page copy with JSON payload, so neither half could be converted safely.** FOUND 2026-09-18 by localisation slice 2 release A (R-557): the file holds 39 Hungarian literals and the inventory classifies them by STATEMENT, not by data flow (`I18N-INVENTORY-2026-09-17.md` §4), so which are section headings the debug page renders and which are values inside a diagnostic dump the operator copies out is not established. Converting a dump value would change what an operator pastes into a report; leaving a heading Hungarian leaves a half-English page. **Fix shape:** walk the file once and label every literal page-copy or payload in the same table slice 2 release A used, then convert only the page-copy half. Belongs to slice 2 release B or C. | **READY - rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3->P4: debug page is operator-facing.** | — | — | CC | diff --git a/hub/CHANGELOG.md b/hub/CHANGELOG.md index 9dbbbbd4..41d9e4c2 100644 --- a/hub/CHANGELOG.md +++ b/hub/CHANGELOG.md @@ -1,3 +1,35 @@ +## v0.138.0 — secrets sealed at rest, a cookie no sibling can plant, the box's own claim state, agent-binary drift (burn-down night: R-879, R-136, R-283, R-349, R-276) (2026-10-05) + +**Operator action on deploy: log in again once** (the session cookie is renamed). **Roll-back needs one extra step +first:** `felhom-hub -unseal-box-secrets` in the running pod (see R-879 below and `05` §16.2), then the old image. + +- **R-879 (security):** the box API keys (`hosts.api_key`), the controller API keys and owner passphrases + (`customer_configs.api_key` / `retrieval_password`) and the PBS-DR token values (`host_pbs_secrets.value`) are sealed at + rest with the R-821/R-133 seal (`OFFSITE_SECRET_KEY`). The API keys get an `api_key_hash` lookup column (SHA-256, + backfilled at every start without the key), so a box authenticates even when the sealing key is missing or wrong; + legacy rows are sealed at start-up by `SealLegacyBoxSecrets` (idempotent, non-fatal). A sealed value that does not open + marks the record unreadable: the config/recovery/artifacts/enroll/appliance paths answer 500, saves refuse it (no + blanking), and a PBS token is not consumed. No key: new secrets are refused, never written in plaintext (`09` §3 + decision 132). Tests `TestR879_*` (store, api, configgen, cmd/hub); six red-proofs. +- **R-879 roll-back:** `felhom-hub -unseal-box-secrets` (same `-config` and `OFFSITE_SECRET_KEY` as the server) opens the + four columns back to plaintext and exits — all or nothing (exit 1 = a value did not open, nothing changed; exit 2 = key + missing), idempotent, counts only in the log; `api_key_hash` is kept, so 0.137.0 can be put back. +- **R-136 (security):** the operator session cookie is `__Host-hub_session` — always Secure, `Path=/`, no Domain — so a + sibling subdomain cannot plant a session cookie the hub would read first. Every operator is logged out once; a browser + login over plain HTTP no longer keeps a session (accepted); Basic auth for scripts is unchanged. +- **R-283:** the customer page's claim card shows the box's own claim state from its latest report (claimed / not claimed + yet / unknown), with an amber warning when the hub says claimed and the box says not — the shape a guest rebuild + leaves. Nothing un-claims: the hub's claim stays set-only by design. +- **R-349 (hub half):** the host page compares the agent's reported `agent_sha256` (the running binary, agent R-349) with + the vouched agent sha when the box reports the vouched agent version — „matches vouched" or an amber DRIFT naming both + hashes. Empty hash (older agent) = unknown, never drift. +- **R-276 (hub half):** deleting a host asks the WireGuard peer sync for an immediate full-list push (as the customer + delete does since R-600), so the off-site endpoint drops the removed box's peer at once instead of on the next tick. +- **scripts:** `wire_contract_gate.py` checks the mirrored escrow roots field by field and prints each root's check + (R-315); `hub_copy_gate.py` accepts the controller gate importing the shared vocabulary (R-325, felhom.eu half); + `reuse_refs_check.py` checks `.md` citations (R-422); the instructions gate's decoy is in the suite, exemptions 20 → 16 + (R-426, partial). + ## v0.137.0 — small fixes from the burn-down: right numbers, right flashes, one create per press, the delete says when it opens (R-277, R-581, R-600, R-544, R-855, R-134, R-92, R-292, R-599, R-725, R-728, R-208; R-124 fixture) (2026-10-05) **Operator action on deploy: none.** One new operator flash (`artifact_version_missing`) and one new customer i18n key