diff --git a/REUSE.md b/REUSE.md index 5b5d0d45..ce8cd451 100644 --- a/REUSE.md +++ b/REUSE.md @@ -198,6 +198,7 @@ | `api.ConfigTemplateProvider` | hub/internal/api/handler.go (~L24) | `web.TemplateFetcher` (Gitea-pulled controller.yaml template) | stub providers in api tests | | `api.LatestVersionProvider` | hub/internal/api/handler.go (~L31) | `web.VersionChecker` (registry poll) | hub/internal/api/config_version_ack_test.go | | `mailRateLimiter.now` (func seam) | hub/internal/api/mail.go (~L27) | `time.Now` | hub/internal/api/mail_test.go clock injection | +| Installer behaviour tests (function lift + recording PATH stubs) | scripts/test_hostinstall.py | functions copied word for word out of `scripts/felhom-host-install.sh` | the same file — use it for any new installer function (BusyBox-safe; the `script-tests` gate runs it) | | `web.tenancyProvisioner` | hub/internal/web/pbsdr.go | `*tenantsync.Client` (pinned SSH to ep0's felhom-tenantsync) | `fakeTenancy` in hub/internal/web/pbsdr_test.go; in-process SSH server in hub/internal/tenantsync/client_test.go | | Cross-repo: ep0 tenancy surface | `scripts/felhom-tenantsync.sh` (JSON stdin/stdout forced command) | installed on ep0 per runbook offsite-endpoint.md §10 | provision/reissue/fingerprint ops; token secret rides stdout ONLY; the peersync script/key are untouched | | Cross-repo: controller → hub | `POST /api/v1/report` (frozen) + `POST /api/v1/event` | felhom-controller repo | new event types MUST enter `allowedEventTypes` (hub/internal/api/handler.go ~L1063) or the controller gets 400 | diff --git a/documentation/architecture/09-update-architecture.md b/documentation/architecture/09-update-architecture.md index f9565507..571214d3 100644 --- a/documentation/architecture/09-update-architecture.md +++ b/documentation/architecture/09-update-architecture.md @@ -905,6 +905,14 @@ its length, and both fixes cost something the household would notice — operato 127. **The agent's three by-design abilities (`03` §3.1) stay for now**; revisited before the first paying customer. *Operator ruling 2026-10-05.* (R-861) +### 2026-10-05 (night) — decided by CC unattended — operator may reverse (the burn-down night) + +131. **What the installer's 120 GiB local-lvm check is (R-130).** Options: (a) rename it a recommendation and say in the + warning that the install continues — matches what `runbooks/day0-install.md` already documents („proceed only if you + sized the grows deliberately"), costs nothing; (b) make it refuse — blocks small boxes that install and run today + (a 75 GiB drill box) and changes a behaviour operators rely on, with no measured floor behind 120. **Chosen (a)**; + reversible by renaming back. Whether a REAL floor exists is unmeasured. Installer 1.32.0 (`installer-v1.32.0`, not cut). + ### 2026-10-05 (~21:00) — three rulings, the reviewer's picks given to the operator (recorded before the work; the burn-down night) 128. **R-126 — a `.fab` export onto a network drive.** **Refuse an export WITHOUT a password to any network drive; with diff --git a/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md b/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md index cbda53ab..8e15aeed 100644 --- a/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md +++ b/documentation/audits/night-burndown-2026-10-05/NIGHT-LOG.md @@ -19,3 +19,16 @@ cherry-picks onto `main`, writes CHANGELOG, closes rows, pushes and watches CI. | R-375 | closed — ruling (decision 130) | 2 | (this commit) | | R-126 | ruling recorded (decision 128); build → helper ctrl-a | 2 | (this commit) | | R-856 | ruling recorded (decision 129); build → helper ctrl-a | 2 | (this commit) | +| R-349 | agent half fixed (`64f704d`); hub half → helper hub; row open | 15 | felhom-agent `64f704d` | +| R-25 | agent half fixed (`769c4c3`); controller half → helper ctrl-b; row open | 25 | felhom-agent `769c4c3` | +| R-531 | closed-stale (all three facts measured, ruled, built) | 10 | this batch | +| R-444 | moved to A — needs a new root grant + a scheduled action on customer disks | 10 | — | +| R-99 | moved to A — deleting on a customer datastore is the row's own separate ruling; ep0 fenced | 5 | — | +| R-270 | agent: no agent half (controller lane) | 5 | — | +| R-275, R-276, R-881 | fixed on main (installer 1.32.0); ships with tag `installer-v1.32.0` | 35 | `85de3f9b` | +| R-306 | fixed on main (installer 1.32.0) | 15 | `17af3b65` | +| R-130 | fixed on main — decision 131 (renamed, not enforced) | 15 | `efe76093` | +| R-180 | fixed on main | 20 | `7f3944ef` | +| R-179 | fixed on main (NAS units; data-drive units left, GL6-F2) | 25 | `c11d4fbf` | +| R-310 | fixed on main + runbook line | 15 | `b43705ac` | +| R-274 | fixed on main (disclosure + pinning tests; check shipped with R-297) | 15 | `f790734e` | diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index c020d13b..64d863c1 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,14 @@ --- +## 2026-10-05 (night) — the burn-down night: stale rows + +The full text of every row below: `git show c8da8e04:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-531** | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** (P3) | CLOSED 2026-10-05 — STALE: all three facts are measured, ruled and built (the burn-down night) | Deploy-kill timing measured (37 s, 2026-09-16, the row's F9'); the budget ruled and the slow 24-h counter shipped (R-539, closed proven live 2026-09-17; `03` §4 RULING 2026-09-16 + BUILT 2026-09-17); deliberate kills count by stated design (felhom-agent `internal/localapi/controllersupervisor.go:73-81`, tests `TestControllerSupervisor_SlowCrashloop` / `_SpreadRestartsNeverSlowCrashloop` / `_SlowCounterSurvivesAgentRestart`); the hub's silent first-sight seed is pinned by `TestControllerSupervisorChecker_EmitsOnMovement` (hub `internal/monitor/controller_supervisor_test.go:47-52`). | + ## 2026-10-05 (night) — the burn-down night: rulings §1 The full text of every row below: `git show 30650cad:documentation/backlog/OPEN-ITEMS.md`. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 40bb9019..799ba97e 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -116,18 +116,18 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-130** | Install & onboarding | P3 | **A "hard min" that only warns.** A fresh box's `local-lvm` was ~75 GiB against `HARD_MIN_LVM_GIB=120` (`scripts/felhom-host-install.sh`); the installer logged `[WARN] local-lvm free ~75 GiB < hard min 120 GiB` and went on to a **fully successful** install | READY (S) | — | Either the minimum is not hard (rename it and state the real floor) or it is wrong (and 120 GiB is not what a working appliance needs). Leaving it is the R-29 shape: a check that reads as coverage while providing none. Evidence: same audit §8 | CC | -| **R-179** | Install & onboarding | P3 | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** The teardown's residue-diff provenance (`day0-install.md` Part E: *"a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers"*) is from **v1.9.1**, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps `/etc/systemd/system/mnt-felhom\x2ddrives-.mount` and `.automount` after a full uninstall | **READY (S) — NEW 2026-08-03** | — | **Observed on demo-hp 2026-08-03** after `--uninstall --vmid 9201`: `mnt-felhom\x2ddrives-Felhom\x2dShare.automount` **loaded failed failed**, its `.mount` `loaded inactive dead`, and `mnt-felhom\x2ddrives.mount` still `active mounted` — the uninstall's own output had warned `/mnt/felhom-drives/Felhom-Share is busy — NOT forcing` and `/mnt/felhom-drives root bind left mounted`, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, `daemon-reload`, unmount the autofs then the parent. **NEGATIVE CONTROL, same day:** demo-felhom's uninstall left **nothing** (`ls /etc/systemd/system | grep -i felhom` → only the unrelated `felhom-bootstrap.service`; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** `felhom-bootstrap.service` is NOT residue — it is the ISO first-boot unit, `disabled`+`inactive`, exactly-once and already fired | CC | -| **R-180** | Install & onboarding | P3 | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** `felhom-host-install.sh` validates the archive storage EXISTS (`pvesm status --storage`, `:1583`) and that the golden volid RESOLVES on it (`:1661`), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default `local local-lvm felhom-pbs` (`--acl-storages`, which `runbooks/day0-install.md` tells the operator **not** to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step | **READY (S) — NEW 2026-08-03** | — | **Hit live on demo-hp 2026-08-03** during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on `felhom-backup` (the enrolled NVMe, where the box's vzdumps live) and `--archive-storage felhom-backup` passed. Pre-flight passed; steps 1–7 ran; step 8 returned `reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)`. **The cost is the ORDER, not the error** — by the time it fires, step 2 has minted the PVE token, step 4b has **rotated root@pam and vaulted it** (so the old console password is already dead), and step 5 has installed the agent. Recovery was `--resume` after moving the golden to `local`, which worked cleanly. **This is statically checkable in pre-flight**: `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a one-line assertion over two variables both known at `:1583`. Same class as R-29 — the checkable thing that nothing checks | CC | +| **R-130** | Install & onboarding | P3 | **A "hard min" that only warns.** A fresh box's `local-lvm` was ~75 GiB against `HARD_MIN_LVM_GIB=120` (`scripts/felhom-host-install.sh`); the installer logged `[WARN] local-lvm free ~75 GiB < hard min 120 GiB` and went on to a **fully successful** install | READY (S) **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `efe76093`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). Decided by CC unattended — operator may reverse (`09` §3 decision 131): renamed, not enforced; no real floor has been measured. **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Either the minimum is not hard (rename it and state the real floor) or it is wrong (and 120 GiB is not what a working appliance needs). Leaving it is the R-29 shape: a check that reads as coverage while providing none. Evidence: same audit §8 | CC | +| **R-179** | Install & onboarding | P3 | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** The teardown's residue-diff provenance (`day0-install.md` Part E: *"a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers"*) is from **v1.9.1**, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps `/etc/systemd/system/mnt-felhom\x2ddrives-.mount` and `.automount` after a full uninstall | **READY (S) — NEW 2026-08-03** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `c11d4fbf`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). Left: an enrolled DATA drive's mount unit still survives an uninstall (removing it unmounts customer drives — GL6-F2, its own question). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | **Observed on demo-hp 2026-08-03** after `--uninstall --vmid 9201`: `mnt-felhom\x2ddrives-Felhom\x2dShare.automount` **loaded failed failed**, its `.mount` `loaded inactive dead`, and `mnt-felhom\x2ddrives.mount` still `active mounted` — the uninstall's own output had warned `/mnt/felhom-drives/Felhom-Share is busy — NOT forcing` and `/mnt/felhom-drives root bind left mounted`, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, `daemon-reload`, unmount the autofs then the parent. **NEGATIVE CONTROL, same day:** demo-felhom's uninstall left **nothing** (`ls /etc/systemd/system | grep -i felhom` → only the unrelated `felhom-bootstrap.service`; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** `felhom-bootstrap.service` is NOT residue — it is the ISO first-boot unit, `disabled`+`inactive`, exactly-once and already fired | CC | +| **R-180** | Install & onboarding | P3 | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** `felhom-host-install.sh` validates the archive storage EXISTS (`pvesm status --storage`, `:1583`) and that the golden volid RESOLVES on it (`:1661`), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default `local local-lvm felhom-pbs` (`--acl-storages`, which `runbooks/day0-install.md` tells the operator **not** to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step | **READY (S) — NEW 2026-08-03** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `7f3944ef`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | **Hit live on demo-hp 2026-08-03** during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on `felhom-backup` (the enrolled NVMe, where the box's vzdumps live) and `--archive-storage felhom-backup` passed. Pre-flight passed; steps 1–7 ran; step 8 returned `reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)`. **The cost is the ORDER, not the error** — by the time it fires, step 2 has minted the PVE token, step 4b has **rotated root@pam and vaulted it** (so the old console password is already dead), and step 5 has installed the agent. Recovery was `--resume` after moving the golden to `local`, which worked cleanly. **This is statically checkable in pre-flight**: `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a one-line assertion over two variables both known at `:1583`. Same class as R-29 — the checkable thing that nothing checks | CC | | **R-250** | Install & onboarding | P3 | **A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time.** Found 2026-08-07 creating the fifth walk's venue. `POST /configs/new` with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — **fail-closed by design** (`offsite.go:111-121`, *"don't serve a descriptor the controller can't verify"*), with `defaultScanBackoff` = 2+4+8+16+30 ≈ **60 s**, sized by its own comment to *"the observed DNS propagation lag"*. **Measured, both halves:** the first create exhausted the ladder — five `no such host`, then `dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable` (the name had just begun resolving, **AAAA-first, into a pod with no IPv6 route**) — and the hub logged `[ERROR] offsite provision for walk5`. An **identical second POST succeeded** ~70 s later on its own final rung (`shared already provisioned … (subaccount 285351)`). **Total settle ≈ 100 s against a 60 s budget.** The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, `SSH-2.0-OpenSSH_9.6p1`. **Two distinct things are wrong and should not be merged:** (1) the budget is sized against DNS *existence*, but what actually bit is the **AAAA-before-A** window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) **the operator is told nothing actionable** — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — `one_time_secrets` 1, one sub-account, no double-provision) **but that safety is invisible to the person deciding whether pressing again will double-charge them.** **Severity LOW-MEDIUM:** self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | **READY** — owner Viktor | — | — | operator | | **R-283** | Install & onboarding | P3 | **After a rebuild the hub says "Claimed 18d ago" while the box serves its first-run setup page.** `customer_claims` for demo-hp still read `claimed_at 2026-07-21 16:29:25`, `generation 2`, `issued_at 2026-08-03` while the freshly provisioned guest — whose `settings.json` is new — correctly showed „A szerver beállítása". The two sides never reconcile: the hub's claim state survives a guest rebuild and the box's does not. Consequences: the operator's screen says the box is claimed when it is not, a resend produces a RESET code instead of a SETUP code (→ **R-282**), and any previously issued code fails with *„Hibás vagy lejárt kód"* — a message that is technically true and tells the customer nothing about the real cause, namely their own reinstall. Mirror image of **R-214/R-235** (an already-paired box still told to pair itself) | **READY (S) — NEW 2026-08-09** | — | Let a report from a box carrying no claim state clear the hub's, or show both sides on the operator page | CC | -| **R-306** | Install & onboarding | P3 | **`--preflight-only` says "no state written" and writes state — with an answer that can be wrong.** `_state_put` short-circuits on `DRY_RUN` only (`felhom-host-install.sh:418`), so a preflight-only run creates `/var/lib/felhom-install/state.json`. Observed live: after a run whose banner read `PRE-FLIGHT PASS (mode=byo) — no state written, no install step executed`, the file existed containing `{"completed": [], "dnsmasq_preexisting": "yes"}`. Both the banner and the flag's own comment at line 226 assert the opposite. **The harm is not the file, it is the value**: the runbook recommends preflight-only first, then the same command without the flag, so on a box carrying a Felhom leftover the wrong ownership answer is baked in before the real install begins | **READY (S) — NEW 2026-08-12, RANK 3** | R-300, R-305 | Either make `_state_put` a no-op under `PREFLIGHT_ONLY` (and record ownership at install instead), or correct both claims. A comment asserting an invariant needs a test pinning it | CC | +| **R-306** | Install & onboarding | P3 | **`--preflight-only` says "no state written" and writes state — with an answer that can be wrong.** `_state_put` short-circuits on `DRY_RUN` only (`felhom-host-install.sh:418`), so a preflight-only run creates `/var/lib/felhom-install/state.json`. Observed live: after a run whose banner read `PRE-FLIGHT PASS (mode=byo) — no state written, no install step executed`, the file existed containing `{"completed": [], "dnsmasq_preexisting": "yes"}`. Both the banner and the flag's own comment at line 226 assert the opposite. **The harm is not the file, it is the value**: the runbook recommends preflight-only first, then the same command without the flag, so on a box carrying a Felhom leftover the wrong ownership answer is baked in before the real install begins | **READY (S) — NEW 2026-08-12, RANK 3** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `17af3b65`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | R-300, R-305 | Either make `_state_put` a no-op under `PREFLIGHT_ONLY` (and record ownership at install instead), or correct both claims. A comment asserting an invariant needs a test pinning it | CC | | **R-516** | Install & onboarding | P3 | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** | — | — | CC | | **R-554** | Install & onboarding | P3 | **[P3-LOW] Delete the first-boot setup wizard — obsolete by design, still reachable.** OPERATOR DECISION 2026-09-17 (localisation starter, decision 4: „out of scope, obsolete"). `02-controller-module-map.md` L56 calls `internal/setup/` obsolete; `cmd/controller/main.go` L322 still enters it when `setup.NeedsSetup(cfg)` — `customer.id` empty after bootstrap ingestion, or a `.needs-setup` marker (`internal/setup/setup.go` L17-25). Ingestion leaves `customer.id` empty on a missing/invalid `bootstrap.json`, a failed hub pull, or a failed merge/write/reload (`internal/bootstrap/bootstrap.go` L109-163) — so a box whose first boot cannot reach the hub shows a household an 8-page wizard (95 Hungarian strings, its own template set and CSRF). **Fix shape:** decide what such a box shows instead (a single „cannot reach Felhom yet, retrying" page — needs no decision beyond copy), then delete `internal/setup/` and `runSetupMode`; red-proof that a failed ingestion renders the waiting page, not a 404. **Check first** whether any drill/golden path still relies on `.needs-setup`. | **READY - rank P3-LOW; owner: CC** | — | — | CC | -| **R-310** | Install & onboarding | P4 | **Two small edges on the installer, neither costing more than a moment.** (1) The R-297 operator-named refusal states the vouched version twice in consecutive sentences (*"…but the vouched golden is 0.213.0. The vouched golden is 0.213.0."*). (2) `--uninstall` reads its typed vmid confirmation from `/dev/tty` and `--force` deliberately does **not** bypass it, so teardown cannot be scripted without a pty — correct for an irreversible destroy, but undocumented; it surfaces as `line 891: /dev/tty: No such device or address` and an rc=1 that looks like a failure rather than a refusal to proceed unattended | **READY (S) — NEW 2026-08-12, RANK 4** | R-297 | Drop the duplicated sentence; add one runbook line naming the pty requirement | CC | +| **R-310** | Install & onboarding | P4 | **Two small edges on the installer, neither costing more than a moment.** (1) The R-297 operator-named refusal states the vouched version twice in consecutive sentences (*"…but the vouched golden is 0.213.0. The vouched golden is 0.213.0."*). (2) `--uninstall` reads its typed vmid confirmation from `/dev/tty` and `--force` deliberately does **not** bypass it, so teardown cannot be scripted without a pty — correct for an irreversible destroy, but undocumented; it surfaces as `line 891: /dev/tty: No such device or address` and an rc=1 that looks like a failure rather than a refusal to proceed unattended | **READY (S) — NEW 2026-08-12, RANK 4** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `b43705ac`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | R-297 | Drop the duplicated sentence; add one runbook line naming the pty requirement | CC | | **R-494** | Install & onboarding | P4 | **NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (`architecture/01-topology-and-trust.md`).** *Original finding, kept:* **[P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead.** MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention **I1**): the claim mail points at `https://felhom.drill0242.felhom.eu`; that name has **no A and no AAAA** record (`dig @1.1.1.1`, control `felhom.enkisfelhom.hu` resolves); the hub has **no tunnel- or DNS-creation code** (`hub/internal/cloudflare/` holds only geo-rule removal; `cf_tunnel_token` is a pasted, optional form field, `configs.go:1478`) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered `google.com` but not the dashboard name at 13:27:39Z; the agent applied the record at **13:27:44Z** (`lanresolver: applied split-horizon record … ip=192.168.0.158`, 3 m 46 s after the controller started), so the box CAN answer the name — **but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that**; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (`curl --resolve …:443:192.168.0.158`). **A volunteer could not have done that.** **What it needs:** an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: operator-side automation; the operator creates the tunnel by hand per the day-0 runbook.** | — | — | CC | | **R-504** | Install & onboarding | P4 | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** **Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic.** | — | — | operator | -| **R-881** | Install & onboarding | P4 | **The installer's uninstall does not remove `/usr/local/sbin/felhom-priv-apply`** (added to the bundle by agent v0.146.1, R-861), and its comment still says the guest hook is agent-installed at runtime (`scripts/felhom-host-install.sh` ~line 1679) — since v0.146.1 the hook is a bundle file. Found 2026-10-05 while checking the installer against the new bundle. | READY | — | Add the file to the uninstall list and fix the comment at the next installer tag | CC | +| **R-881** | Install & onboarding | P4 | **The installer's uninstall does not remove `/usr/local/sbin/felhom-priv-apply`** (added to the bundle by agent v0.146.1, R-861), and its comment still says the guest hook is agent-installed at runtime (`scripts/felhom-host-install.sh` ~line 1679) — since v0.146.1 the hook is a bundle file. Found 2026-10-05 while checking the installer against the new bundle. | READY **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `85de3f9b`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Add the file to the uninstall list and fix the comment at the next installer tag | CC | ## Apps & catalog — 21 rows (P3 11, P4 10) @@ -196,7 +196,7 @@ stopping line that lies. | **R-698** | Backup & restore | P3 | **[P3-LOW] A backup stores the image's NAME, not the image — a restore of a version its maker has deleted cannot start.** `RecoveryManifest.image_pins` ("image NOT stored — re-pulled on restore"); since controller v0.275.0 each data file also records its running `ref@digest`, and a restore brings the data back AT ITS OWN VERSION (`07` §6.6) — so a restore asks for exactly the old image. **Measured 2026-09-26** (`audits/version-travel-2026-09-26/A7/`, registry HEADs, no pulls): the catalog's 42 ladder `ref@digest` pairs all resolve (200); an invented digest answers 404 on Docker Hub and ghcr.io (negative control). Not measured: the digests recorded on boxes (older than any ladder entry), how often makers delete versions, the catalog's 66 digest-less compose lines. **Options (decide nothing yet):** (a) keep — a restore of a deleted version fails at the pull and the household uses the next copy or a newer version; (b) mirror every INSTALLED image into the DooPlex registry, restore falls back to it — storage + bandwidth on DooPlex, a new part on the recovery path; (c) mirror only ladder-named versions — bounded, misses pre-ladder boxes; (d) `docker save` into the unit — hundreds of MB per app per copy on every tier. **-- 2026-09-30 late (decision 53):** a box now keeps only an app's running and previous image; a restore to an older version re-pulls it — as every restore already did. The limit above is unchanged. | **OPEN — P3; owner: operator (a decision), CC measures** **Operator ruling 2026-10-05 18:23: kept OPEN as a known risk to a household's restore; owner the operator; not worked on in the burn-down.** | — | — | operator | | **R-729** | Backup & restore | P3 | **[P3-LOW] An off-site target, once saved on the page, cannot be removed through the product.** MEASURED 2026-09-30 on 9202: `/backup/offbox/config` refuses an empty address and no route clears the target; the session removed its throwaway target from `settings.json` by hand, with the controller stopped (harness teardown on a scratch guest). A household that tries its own NAS and gives up keeps a disabled target forever. **Fix direction:** a „Távoli mentési cél törlése" press that clears the target (never the repository). | **READY — rank P3-LOW; owner: CC (controller)** | — | — | CC | | **R-91** | Backup & restore | P4 | Old 13 GB datastore copy at `/srv/pbs-felhom` on ep0's root disk **Checked from source 2026-10-05 (burn-down round 2):** Gate is long past (row waits on demo-felhom's first post-migration PBS backup, migration 2026-07-27). Last positive record of the copy: audits/CAMPAIGN-9-restore-proof-2026-07-28.md:759 'ep0 : /srv/pbs-felhom rollback copy intact (13G)'; CONTEXT.md:3666 still says it is 13 G of dead weight awaiting R-91. No later record of deletion found (grep srv/pbs-felhom across felhom.eu). Deleting is on ep0 (protected) and needs an operator word. | WATCHING | demo-felhom's first **post-migration** PBS backup | Delete once it lands; fix `CONTEXT.md:1018` same commit | CC | -| **R-99** | Backup & restore | P4 | Server-side prune **never removes** a phantom snapshot. Confirmed it does NOT count them toward `keep-last` (dry-run kept 2 real + the phantom) so there is **no retention/data-loss bug** — but one accumulates per aborted upload, forever | READY (S) | — | Decide a cleanup path. Deletion on a **customer** datastore is a separate ruling — detection shipped, removal deliberately not automated | CC | +| **R-99** | Backup & restore | P4 | Server-side prune **never removes** a phantom snapshot. Confirmed it does NOT count them toward `keep-last` (dry-run kept 2 real + the phantom) so there is **no retention/data-loss bug** — but one accumulates per aborted upload, forever | READY (S) **2026-10-05 (burn-down night): NEEDS THE OPERATOR.** Removing a phantom means deleting on a customer datastore (the row's own separate ruling), and the prune runs on ep0 (fenced). Next: the ruling. | — | Decide a cleanup path. Deletion on a **customer** datastore is a separate ruling — detection shipped, removal deliberately not automated | CC | | **R-164** | Backup & restore | P4 | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | | **R-213** | Backup & restore | P4 | **Putting files back in place — the half the recovery screen deliberately does not do.** The screen (R-193, controller v0.200.0) unlocks the repository and LISTS what is in it; restoring is per-app and lives in the backups area, and the operator ruled the two separate on 2026-08-05: *a screen that unlocks and then offers to overwrite is two decisions wearing one button*. What is missing is the step after the listing — a customer who can now SEE their files still has to work out, per app, which restore to choose. **The operator named its requirement: a live-versus-backup comparison** — the customer must be able to see what would change before anything is overwritten | **OPEN — not started, deliberately** | the comparison design (nothing exists for it yet) | Design the live-vs-backup comparison, then the put-back flow on top of it. Do NOT fold it into the recovery screen | Operator + CC | | **R-246** | Backup & restore | P4 | **A leftover staleness flag has silently disabled the new recovery discriminator on `demo-hp` since 2026-08-04, and the flag is WRONG.** Found by a read-only spike, 2026-08-08. **Q1 — traced to an act, to the second:** at `2026-08-04 20:15:49` the hub emitted `offsite_reissued` and `escrow_stale` in the same second — an operator **Re-issue** pressed during the R-201 drill, three minutes after `escrow_blob_served` at 20:12:40/20:12:54. That was `offsite.ReissueCredentials`'s **precautionary** `MarkEscrowStale` call, which **hub v0.95.0 REMOVED the very next day** (R-196 / R-204 item 2) precisely because it marked healthy escrows stale. **Q2 — the flag is wrong, measured on both sides:** the hub's blob seals `restic_pw_sha256 = 8a9e33aa4da6769c…d080a`, and the key the box is actually using hashes to **the identical value**. The blob covers the key. **Q3 — nothing clears it by itself:** the ONLY writer of `stale_at = NULL` is `SaveHostEscrow`'s `ON CONFLICT` — i.e. a fresh escrow ceremony, **which is the one act that would supersede the good blob**. So the only exit from a false alarm is the destructive act the false alarm recommends. **Q5 — the blast radius, enumerated rather than assumed:** (1) `GetEscrowStatusForCustomer` withholds `restic_pw_sha256` from the ACK; (2) a **pending** box can never auto-confirm, so (3) **every off-site run is refused indefinitely** — neither bites `demo-hp`, which was already `escrowed` and is backing up healthily (12 snapshots, last success 2026-08-07T02:15:35Z); (4) the customer is told to create a new code; (5) **NEW — v0.206.0's shape (c) is inert**, because the box records an empty hub hash and falls back to (a)/(b), so the recovery screen would stay silent even if recovery were needed. **Q6 — a fresh box CANNOT reach this state:** `MarkEscrowStale` has **no production caller anywhere in the tree** (census: only its own definition, two comments and two test references). **The next walk cannot meet it.** **NOT CLEARED, deliberately** — Q2's answer says the fix is to clear this instance, but whether to also stop the column being settable at all is a separate ruling, and the spike was scoped read-only. **✅ THE FLAG IS CLEARED — operator-approved and applied 2026-08-08.** One row, identity-matched on `host_id` and guarded on `stale_at IS NOT NULL`; `changes()` returned **1**. **Verified end to end, not just in the database:** the hub now serves the hash again, the box recorded `hub_escrow_key_sha256 = 8a9e33aa4da6769c…d080a` at `11:10:19Z`, and that is **byte-identical to the key it is using** — so shape (c) compares, matches, and correctly stays silent. **The false stale warning is gone, proven with a positive control** rather than an absent line: **0** `escrow-confirm` lines since the restart while **5** scheduler lines in the same window prove the box was logging, and the recorded hash proves an ACK was processed. *(Method note: the hub pod is Alpine with no `sqlite3`; it was installed into the container's ephemeral writable layer — image and node untouched, gone on restart. SQLite's own file locking coordinated the write with the live hub; an earlier attempt failed cleanly on quoting and changed nothing, which is the fail-safe working.)* **STILL OPEN under this ID: the ruling on whether `stale_at` keeps a live setter at all.** It currently has NO production caller, so the column is write-only-by-accident — a field that changes behaviour, that nothing sets and nothing can see (R-248). Either give it an evidential setter or retire it; **do not leave it as a trap that only a database read can spring.** | **READY — flag cleared; the column ruling is still owed** — owner Viktor **Folded R-248 2026-10-03** (the same ruling: give `stale_at` a visible, evidence-bearing setter, or retire it). | — | — | operator | @@ -220,7 +220,7 @@ stopping line that lies. | **R-332** | Storage & devices | P3 | **The new Hiba-from-counters path has never fired on real hardware.** v0.215.0's whole point is a verdict the product could not previously reach, and it is proven only against the committed fixture's values in unit tests (12 scenario groups, 11 of 12 red-proofs failing as required). The live validation on demo-hp proved the **negative** — three healthy disks still read Rendben across the deploy, no false alert — and the **severity wire** end to end, but no live disk has actually reached Hiba. **This is the honest gap and it must not be closed by pointing at the fixture tests**: the drive that produced the fixture is in DooPlex, which is Tier 2 and never a drill target, and the demo boxes are all-flash and healthy | **WATCHING — NEW 2026-08-14, NARROWED same day.** One item originally in this gap is now PROVEN LIVE: the **persisted state surviving a controller restart**. The v0.215.0→v0.216.0 redeploy destroyed and rebuilt the container, and the new one read back a `changed_at` written by the PREVIOUS version (`2026-08-14T07:23:14.640216851Z`, still intact at 09:31:35Z) instead of re-baselining — Scenario L on real hardware, not just the production-path unit test. **What remains unproven is the verdict itself, plus the stronger restart half: an already-ALERTED disk not re-alerting** | a real degrading disk, or an injection harness | **Closing condition:** a live disk reaching Hiba from counters, OR a deliberate injection through the REAL pipeline (agent `/disks` → controller check → hub event), not a hand-set verdict | CC | | **R-542** | Storage & devices | P3 | **[P3-LOW] `/api/disks/candidates` offers a REGISTERED, in-use drive under „initialize".** MEASURED 2026-09-16 on the fresh box (controller 0.244.0): after `/dev/sdb` was formatted, mounted at `/mnt/felhom-drives/adatlemez` and registered as the default data drive, the endpoint still listed it under `initialize` (and again under `attach` with `already_mounted: null`). **Customer-invisible today:** the Meghajtók page filters correctly — its „Nem regisztrált meghajtók" section lists none, and the drive shows as „Adatlemez · Alapértelmezett · Aktív". So the defect is in the raw endpoint that feeds a FORMATTING flow, not in the page. **It also misleads a session:** I read „not mounted" off this endpoint and briefly filed a false finding against the product (corrected in `audits/evidence-backup-promise-2026-09-16/phaseE-freshbox.txt`). **Fix shape:** exclude paths the controller has registered from `initialize`, and set `already_mounted` from the real mount state rather than null. | **READY — rank P3-LOW; owner: CC (controller/agent)** | — | — | CC | | **R-756** | Storage & devices | P3 | **[P3-LOW] On 9202, "remove with drive data" refuses calibre-web with 409 „…/scratch_hdd/userdata/calibre-web tárhely jelenleg nem elérhető", while the controller container lists that folder.** MEASURED twice on 2026-10-01 (`audits/lockouts-2026-10-01/B/B1…`, `audits/calibre-name-and-prune-2026-10-01/A/A1…`): `POST /api/stacks/calibre-web/remove` with `remove_hdd_data` → 409; `docker exec felhom-controller ls -ld /mnt/felhom-drives/scratch_hdd/userdata/calibre-web` → the directory (dated 2026-09-22). The walk then removed the app keeping the data (R-442's fail-closed answer). Either the drive is not a registered drive on this scratch box (a test-venue artefact) or the resolver reads another path than the one it names. Not measured which. **-- 2026-10-01 (night, new apps):** the same 409 for Grimmory (`remove_hdd_data`), and a drive app's per-app backup is NOT offered for restore after a remove — the controller logged `drive … not mounted — skipping ensure (held by drive gate)`; the unit lives on that drive. So on 9202 checklist row 2.5 cannot be shown for any drive app until its scratch drive is a registered drive (`audits/new-apps-2026-10-01/box/grimmory/restore-why.txt`). **Checked from source 2026-10-05 (burn-down round 2):** Mechanism found in source: the refusal comes from felhom-controller@114ff27 controller/internal/stacks/delete.go:147-149 if !m.DriveLive(hddPath) -> msgDriveAbsentFmt with hddPath, and DriveLive is deploy.go:1007-1011 return m.isMountPoint(hddPath) -- it requires HDD_PATH ITSELF to be a mount point. Everywhere else HDD_PATH is compared to a registered storage path (api/router.go:574-575, web/handlers.go:3049, storage_handlers.go:528), i.e. a drive root. The message names .../scratch_hdd/use | **OPEN — rank P3-LOW; owner: CC** | — | — | CC | -| **R-25** | Storage & devices | P4 | **Device-node TOCTOU hardening (drive init).** Graduate the controller v0.141.0 Observation: the `format → resolveEnrollUUID(path) → AssignDisk(uuid)` sequence has a narrow /dev-re-enumeration window (agent-guarded on the destructive format via anti-retarget durable-id; benign fs-UUID mount). Bind resolve+assign to the format's durable-id so the mount can't target a moved node. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | From the v0.141.0 F6 commit's security-review finding (`felhom-controller` REPORT). Low real risk (single-operator, agent-guarded), but cheap to close | CC | +| **R-25** | Storage & devices | P4 | **Device-node TOCTOU hardening (drive init).** Graduate the controller v0.141.0 Observation: the `format → resolveEnrollUUID(path) → AssignDisk(uuid)` sequence has a narrow /dev-re-enumeration window (agent-guarded on the destructive format via anti-retarget durable-id; benign fs-UUID mount). Bind resolve+assign to the format's durable-id so the mount can't target a moved node. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-19, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): agent half on `main`** (felhom-agent `769c4c3`: the format answer carries the new filesystem's UUID, only while the durable id still names the formatted device; tests + 3 red-proofs). Left: the controller mounts that UUID; then a release. | — | From the v0.141.0 F6 commit's security-review finding (`felhom-controller` REPORT). Low real risk (single-operator, agent-guarded), but cheap to close | CC | | **R-331** | Storage & devices | P4 | **Disk health Phase 3 — growth-rate detection, and retiring the static 64.** The v0.215.0 count backstop (64 unreadable sectors → Hiba) is **a judgement from ONE drive**: the observed benign excursion peaked at 16 and cleared inside an hour, and the terminal run passed 64 at 13 Aug 11:28 and never came back. It is deliberately a backstop BEHIND the sustain rule, not the primary signal, but it is still a magic number tuned on a single sample and it will be wrong for some drive. With Phase 2's history the box can ask the question that actually matters — *is this count climbing, and how fast* — which distinguishes a drive with eight stable aging sectors from one adding forty a day, something no static threshold can do. Revisit 64 when that exists | **READY (M) — NEW 2026-08-14** | R-330 | Growth-rate rule over persisted samples; re-derive or delete the static 64 | CC | | **R-352** | Storage & devices | P4 | **Four screens state something untrue about where an app's data goes, and the configured default is consulted by nothing that places data.** Measured on `demo-hp` 2026-08-21. **(1)** **40 of 53** catalogue templates declare no data path (`grep -rl 'env_var: HDD_PATH' --include='.felhom.yml'` → 13; total 53); those apps get **no storage field and no default** — their data lands in a named Docker volume on the system drive. **(2)** `GetDefaultStoragePath()` has exactly **three** non-test callers — the metrics collector (`cmd/controller/main.go:410`), the dashboard SystemInfo panel (`web/server.go:733`) and `.fab` import landing (`handler_export_upload.go:154`). **The deploy route never reads it.** Its field comment `// new apps use this by default` (`internal/settings/settings.go:453`) has never been true — an invariant with no test pinning it. **(3)** The first-tier backup follows the data onto the same disk (`backup/backup.go:324-334` → `systemDataPath`), so data and nearest copy share one device for a customer doing nothing wrong — the posture Tier 2 refuses outright at `tier2.go:329`. **(4)** „1 alkalmazás használja" on the Drives page counts only `Env["HDD_PATH"] == path` (`web/handlers.go:2118`), so it can never include the 40-class; it truthfully means *„1 of the apps that CAN use a drive does"*. | **NARROWED** — **PARTLY CLOSED 2026-08-21** — visibility shipped; **placement OPEN** | — | **Shipped tonight (visibility only, no placement change, nothing migrated):** the deploy page now states where the app's data will live before the button is pressed, naming the system drive for the 40-class and the selected drive for the 13. **⚠ RE-FRAMED 2026-08-22 — THE FOUR MEASUREMENTS STAND; TWO OF THE CONCLUSIONS DRAWN FROM THEM DO NOT.** (1) is a measurement and is correct, but "those apps get no storage field and no default" is not a deprivation: the architecture places **hot** data (DB/config/cache) on fast storage inside the guest and states that placement is **ENFORCED** (`documentation/architecture/01-topology-and-trust.md:150-152`). The 40 are all-hot apps; the 13 are the ones with **bulk** content, which belongs on an attached drive. There is no choice being denied. (3) **overstated one risk and understated a distinction.** Since R-165 the guest carries a small OS rootfs plus **ONE** data volume at `/var/lib/felhom`; `/var/lib/docker` and `/mnt/sys_drive` are two **binds of that same volume** (`felhom-agent/configs/build-golden.sh:29-40, 99`) — the `mp0`/`mp1` split assumed here was retired 2026-08-03. **Real risk:** a physical-disk failure loses the data and its first-tier copy together — which is what the off-site and whole-machine tiers exist for, and which is equally true of a drive-resident app whose unit sits beside its data by design. **Overstated risk:** a full data volume stopping the operating system — the OS rootfs is a separate volume and the capture floor refuses per app before exhaustion (`00-capability-map.md:94`), watched working 2026-08-21 with the volume at 99% and all 15 containers healthy. **The comparison to Tier 2's same-disk refusal (`tier2.go:329`) is withdrawn:** Tier 2 refuses a SECOND copy on the same disk; Tier 1's unit is meant to sit beside the data. **(2) and (4) are untouched and remain correct** — (2) is now filed on its own as **R-368** with its scope measured, and (4) needs no ruling: the count is honest and only easy to misread. **The specification for the rest is filed at `documentation/backlog/SPEC-app-data-placement-2026-08-21.md`** (corrected 2026-08-22, framing marked inline, measurements kept) and lists the five points a ruling must settle (compose-template vs controller, existing deployments, when the SSD is legitimately right, `IsDefault` must become true or go away *with a test*, and the Drives-page count). **An earlier recommendation to refuse deployment until a drive is registered was WITHDRAWN** — it assumed the customer had failed to choose; they had no choice to make. | **Viktor rules**, CC executes | @@ -237,8 +237,8 @@ stopping line that lies. | **R-138** | Security & access | P3 | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) | READY (S) | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | | **R-255** | Security & access | P3 | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator | | **R-270** | Security & access | P3 | **R-268's stated rotation recipe is incomplete: the controller never re-reads `bootstrap.json`'s `local_api`, so a rotation leaves the agent channel dead across restarts.** `bootstrap.ensureLocalAPI` returns early when `cfg.LocalAPI.Endpoint != ""` — by design it FILLS an absent block and never refreshes a present one — so the token the controller uses lives in its own `controller.yaml`, not in the mount. Proved live 2026-08-09: two controller restarts after a correct `bootstrap.json` rotation, still `HTTP 401`; the channel came up only once `local_api.token` was written into `controller.yaml`. The neighbouring `DetectEndpointDrift` compares the ENDPOINT and deliberately does not compare the token (*"a token mismatch is a different failure"*), so this shape is knowingly unmodelled. Parent question — which file is authoritative — is **R-78** | **READY (S) — NEW 2026-08-09** | — | Either teach the drift detector the token, or make the rotation path write both files. Correct the R-268 row's recipe either way | CC | -| **R-275** | Security & access | P3 | **`--uninstall` leaves five 0600 `agent.json.*` credential backups, and the reinstall hands them to the new service account.** `/etc/felhom-agent/` survives with `agent.json.{campaign8-before,campaign9-before,campaign9-prev,pre-e-target-move,pre-prunegate.bak}`, each carrying a 64-char `hub.api_key` and a 59-char `proxmox.token`. The teardown claims to remove *"config (+ its .bak backups)"* and `scripts/CHANGELOG` F1 records *"uninstall now purges the agent config's `.bak*` siblings (one held a live hub api_key)"* — **that fix does not match the filenames in use, and it misses `agent.json.pre-prunegate.bak`, a file that literally ends in `.bak`.** **Exposure assessed, not assumed:** these are SUPERSEDED — the orphaned key hashes to `a5d2222a…`, the hub's current demo-hp key to `8c59d1b6…`, and the Proxmox token was deleted by the same uninstall. **But the reinstall recreates `felhom-agent` at uid 999, the same uid the deleted account had**, so three of the backups become the new account's files — verified readable as `felhom-agent`. A fresh install's service account inherits read access to the prior install's credentials; superseded today, live if the backups were recent (R-179's precedent). **Also left, undeclared:** `/etc/felhom/{.bootstrap-done,appliance-pairing-code}`, `felhom-bootstrap.service` + `/usr/local/sbin/felhom-bootstrap.sh`, the `vmbr9` stanza in `/etc/network/interfaces`, and `/etc/sudoers.d/felhom-agent.bak-pre-e2a` (21 KB — **INERT: sudo skips dotted filenames, verified with `sudo -l -U felhom-agent`; `visudo -c -f` parsing it OK is NOT evidence sudo loads it**) | **READY (S) — NEW 2026-08-09** | — | Purge by directory, not by glob; and do not let a new service account reuse a uid that owns old secrets | CC | -| **R-276** | Security & access | P3 | **RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing.** After `--uninstall` on demo-hp, `wg-quick@wg-felhom` is **enabled and active**, `/etc/wireguard/wg-felhom.conf` present, handshake to `167.233.158.164:443` **52 s old**, counters 5.86 GiB in / 2.48 GiB sent. It appears in **neither** the WIPED nor the KEPT list, though the BYO install disclosure names it prominently on the way in (*"an OUTBOUND WireGuard tunnel to the Felhom hub"*). A host told to leave Felhom retains a live network path into Felhom infrastructure, its hub-side peer registration intact, and nobody is told | **READY (S) — NEW 2026-08-09** | — | Tear the tunnel down and deregister the peer, or list it under KEPT with the reason and the removal command | CC | +| **R-275** | Security & access | P3 | **`--uninstall` leaves five 0600 `agent.json.*` credential backups, and the reinstall hands them to the new service account.** `/etc/felhom-agent/` survives with `agent.json.{campaign8-before,campaign9-before,campaign9-prev,pre-e-target-move,pre-prunegate.bak}`, each carrying a 64-char `hub.api_key` and a 59-char `proxmox.token`. The teardown claims to remove *"config (+ its .bak backups)"* and `scripts/CHANGELOG` F1 records *"uninstall now purges the agent config's `.bak*` siblings (one held a live hub api_key)"* — **that fix does not match the filenames in use, and it misses `agent.json.pre-prunegate.bak`, a file that literally ends in `.bak`.** **Exposure assessed, not assumed:** these are SUPERSEDED — the orphaned key hashes to `a5d2222a…`, the hub's current demo-hp key to `8c59d1b6…`, and the Proxmox token was deleted by the same uninstall. **But the reinstall recreates `felhom-agent` at uid 999, the same uid the deleted account had**, so three of the backups become the new account's files — verified readable as `felhom-agent`. A fresh install's service account inherits read access to the prior install's credentials; superseded today, live if the backups were recent (R-179's precedent). **Also left, undeclared:** `/etc/felhom/{.bootstrap-done,appliance-pairing-code}`, `felhom-bootstrap.service` + `/usr/local/sbin/felhom-bootstrap.sh`, the `vmbr9` stanza in `/etc/network/interfaces`, and `/etc/sudoers.d/felhom-agent.bak-pre-e2a` (21 KB — **INERT: sudo skips dotted filenames, verified with `sudo -l -U felhom-agent`; `visudo -c -f` parsing it OK is NOT evidence sudo loads it**) | **READY (S) — NEW 2026-08-09** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `85de3f9b`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Purge by directory, not by glob; and do not let a new service account reuse a uid that owns old secrets | CC | +| **R-276** | Security & access | P3 | **RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing.** After `--uninstall` on demo-hp, `wg-quick@wg-felhom` is **enabled and active**, `/etc/wireguard/wg-felhom.conf` present, handshake to `167.233.158.164:443` **52 s old**, counters 5.86 GiB in / 2.48 GiB sent. It appears in **neither** the WIPED nor the KEPT list, though the BYO install disclosure names it prominently on the way in (*"an OUTBOUND WireGuard tunnel to the Felhom hub"*). A host told to leave Felhom retains a live network path into Felhom infrastructure, its hub-side peer registration intact, and nobody is told | **READY (S) — NEW 2026-08-09** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `85de3f9b`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). Left besides the tag: the hub deletes a removed host's WireGuard peer and pushes the peer list to the off-site endpoint (a hub change). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Tear the tunnel down and deregister the peer, or list it under KEPT with the reason and the removal command | CC | | **R-338** | Security & access | P3 | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it **Checked from source 2026-10-05 (burn-down round 2):** nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246). | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-616** | Security & access | P3 | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** | — | — | CC + operator | | **R-717** | Security & access | P3 | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). | **OPEN — P3; owner: CC** | — | — | CC | @@ -262,11 +262,10 @@ stopping line that lies. | **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC | | **R-242** | Box system & updates | P3 | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. **⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit message, in `hub/CHANGELOG.md` and in `REPORT.md` — a BYPASS, not a waiver**, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and these need one. **The operator was asked and ruled bypass-now-bake-later on 2026-08-30**, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. **That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it.** **OWED: bake a golden carrying 0.225.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change — `golden_version` + `agent_version` + `min_agent`; MinAgent is 0.129.0 per both CHANGELOG headers). **Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.** **⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0.** The `felhom.eu` push carrying that release's documentation used `git push --no-verify` on the operator's standing ruling from earlier the same day, declared in the commit and in `REPORT.md`. **The day-0 ground still holds for all three and was re-checked rather than assumed:** R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; **R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore.** **The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence.** **Owed: ONE bake carrying 0.226.0 covers all three** (`RUNBOOK-manual-build.md` §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. **✅ PAID THE SAME DAY — golden `0.226.1` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30).** Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`. `golden_currency_gate.py` went **red → green** on the same command, which is its proof that it measures something real. **The three declared bypasses above are now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back — **657 197 592 B, sha256 `70ed8e93…baefe69`**, both identical to what the bake reported — and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.226.1`, i.e. the delivered artifact naming the controller it will start. **The three-field vouch was checked deliberately, not assumed:** `MinAgent 0.129.0` read from the golden's controller CHANGELOG header, `agent_version 0.130.0 ≥ min_agent 0.129.0` (so NOT the R-216 shape), and the result verified by **re-reading the manifest** rather than trusting the flash — golden option `0.226.1 SELECTED`, all four shas matching. **The floor is proven ACTING, not merely set:** `demo-felhom` self-updated within 30 s, logging `[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1)`. **⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN.** v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that `felhom.eu` push used `--no-verify` and declared it, and golden **0.227.1** was baked, published, round-trip verified, **VOUCHED** and the floor **RAISED to 0.227.1** within the hour. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`. **THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day.** Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. **The floor was proven ACTING both times:** `demo-felhom` self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new `offsite-integrity` job **by itself, on a box nobody deployed to**, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. **This row's OTHER half is still open and untouched: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. **2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day.** v0.230.0 shipped in the morning with the newest golden at 0.229.0, **which is the build R-403 says deletes a good copy**, so the gate was red across four commits (`dddcc80`, `6e550ae`, `130f7a6`, `32a4c35`). Golden **0.230.0** baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; `demo-felhom` moved itself off the defective build unattended (`controller-swap: new controller healthy`, 16:21:40 CEST). Evidence: `documentation/tests/golden-0.230.0-2026-08-31/`. **The gate did its job and its own weakness surfaced doing it — R-410.** | **READY — the vouch half only** — owner Viktor **⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - **a BYPASS, not a waiver**, on the same reasoning as the five entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **The day-0 ground was RE-CHECKED rather than reused:** R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and `MinAgent` is unchanged at 0.129.0. **The ground still expires the moment a release changes first-boot behaviour.** **OWED: bake a golden carrying 0.229.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change - `golden_version` + `agent_version` + `min_agent`, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. **Golden and fleet delivery are the operator's (this row).** **✅ PAID THE SAME DAY — golden `0.229.0` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31).** Evidence: `documentation/tests/golden-0.229.0-2026-08-31/`. `golden_currency_gate.py` went **red to green** on the same command, which is its proof that it measures something real. **The `--no-verify` bypass declared above is now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back - **656 864 331 B, sha256 `39aa886d…d7bdae87`**, both identical to what the bake reported - and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.229.0`. **A THIRD independent reader agreed before anything was vouched:** the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. **The three-field vouch was checked deliberately, not assumed** (`MinAgent 0.129.0` read from the golden's controller CHANGELOG header; `agent_version 0.130.0` >= `min_agent 0.129.0`, so NOT the R-216 shape; `agent_sha256` and `wrapper_sha256` carried through explicitly because the handler clears a field it is not sent), and verified by **re-reading the manifest** rather than trusting the flash. The **R-120 gate passed rather than being bypassed** - fleet newest 0.229.0, golden 0.229.0. **The floor is proven ACTING:** `demo-felhom` self-updated `0.228.0 -> 0.229.0` and logged `settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0)` - **nobody deployed to that box.** **Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour.** This row's OTHER half is still open and untouched: **nothing gates the VOUCH itself** - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. **⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused.** Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). **The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here.** The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - a BYPASS, not a waiver. **OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor.** `demo-hp` was updated by hand; `demo-felhom` is still on 0.229.0 and still carries the defect. **See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported.** **2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN.** R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. **Nothing gates the VOUCH.** A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's `hub_settings` table, there is no copy in git, and a hub-reading gate could not be `--fast` so it would run in neither the hook nor CI. **Do not read R-404's closure as closing this.** **2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468).** The docstring's *"honest fix is a recorded waiver in the register, never a habit of bypassing"* is now a mechanism: `golden_currency_gate.py` reads `documentation/tests/golden-waiver.yml` (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. **What stays open on THIS row is exactly one thing: nothing gates the VOUCH.** The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). | — | — | operator | -| **R-274** | Box system & updates | P3 | **A local golden is adopted with NO version and NO checksum check, so a reinstall can silently come up releases behind.** `felhom-host-install.sh` step 7: `if [[ -n "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then log_skip "using local golden"; return 0; fi` — **the hub manifest's `golden.sha256`, whose entire purpose is to vouch from a different trust root than Gitea, is consulted only on the FETCH path.** A locally-present archive bypasses the vouch: no version compare, no digest, no warning. On demo-hp 2026-08-09 the preflight selected `local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst`, whose baked marker reads `felhom-controller:`**`0.192.0`**, against a vouched golden of **0.210.0** — 18 releases stale. **The sharp consequence:** 0.192.0 is **below 0.200.0, where R-193's off-site recovery SCREEN shipped**, so a customer reinstalled today returns on a controller that cannot run the recovery ceremony their data depends on; it is also born below the managed-update floor (0.200.0), and the updater's auto-target is the floor, never the newest. **It compounds with the teardown**, which deliberately keeps the old golden (*"golden vzdump left in place"*). This is the R-111/R-115/R-120 drift family one layer down: the R-120 gate guards what may be VOUCHED, nothing guards what an install TAKES. **OBSERVED 2026-08-09, AND THE RESULT NARROWS THE ROW — recorded because it partly refutes what was written above.** On the RESUME path step 7 **fetched the vouched 0.210.0 correctly** (`fetching golden v0.210.0 from Gitea`), because `--resume` skips preflight and preflight is where local auto-discovery sets `GOLDEN_VOLID` (the GL6-F4 comment says so). **So the fresh-install and resume paths disagree on golden selection, and the resume path is the safe one.** Discovery is `… | sort | tail -1`, i.e. the NEWEST local archive by filename — a sensible heuristic, **and still no comparison against the manifest's version or sha**. The defect therefore stands as: *a box whose newest local golden predates the vouched one installs stale, silently* — which is exactly the state demo-hp was in before this run (newest local 0.192.0 vs vouched 0.210.0). It is now masked on this box because the freshly fetched 0.210.0 is the newest — **correct by recency, not by verification**. There are now **three** goldens on `local` (07-21, 08-03, 08-09), because the teardown keeps them. **Still not observed: a FRESH (non-resume) install taking a stale local golden.** | **VERIFY** (2026-10-03 triage: Local golden is now checked against the hub manifest's version and sha before use (GOLDEN_CHECK_WHY, R-297) — felhom.eu/scripts/felhom-host-install.sh:2855-2905) — **READY (S) — NEW 2026-08-09, NARROWED same day** | — | Compare the local golden's version/sha against the manifest and refuse or re-fetch on mismatch; say so in the BYO disclosure, which today lists only what the install CREATES, never what it REUSES | CC | -| **R-349** | Box system & updates | P3 | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC | -| **R-444** | Box system & updates | P3 | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** | — | — | CC | +| **R-274** | Box system & updates | P3 | **A local golden is adopted with NO version and NO checksum check, so a reinstall can silently come up releases behind.** `felhom-host-install.sh` step 7: `if [[ -n "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then log_skip "using local golden"; return 0; fi` — **the hub manifest's `golden.sha256`, whose entire purpose is to vouch from a different trust root than Gitea, is consulted only on the FETCH path.** A locally-present archive bypasses the vouch: no version compare, no digest, no warning. On demo-hp 2026-08-09 the preflight selected `local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst`, whose baked marker reads `felhom-controller:`**`0.192.0`**, against a vouched golden of **0.210.0** — 18 releases stale. **The sharp consequence:** 0.192.0 is **below 0.200.0, where R-193's off-site recovery SCREEN shipped**, so a customer reinstalled today returns on a controller that cannot run the recovery ceremony their data depends on; it is also born below the managed-update floor (0.200.0), and the updater's auto-target is the floor, never the newest. **It compounds with the teardown**, which deliberately keeps the old golden (*"golden vzdump left in place"*). This is the R-111/R-115/R-120 drift family one layer down: the R-120 gate guards what may be VOUCHED, nothing guards what an install TAKES. **OBSERVED 2026-08-09, AND THE RESULT NARROWS THE ROW — recorded because it partly refutes what was written above.** On the RESUME path step 7 **fetched the vouched 0.210.0 correctly** (`fetching golden v0.210.0 from Gitea`), because `--resume` skips preflight and preflight is where local auto-discovery sets `GOLDEN_VOLID` (the GL6-F4 comment says so). **So the fresh-install and resume paths disagree on golden selection, and the resume path is the safe one.** Discovery is `… | sort | tail -1`, i.e. the NEWEST local archive by filename — a sensible heuristic, **and still no comparison against the manifest's version or sha**. The defect therefore stands as: *a box whose newest local golden predates the vouched one installs stale, silently* — which is exactly the state demo-hp was in before this run (newest local 0.192.0 vs vouched 0.210.0). It is now masked on this box because the freshly fetched 0.210.0 is the newest — **correct by recency, not by verification**. There are now **three** goldens on `local` (07-21, 08-03, 08-09), because the teardown keeps them. **Still not observed: a FRESH (non-resume) install taking a stale local golden.** | **VERIFY** (2026-10-03 triage: Local golden is now checked against the hub manifest's version and sha before use (GOLDEN_CHECK_WHY, R-297) — felhom.eu/scripts/felhom-host-install.sh:2855-2905) — **READY (S) — NEW 2026-08-09, NARROWED same day** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `f790734e`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). The version/checksum check itself shipped with R-297; tonight added the disclosure line and the tests that pin the check. **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Compare the local golden's version/sha against the manifest and refuse or re-fetch on mismatch; say so in the BYO disclosure, which today lists only what the install CREATES, never what it REUSES | CC | +| **R-349** | Box system & updates | P3 | **"Prove it by hand, then publish" leaves the fleet running a DIFFERENT binary under the SAME version name — and self-update cannot notice.** Hit on 2026-08-20 during the R-344 train, caught and corrected the same hour, filed because the next prove-then-publish train will hit it identically. **The mechanism:** a proof deploy is a hand build (`go build -ldflags "-X main.version=0.130.0"`), while `scripts/release-agent.sh` deliberately builds with **`-trimpath -buildvcs=false`** so the published artifact is reproducible (R-186). Same source, same version string, **different bytes**: `256e0829...` on the boxes vs **`a56a92a7...`** published and vouched. **Nothing corrects it automatically**, and that is the sharp edge: the boxes already report `0.130.0`, so the self-update path sees the vouched version as already installed and does nothing, **forever**. The divergence is invisible to every version check in the system — the hub, `--version`, and the artifact manifest all agree, because they all compare the version STRING. **Consequence if unnoticed:** the binary a customer box runs is not the binary the operator vouched, and not the one a reinstall would fetch — so a bug reproduced on the fleet may not exist in the published artifact, or vice versa. It is the same "one version name, two binaries" hazard `publish-agent.sh` already carries a comment about for `CGO_ENABLED`; that comment fixed the two ENTRY POINTS and does not cover a hand build during a proof. **Corrected here** by downloading the published artifact from the registry (not rebuilding it locally — the boxes get the bytes a fresh install would get) and installing it on both; both now report `sha256 a56a92a7...`, matching the vouch. | **READY (S) — NEW 2026-08-20** **2026-10-05 (burn-down night): agent half on `main`** (felhom-agent `64f704d`: the host report carries `agent_sha256` of the running binary; test + red-proof). Left: the hub compares it with the vouched sha and shows drift; then a release. | — | Make the reconciliation a step, not a memory: the honest fix is for the agent to REPORT the sha256 of its own binary in the host report, so the hub can compare it against the vouched `agent_sha256` and flag drift — **exactly the mechanism `wrapper_sha256` already implements for the PBS wrapper** (R-50b), whose manifest help text says it *"makes host drift visible: agents report the installed file's hash and a mismatch is surfaced on the host page"*. The pattern exists and is proven; it simply was never extended to the agent's own binary. Cheaper interim: end every prove-then-publish train by installing the DOWNLOADED artifact. | CC | +| **R-444** | Box system & updates | P3 | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** **2026-10-05 (burn-down night): NEEDS THE OPERATOR.** A scheduled trim needs a new root grant (`pct fstrim `, widening what R-861 narrowed) and a new scheduled action on every customer guest's disks. Next: an operator yes on the grant and a cadence, then one measured run on demo-hp. | — | — | CC | | **R-468** | Box system & updates | P3 | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | — | — | CC | -| **R-531** | Box system & updates | P3 | **[P3-LOW] Three supervisor facts measured live and not pinned: restart timing during a deploy was not measured; restarts before the hub first sees the stanza produce no `controller_restarted_by_agent`; deliberate operator kills spend the crash-loop budget.** MEASURED 2026-09-15 on 9201: after 3 test restarts in 13 minutes the 4th kill tripped the 30-minute pause and the dashboard stayed down (the guard as designed, `A4-kill-middeploy-9201.txt`). The hub checker seeds silently on first sight, so the three restarts before the first v0.131.0 report emitted nothing (only the crash-loop did). **What it needs:** a deploy-kill timing on a fresh budget; the operator's view whether a restart after minutes of uptime should count toward the budget **MEASURED 2026-09-16 on the drill box, both halves.** (1) **Timing during a deploy (F9'):** the controller was killed 5 s into a deploy on an EMPTY budget; the agent saw it on the next sweep, confirmed on the one after, and the dashboard answered 200 again **37 s** after the kill; the interrupted app ended `not_deployed`, not stuck. (2) **The budget's shape (F9''):** three further kills at idle, 20 minutes apart, recovered in **61 s / 41 s / 61 s** - and NONE of them accumulated, because the window is 15 minutes. Four restarts this session, zero pauses, zero crash-loop events. **So the brake catches a FAST loop and is blind to a SLOW one:** a controller dying every 20 minutes is restarted forever and the only trace is an `info` event that mails nobody. That is a design question for the operator (leave it / add a longer second counter / raise the severity of the Nth restart in a day), and this session deliberately measured it without changing it. Evidence: `audits/evidence-drill-0243-2026-09-16/phase2-f9prime.txt` and `phase2-f9dprime.txt`. | **READY — rank P3-LOW; owner: CC (measure) · operator (budget rule)** | — | — | CC + operator | | **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC | | **R-839** | Box system & updates | P3 | **After a Docker restart that stopped every container, the boot sweep HELD an app whose `app.yaml` `HDD_PATH` names its user folder instead of the drive.** MEASURED 2026-10-04 on scratch 9202 (paperless-ngx): `bootrecon` logged `drive /mnt/felhom-drives/scratch_hdd/userdata/paperless-ngx is not a live mountpoint — NOT starting it`, although the drive `/mnt/felhom-drives/scratch_hdd` IS a mountpoint; the app stayed down until started by hand. The customer boxes' 9201s showed no hold. **Not diagnosed:** which writer put a per-app path in `HDD_PATH` on 9202, and whether a household box can get it. `audits/os-updates-spike-2026-10-04/partG/SUMMARY.md` | **READY — diagnose; owner: CC** | — | — | CC | diff --git a/documentation/runbooks/day0-install.md b/documentation/runbooks/day0-install.md index a51405fc..eea356fd 100644 --- a/documentation/runbooks/day0-install.md +++ b/documentation/runbooks/day0-install.md @@ -481,6 +481,9 @@ smoke before claim, drive the in-guest Host-header probe instead — the public ./felhom-host-install.sh --uninstall --vmid ``` +**`--uninstall` needs a terminal** for the typed vmid confirmation, and `--force` does not skip it (an irreversible +destroy is never unattended). Over SSH use `ssh -t`. Without a terminal it refuses in words and destroys nothing (R-310). + Guardrails (all verified in the drill): - **Ours-check**: refuses to destroy a guest without the `/etc/felhom-bootstrap` mount (i.e. not diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index 64e53e38..051f94da 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,3 +1,28 @@ +## felhom-host-install.sh 1.32.0 — the uninstall leaves nothing that holds a secret or a tunnel; pre-flight tells the truth (burn-down night, 2026-10-05; NOT published — ships with the tag `installer-v1.32.0`) + +- **R-275:** the uninstall removes the agent's config directory with every copy of the config in it (the `.bak*` glob + missed `agent.json.campaign9-before` & co.), removes the sudoers' dotted copies, and names the vmbr9 stanza and the ISO + first-boot files under KEPT; a new install whose service user is freshly created makes any old copies root-only. +- **R-276:** the uninstall takes down the WireGuard tunnel to the off-site endpoint (`wg-quick@wg-felhom` disabled + `--now`, conf removed, observed down) and names the hub-side peer under KEPT. +- **R-881:** the uninstall removes `/usr/local/sbin/felhom-priv-apply`; the BYO disclosure says it and the guest hook + come from the agent's config bundle; a test fails when the bundle gains a file the uninstall does not name. +- **R-306:** `--preflight-only` writes no state at all; the banner's „no state written" is now true. +- **R-130:** the local-lvm check is named as what it is — a RECOMMENDED 120 GiB that warns and lets the install continue + (was called a hard minimum; behaviour unchanged, as `runbooks/day0-install.md` documents). +- **R-180:** pre-flight refuses an `--archive-storage` the agent's token will not be granted on (it used to fail at step + 8/8 with HTTP 403 after root@pam was rotated); the backup target counts as granted (step 6, R-185). +- **R-179:** the uninstall removes the NAS network-storage `.automount`/`.mount` units the agent wrote, automounts first, + before the drive umounts; a busy share is never forced and is named under KEPT; drive and foreign units untouched. +- **R-310:** the named-golden refusal states the vouched version once; `--uninstall` without a terminal refuses with a + sentence before the typed confirmation instead of dying on `/dev/tty`. +- **R-274:** the BYO disclosure says a golden already on the archive storage is reused only when it matches the hub's + vouched golden (the check itself shipped with R-297); tests pin the check before the adopt. +- **New:** `scripts/test_hostinstall.py` — 33 behaviour tests that lift functions word for word from the script and run + them with recording PATH stubs (BusyBox-safe; picked up by the `script-tests` gate). Each behaviour fix was red-proved. + `hostinstall-mode-harness.sh`: GL4-D no longer reports present items missing when `grep` stops early; GL8-F1 checks the + new purge. + ## gates + tools — burn-down round 2 (2026-10-05) - **R-819:** `check_stands.py` rule 3 also accepts an id found in `CLOSED-ITEMS.md`; registered as the `stands` gate in