From e6cfe6d6aea966b0a3e33466bbc1659a93271015 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Tue, 6 Oct 2026 08:14:36 +0200 Subject: [PATCH] installer 1.32.0 published and read back (sha = tag); its nine rows closed (164 -> 155); golden 0.299.0 baked (pinned) for controller v0.299.0 (R-889) Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_0159rPz1ZhFKsS53msqPYxtS --- .../installer-publish.txt | 5 + documentation/backlog/CLOSED-ITEMS.md | 16 + documentation/backlog/OPEN-ITEMS.md | 9 - .../02-round-trip.txt | 3 + .../tests/golden-0.299.0-2026-10-06/README.md | 29 ++ .../tests/golden-0.299.0-2026-10-06/bake.log | 340 ++++++++++++++++++ scripts/CHANGELOG.md | 2 +- 7 files changed, 394 insertions(+), 10 deletions(-) create mode 100644 documentation/audits/morning-after-2026-10-06/installer-publish.txt create mode 100644 documentation/tests/golden-0.299.0-2026-10-06/02-round-trip.txt create mode 100644 documentation/tests/golden-0.299.0-2026-10-06/README.md create mode 100644 documentation/tests/golden-0.299.0-2026-10-06/bake.log diff --git a/documentation/audits/morning-after-2026-10-06/installer-publish.txt b/documentation/audits/morning-after-2026-10-06/installer-publish.txt new file mode 100644 index 00000000..cf7f8d4b --- /dev/null +++ b/documentation/audits/morning-after-2026-10-06/installer-publish.txt @@ -0,0 +1,5 @@ +== public installer read-back 2026-10-06T05:53:04Z +version: SCRIPT_VERSION="1.32.0" +public sha256 70a4ea9f750ff70d2ebd8ee2e3c3d053b2f53f78af463602c71b7b73d3ac98b9 +tag sha256 70a4ea9f750ff70d2ebd8ee2e3c3d053b2f53f78af463602c71b7b73d3ac98b9 (installer-v1.32.0 = 32a1520833) +manifest pins: 2 of 2 diff --git a/documentation/backlog/CLOSED-ITEMS.md b/documentation/backlog/CLOSED-ITEMS.md index 06f7e6cc..65220127 100644 --- a/documentation/backlog/CLOSED-ITEMS.md +++ b/documentation/backlog/CLOSED-ITEMS.md @@ -26,6 +26,22 @@ --- +## 2026-10-06 (morning) — installer 1.32.0 published + +The full text of every row below: `git show 32a15208:documentation/backlog/OPEN-ITEMS.md`. + +| Row | What | Closed | Evidence | +|---|---|---|---| +| **R-130** | **A "hard min" that only warns.** (P3) | CLOSED 2026-10-06 — FIXED AND PUBLISHED (installer 1.32.0): the 120 GiB check is named a recommendation | felhom.eu `efe76093` (`09` decision 131). Published 2026-10-06 05:52Z as `installer-v1.32.0` (= `32a1520833`; both `webpage.yaml` pins, `dbede797`); the public URL serves `SCRIPT_VERSION="1.32.0"` with sha256 `70a4ea9f…` = the tag's file (`audits/morning-after-2026-10-06/installer-publish.txt`). Operator ruling `09` §3 decision 137. | +| **R-179** | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** (P3) | CLOSED 2026-10-06 — FIXED AND PUBLISHED (installer 1.32.0): the uninstall removes the NAS network-storage units | felhom.eu `c11d4fbf`. Not done, stated: an enrolled DATA drive's mount unit still survives an uninstall (removing it unmounts customer drives — GL6-F2, its own question). Published 2026-10-06 05:52Z as `installer-v1.32.0` (= `32a1520833`; both `webpage.yaml` pins, `dbede797`); the public URL serves `SCRIPT_VERSION="1.32.0"` with sha256 `70a4ea9f…` = the tag's file (`audits/morning-after-2026-10-06/installer-publish.txt`). Operator ruling `09` §3 decision 137. | +| **R-180** | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** (P3) | CLOSED 2026-10-06 — FIXED AND PUBLISHED (installer 1.32.0): pre-flight refuses an archive storage the token will not be granted on | felhom.eu `7f3944ef`. Published 2026-10-06 05:52Z as `installer-v1.32.0` (= `32a1520833`; both `webpage.yaml` pins, `dbede797`); the public URL serves `SCRIPT_VERSION="1.32.0"` with sha256 `70a4ea9f…` = the tag's file (`audits/morning-after-2026-10-06/installer-publish.txt`). Operator ruling `09` §3 decision 137. | +| **R-306** | **`--preflight-only` says "no state written" and writes state — with an answer that can be wrong.** (P3) | CLOSED 2026-10-06 — FIXED AND PUBLISHED (installer 1.32.0): --preflight-only writes no state | felhom.eu `17af3b65`. Published 2026-10-06 05:52Z as `installer-v1.32.0` (= `32a1520833`; both `webpage.yaml` pins, `dbede797`); the public URL serves `SCRIPT_VERSION="1.32.0"` with sha256 `70a4ea9f…` = the tag's file (`audits/morning-after-2026-10-06/installer-publish.txt`). Operator ruling `09` §3 decision 137. | +| **R-310** | **Two small edges on the installer, neither costing more than a moment.** (P4) | CLOSED 2026-10-06 — FIXED AND PUBLISHED (installer 1.32.0): one vouched version per refusal; an unattended uninstall refuses in words | felhom.eu `b43705ac` + the `day0-install.md` pty line. Published 2026-10-06 05:52Z as `installer-v1.32.0` (= `32a1520833`; both `webpage.yaml` pins, `dbede797`); the public URL serves `SCRIPT_VERSION="1.32.0"` with sha256 `70a4ea9f…` = the tag's file (`audits/morning-after-2026-10-06/installer-publish.txt`). Operator ruling `09` §3 decision 137. | +| **R-881** | **The installer's uninstall does not remove `/usr/local/sbin/felhom-priv-apply`** (P4) | CLOSED 2026-10-06 — FIXED AND PUBLISHED (installer 1.32.0): the uninstall removes felhom-priv-apply | felhom.eu `85de3f9b`; a test fails when the agent's bundle gains a file the uninstall does not name. Published 2026-10-06 05:52Z as `installer-v1.32.0` (= `32a1520833`; both `webpage.yaml` pins, `dbede797`); the public URL serves `SCRIPT_VERSION="1.32.0"` with sha256 `70a4ea9f…` = the tag's file (`audits/morning-after-2026-10-06/installer-publish.txt`). Operator ruling `09` §3 decision 137. | +| **R-275** | **`--uninstall` leaves five 0600 `agent.json.*` credential backups, and the reinstall hands them to the new service account.** (P3) | CLOSED 2026-10-06 — FIXED AND PUBLISHED (installer 1.32.0): the uninstall removes every copy of the agent config | felhom.eu `85de3f9b` (tests in `scripts/test_hostinstall.py`, red-proved). Published 2026-10-06 05:52Z as `installer-v1.32.0` (= `32a1520833`; both `webpage.yaml` pins, `dbede797`); the public URL serves `SCRIPT_VERSION="1.32.0"` with sha256 `70a4ea9f…` = the tag's file (`audits/morning-after-2026-10-06/installer-publish.txt`). Operator ruling `09` §3 decision 137. | +| **R-276** | **RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing.** (P3) | CLOSED 2026-10-06 — FIXED AND PUBLISHED (installer 1.32.0): the uninstall takes the tunnel down; the hub drops the peer | installer `85de3f9b` + hub `0a460bcd` (v0.138.0, deployed 2026-10-05). Published 2026-10-06 05:52Z as `installer-v1.32.0` (= `32a1520833`; both `webpage.yaml` pins, `dbede797`); the public URL serves `SCRIPT_VERSION="1.32.0"` with sha256 `70a4ea9f…` = the tag's file (`audits/morning-after-2026-10-06/installer-publish.txt`). Operator ruling `09` §3 decision 137. | +| **R-274** | **A local golden is adopted with NO version and NO checksum check, so a reinstall can silently come up releases behind.** (P3) | CLOSED 2026-10-06 — FIXED AND PUBLISHED (installer 1.32.0): the BYO disclosure names the golden it reuses; the check is pinned | felhom.eu `f790734e` (the check itself shipped with R-297). Published 2026-10-06 05:52Z as `installer-v1.32.0` (= `32a1520833`; both `webpage.yaml` pins, `dbede797`); the public URL serves `SCRIPT_VERSION="1.32.0"` with sha256 `70a4ea9f…` = the tag's file (`audits/morning-after-2026-10-06/installer-publish.txt`). Operator ruling `09` §3 decision 137. | + ## 2026-10-06 (night) — the burn-down night: the decoy exemptions emptied The full text of every row below: `git show e6fe7319:documentation/backlog/OPEN-ITEMS.md`. diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 7fbc9446..8c8f2eab 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -116,17 +116,11 @@ stopping line that lies. | ID | Category | Sev | What | State | Blocked on | Next action | Owner | |---|---|---|---|---|---|---|---| -| **R-130** | Install & onboarding | P3 | **A "hard min" that only warns.** A fresh box's `local-lvm` was ~75 GiB against `HARD_MIN_LVM_GIB=120` (`scripts/felhom-host-install.sh`); the installer logged `[WARN] local-lvm free ~75 GiB < hard min 120 GiB` and went on to a **fully successful** install | READY (S) **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `efe76093`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). Decided by CC unattended — operator may reverse (`09` §3 decision 131): renamed, not enforced; no real floor has been measured. **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Either the minimum is not hard (rename it and state the real floor) or it is wrong (and 120 GiB is not what a working appliance needs). Leaving it is the R-29 shape: a check that reads as coverage while providing none. Evidence: same audit §8 | CC | -| **R-179** | Install & onboarding | P3 | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** The teardown's residue-diff provenance (`day0-install.md` Part E: *"a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers"*) is from **v1.9.1**, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps `/etc/systemd/system/mnt-felhom\x2ddrives-.mount` and `.automount` after a full uninstall | **READY (S) — NEW 2026-08-03** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `c11d4fbf`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). Left: an enrolled DATA drive's mount unit still survives an uninstall (removing it unmounts customer drives — GL6-F2, its own question). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | **Observed on demo-hp 2026-08-03** after `--uninstall --vmid 9201`: `mnt-felhom\x2ddrives-Felhom\x2dShare.automount` **loaded failed failed**, its `.mount` `loaded inactive dead`, and `mnt-felhom\x2ddrives.mount` still `active mounted` — the uninstall's own output had warned `/mnt/felhom-drives/Felhom-Share is busy — NOT forcing` and `/mnt/felhom-drives root bind left mounted`, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, `daemon-reload`, unmount the autofs then the parent. **NEGATIVE CONTROL, same day:** demo-felhom's uninstall left **nothing** (`ls /etc/systemd/system | grep -i felhom` → only the unrelated `felhom-bootstrap.service`; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** `felhom-bootstrap.service` is NOT residue — it is the ISO first-boot unit, `disabled`+`inactive`, exactly-once and already fired | CC | -| **R-180** | Install & onboarding | P3 | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** `felhom-host-install.sh` validates the archive storage EXISTS (`pvesm status --storage`, `:1583`) and that the golden volid RESOLVES on it (`:1661`), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default `local local-lvm felhom-pbs` (`--acl-storages`, which `runbooks/day0-install.md` tells the operator **not** to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step | **READY (S) — NEW 2026-08-03** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `7f3944ef`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | **Hit live on demo-hp 2026-08-03** during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on `felhom-backup` (the enrolled NVMe, where the box's vzdumps live) and `--archive-storage felhom-backup` passed. Pre-flight passed; steps 1–7 ran; step 8 returned `reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)`. **The cost is the ORDER, not the error** — by the time it fires, step 2 has minted the PVE token, step 4b has **rotated root@pam and vaulted it** (so the old console password is already dead), and step 5 has installed the agent. Recovery was `--resume` after moving the golden to `local`, which worked cleanly. **This is statically checkable in pre-flight**: `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a one-line assertion over two variables both known at `:1583`. Same class as R-29 — the checkable thing that nothing checks | CC | | **R-250** | Install & onboarding | P3 | **A customer create can fail fail-closed because the host-key scan ladder is shorter than the DNS/AAAA settle time.** Found 2026-08-07 creating the fifth walk's venue. `POST /configs/new` with off-site enabled provisions the Storage Box sub-account and then scans its SSH host key to pin it — **fail-closed by design** (`offsite.go:111-121`, *"don't serve a descriptor the controller can't verify"*), with `defaultScanBackoff` = 2+4+8+16+30 ≈ **60 s**, sized by its own comment to *"the observed DNS propagation lag"*. **Measured, both halves:** the first create exhausted the ladder — five `no such host`, then `dial tcp [2a01:4f8:bacc:2:200::d30]:23: connect: network is unreachable` (the name had just begun resolving, **AAAA-first, into a pod with no IPv6 route**) — and the hub logged `[ERROR] offsite provision for walk5`. An **identical second POST succeeded** ~70 s later on its own final rung (`shared already provisioned … (subaccount 285351)`). **Total settle ≈ 100 s against a 60 s budget.** The endpoint was never the problem: verified afterwards from both the node and the hub pod, by name, `SSH-2.0-OpenSSH_9.6p1`. **Two distinct things are wrong and should not be merged:** (1) the budget is sized against DNS *existence*, but what actually bit is the **AAAA-before-A** window, a different and longer phenomenon — lengthen the ladder and/or prefer the A record for the scan dial; (2) **the operator is told nothing actionable** — the create fails with a generic error, the remedy is "press it again", and nothing says so. The retry is genuinely safe (idempotent on the label; confirmed afterwards — `one_time_secrets` 1, one sub-account, no double-provision) **but that safety is invisible to the person deciding whether pressing again will double-charge them.** **Severity LOW-MEDIUM:** self-clearing, no data at risk — but it is the first thing a new customer's provisioning does and it fails looking like an outage. | **READY** — owner Viktor | — | — | operator | -| **R-306** | Install & onboarding | P3 | **`--preflight-only` says "no state written" and writes state — with an answer that can be wrong.** `_state_put` short-circuits on `DRY_RUN` only (`felhom-host-install.sh:418`), so a preflight-only run creates `/var/lib/felhom-install/state.json`. Observed live: after a run whose banner read `PRE-FLIGHT PASS (mode=byo) — no state written, no install step executed`, the file existed containing `{"completed": [], "dnsmasq_preexisting": "yes"}`. Both the banner and the flag's own comment at line 226 assert the opposite. **The harm is not the file, it is the value**: the runbook recommends preflight-only first, then the same command without the flag, so on a box carrying a Felhom leftover the wrong ownership answer is baked in before the real install begins | **READY (S) — NEW 2026-08-12, RANK 3** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `17af3b65`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | R-300, R-305 | Either make `_state_put` a no-op under `PREFLIGHT_ONLY` (and record ownership at install instead), or correct both claims. A comment asserting an invariant needs a test pinning it | CC | | **R-516** | Install & onboarding | P3 | **[P3-LOW] English a customer meets on a fresh box and its apps' first screens — enumerated by the big night.** MEASURED 2026-09-14 (BIGNIGHT, VM 333, ISO 1.27.1, controller 0.242.0). Felhom-owned: (1) the dashboard menu item **„Debug"**; (2) the dashboard CPU tile **„Load: 0.29 / 0.39 / 0.37"**; (3) the launcher tile **„Filebrowser"** opens a login in English with no Felhom text (R-513); (4) the storage page mixes formal „Adjon hozzá / Csatlakoztasson" with the product's „te". App first screens a household meets before any Felhom text helps: (5) **Uptime Kuma 2.4 opens on „Which database would you like to use?"** (SQLite / Embedded MariaDB, „Next") — the app card's „Első lépések" does not mention it; (6) PrivateBin, Gokapi, AdventureLog and FileBrowser UIs are English (the apps' own). Already rows: the Proxmox installer screens (R-495, answered by the guide), `wiki.DOMAIN` (R-498). **Fix shape:** rename „Debug"/„Load" (controller); add the Uptime Kuma database step to its card, or pre-seed `db-config.json` for SQLite in the template (catalog). **Added by F4 (20:04:54Z):** (7) the storage page prints the disconnect time as a raw ISO UTC string „Leválasztva: 2026-09-14T19:58:02Z"; (8) the „Meghajtó leválasztva" banner appears twice on every page; (9) „4 telepített alkalmazás nem fut — nézze meg a rendszermonitort" uses the formal form. **Added by F7 (20:50–21:00Z, system disk at 95 %):** (10) a banner on every page in English, „**SSD disk usage high: 90%**"; (11) the dashboard tile reads „Rendszer (/) 61.8 GB / 68.7 GB (**90%**)" while `df` reports **95 %** (reserved blocks ignored), and „(/)" labels the data volume `/mnt/sys_drive`; the deploy page says nothing about free disk. **Added by the i18n spike (2026-09-17, controller v0.247.0):** (12) **six formal („ön") forms in the converted dashboard copy** — „Olvassa be telefonnal" (launcher QR hint), „Biztosan kikapcsolja a megosztást?" (launcher), „Ha újratelepíti … importálnia kell" (the layout's remove-app modal), „Kérjük, vegye fel a kapcsolatot" (backups empty state). They were NOT fixed — a localisation release may not change Hungarian bytes — and are now COUNTED by `controller/scripts/i18n_missing_gate.py` (`HU_FORMAL_CEILING` = 6, a ratchet: a new one convicts, fixing one lowers it). The inventory (`audits/I18N-INVENTORY-2026-09-17.md`) is the list this row closes against in localisation slice 6 (R-561). **Extended 2026-09-17 by slice 1 (controller v0.248.0–v0.250.0):** the converted copy now counts **16** formal forms (`HU_FORMAL_CEILING` 6 → 10 → 12 → 16, each raise stated, Hungarian unchanged by rule) — and the count is an UNDER-count: the gate's stem list sees 4 on the release C pages (storage, drive wizards, debug), while a wider list („írja be”, „adja meg”, „adjon hozzá”, „válassza ki”, „biztosan eltávolítja”, „engedélyezze”, „hozzon létre”, „kattintson”) finds at least **22** keys there alone (login „Adja meg a jelszavát”, the NAS guide, the storage confirms). Widen the stems when this row is worked; the ceiling rises with them. **NARROWED 2026-09-20 by localisation slice 6's walk, item by item** (`audits/i18n-slice6-2026-09-20/R-516-item-by-item.md`): item 2 FIXED (slice 1), items 3 and 6 APP-OWNED (and for an English household item 6 is an advantage), item 11 is a wrong NUMBER rather than a language and needs its own row, item 1 is correct English for an English reader and still open for a Hungarian one. **Items 4, 7, 8, 9 and 10 are NOT DECIDABLE FROM AN ENGLISH WALK** - they are about Hungarian copy quality, a disconnected drive and a full disk, none of which this walk had. **What this row is now waiting for is a HUNGARIAN walk on a box with a second drive**, not another English one; saying it closed would be the kind of closure that makes a register stop meaning anything. | **NARROWED 2026-09-20 - rank P3-LOW; owner: CC; needs a Hungarian walk** **2026-10-05 (burn-down night): items 1 and 4 FIXED on controller `main`** (`d7efa7b`): the menu says „Hibakeresés” (English „Debug”); the storage pages use the te-form (formal-form ceiling 17 → 14; 116 parity fixtures changed by exactly those bytes). Item 2 was already fixed. **Left:** items 5–6 (apps' own screens) and the rest. **2026-10-06: items 1 and 4 DELIVERED** (controller v0.298.0). **2026-10-06 (burn-down night, later): items 7–10 and 12 FIXED on controller `main`** (`e4774e6`, `2e5ef7e`): the disconnect time in local time; one banner per unplugged drive; „nézd meg”; the disk/memory/CPU/temperature banners in the household's language (`health.*` keys; the wire text to the hub unchanged and pinned); the last eleven counted formal forms are te-form (ceiling 14 → 0; 119 parity fixtures changed by exactly those bytes). **Left:** items 5–6 (apps' own screens); about 20 formal forms the gate does not count yet („írja be”, „adja meg”, „hozzon létre”, …) on the debug, storage, network-storage and security pages. Item 11 moved to its own row **R-889**. **2026-10-06 02:40: the bundle has NO formal form left** (controller `4c3c203`, unreleased): 49 sentences in the te-form; the gate's stem list widened (it counted 61 on the previous bundle, 0 now) with decoys and one listed third-person exception. **Left (narrowed):** Felhom-owned formal forms in Go and template literals the gate cannot read yet — the network-storage attach errors, escrow and share handlers, the single-copy notice, a settings refusal, the SMART mail, a fill-watch „Kérjük”, and the setup wizard pages; each moves into the bundle in the te-form. Items 5–6 are the apps' own screens (not ours). | — | — | CC | | **R-554** | Install & onboarding | P3 | **[P3-LOW] Delete the first-boot setup wizard — obsolete by design, still reachable.** OPERATOR DECISION 2026-09-17 (localisation starter, decision 4: „out of scope, obsolete"). `02-controller-module-map.md` L56 calls `internal/setup/` obsolete; `cmd/controller/main.go` L322 still enters it when `setup.NeedsSetup(cfg)` — `customer.id` empty after bootstrap ingestion, or a `.needs-setup` marker (`internal/setup/setup.go` L17-25). Ingestion leaves `customer.id` empty on a missing/invalid `bootstrap.json`, a failed hub pull, or a failed merge/write/reload (`internal/bootstrap/bootstrap.go` L109-163) — so a box whose first boot cannot reach the hub shows a household an 8-page wizard (95 Hungarian strings, its own template set and CSRF). **Fix shape:** decide what such a box shows instead (a single „cannot reach Felhom yet, retrying" page — needs no decision beyond copy), then delete `internal/setup/` and `runSetupMode`; red-proof that a failed ingestion renders the waiting page, not a 404. **Check first** whether any drill/golden path still relies on `.needs-setup`. | **READY - rank P3-LOW; owner: CC** | — | — | CC | -| **R-310** | Install & onboarding | P4 | **Two small edges on the installer, neither costing more than a moment.** (1) The R-297 operator-named refusal states the vouched version twice in consecutive sentences (*"…but the vouched golden is 0.213.0. The vouched golden is 0.213.0."*). (2) `--uninstall` reads its typed vmid confirmation from `/dev/tty` and `--force` deliberately does **not** bypass it, so teardown cannot be scripted without a pty — correct for an irreversible destroy, but undocumented; it surfaces as `line 891: /dev/tty: No such device or address` and an rc=1 that looks like a failure rather than a refusal to proceed unattended | **READY (S) — NEW 2026-08-12, RANK 4** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `b43705ac`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | R-297 | Drop the duplicated sentence; add one runbook line naming the pty requirement | CC | | **R-494** | Install & onboarding | P4 | **NARROWED 2026-09-14 by operator ruling → [P3-LOW] the hub COULD create the tunnel at customer creation, for a domain already on Cloudflare. Not blocking: every customer has their own domain and the operator creates the tunnel per day-0 A.1 (`architecture/01-topology-and-trust.md`).** *Original finding, kept:* **[P1-HIGH] A new customer's dashboard has NO reachable address unless the operator hand-makes a Cloudflare tunnel — the link in the setup-code mail is dead.** MEASURED 2026-09-14 on a fresh install from the public ISO (drill intervention **I1**): the claim mail points at `https://felhom.drill0242.felhom.eu`; that name has **no A and no AAAA** record (`dig @1.1.1.1`, control `felhom.enkisfelhom.hu` resolves); the hub has **no tunnel- or DNS-creation code** (`hub/internal/cloudflare/` holds only geo-rule removal; `cf_tunnel_token` is a pasted, optional form field, `configs.go:1478`) — day-0 runbook A.1 makes it a manual Cloudflare-dashboard step that nothing on the customer-create page asks for; the box's own split-horizon resolver on the appliance LAN IP answered `google.com` but not the dashboard name at 13:27:39Z; the agent applied the record at **13:27:44Z** (`lanresolver: applied split-horizon record … ip=192.168.0.158`, 3 m 46 s after the controller started), so the box CAN answer the name — **but only to a device that uses the box as its DNS server, and no document, screen or mail tells a household to do that**; the router and the installer-offered DNS answer nothing. The page was reachable only at the guest's LAN address with the name forced (`curl --resolve …:443:192.168.0.158`). **A volunteer could not have done that.** **What it needs:** an operator ruling — the hub creates the tunnel and DNS at customer creation, or the product gives a household a LAN address that works with no DNS change. | **READY — rank P3-LOW; owner: CC** **Re-ranked 2026-10-03: P3→P4: operator-side automation; the operator creates the tunnel by hand per the day-0 runbook.** | — | — | CC | | **R-504** | Install & onboarding | P4 | **[P3-LOW] `iso.felhom.eu` cannot show an index page on its own — its root returns 404, and the download page lives on the website instead.** MEASURED 2026-09-14: `https://iso.felhom.eu/` and `/index.html` → 404; only named objects answer. The host is an R2 bucket behind a custom domain; whether R2 would serve an uploaded `index.html` at `/` was **not measured** (uploading anything to the public bucket is a publication). The ISO v1.27.0 task puts the Hungarian download page at `felhom.eu/letoltes` (published with the ISO, after the operator's yes). **Remaining:** a redirect from `iso.felhom.eu/` to that page needs a Cloudflare rule the session has no credential for. | **WAITING-ON-OPERATOR — rank P3-LOW; owner: operator (Cloudflare rule)** **Re-ranked 2026-10-03: P3→P4: households are sent to the website's download page; the bare address is cosmetic.** | — | — | operator | -| **R-881** | Install & onboarding | P4 | **The installer's uninstall does not remove `/usr/local/sbin/felhom-priv-apply`** (added to the bundle by agent v0.146.1, R-861), and its comment still says the guest hook is agent-installed at runtime (`scripts/felhom-host-install.sh` ~line 1679) — since v0.146.1 the hook is a bundle file. Found 2026-10-05 while checking the installer against the new bundle. | READY **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `85de3f9b`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Add the file to the uninstall list and fix the comment at the next installer tag | CC | ## Apps & catalog — 14 rows (P3 6, P4 8) @@ -222,8 +216,6 @@ stopping line that lies. | **R-137** | Security & access | P3 | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | | **R-138** | Security & access | P3 | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) | READY (S) **2026-10-05 (burn-down night): NEEDS A DESIGN.** No notion of a „shared zone” exists anywhere; the token is typed into the hub form, so the guard belongs on the hub side with that notion defined. | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | | **R-255** | Security & access | P3 | **The check that would catch a fourth secret-in-the-body covers 4 of 27 pages, and the cheap gate that covers all 36 templates is blind to the shape that actually shipped.** Filed 2026-08-08 while closing R-254, **because a partial guard reported as complete is worse than no guard — it stops the next person looking.** **Two nets, both measured.** **(1) `scripts/secret_in_markup_gate.py`** reads all 36 templates and convicts any `{{ … }}` naming a secret unless allowlisted with a reason. It catches `{{.RetrievalPassword}}` and `{{.InitialCreds.Password}}`, **and it catches a launder through a local variable** because the assignment itself names the secret (`{{$v := .InitialCreds.Password}}` is convicted — verified). **It is blind to a secret arriving under a NEUTRAL PAGE-DATA KEY** — `data["Tagline"] = creds.Password` then `{{.AppInfo.Tagline}}` passes it cleanly, also verified. **That is exactly the shape of R-254 site two** (`value="{{$val}}"` inside an `{{if eq .Type "secret"}}` branch), so the gate **would not have caught one of the three instances it was written for.** **(2) The runtime body assertion** — render the page and grep the response for a sentinel — catches every shape, including that one (demonstrated on the same planted leak the gate missed). But it needs each page's data to be constructible in a test, and **only 4 of 27 page templates have that today**: `settings_security`, `app_info`, `deploy`, `backups_restore` — the four that were touched by R-249/R-252/R-253/R-254 and therefore got their own tests. **The other 23 pages have no runtime coverage at all.** **What closing this needs, so the cost is not re-estimated:** a per-page data fixture for the remaining 23 (most need a wired `Server` — `stackMgr`, `backupMgr`, agent seams), then one table-driven test that renders each with a sentinel substituted for every string in its data and asserts the sentinel is absent. **That is real scaffolding, which is why it was NOT built inside R-254's session** rather than half-built and declared done. | **READY** — owner Viktor | — | — | operator | -| **R-275** | Security & access | P3 | **`--uninstall` leaves five 0600 `agent.json.*` credential backups, and the reinstall hands them to the new service account.** `/etc/felhom-agent/` survives with `agent.json.{campaign8-before,campaign9-before,campaign9-prev,pre-e-target-move,pre-prunegate.bak}`, each carrying a 64-char `hub.api_key` and a 59-char `proxmox.token`. The teardown claims to remove *"config (+ its .bak backups)"* and `scripts/CHANGELOG` F1 records *"uninstall now purges the agent config's `.bak*` siblings (one held a live hub api_key)"* — **that fix does not match the filenames in use, and it misses `agent.json.pre-prunegate.bak`, a file that literally ends in `.bak`.** **Exposure assessed, not assumed:** these are SUPERSEDED — the orphaned key hashes to `a5d2222a…`, the hub's current demo-hp key to `8c59d1b6…`, and the Proxmox token was deleted by the same uninstall. **But the reinstall recreates `felhom-agent` at uid 999, the same uid the deleted account had**, so three of the backups become the new account's files — verified readable as `felhom-agent`. A fresh install's service account inherits read access to the prior install's credentials; superseded today, live if the backups were recent (R-179's precedent). **Also left, undeclared:** `/etc/felhom/{.bootstrap-done,appliance-pairing-code}`, `felhom-bootstrap.service` + `/usr/local/sbin/felhom-bootstrap.sh`, the `vmbr9` stanza in `/etc/network/interfaces`, and `/etc/sudoers.d/felhom-agent.bak-pre-e2a` (21 KB — **INERT: sudo skips dotted filenames, verified with `sudo -l -U felhom-agent`; `visudo -c -f` parsing it OK is NOT evidence sudo loads it**) | **READY (S) — NEW 2026-08-09** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `85de3f9b`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Purge by directory, not by glob; and do not let a new service account reuse a uid that owns old secrets | CC | -| **R-276** | Security & access | P3 | **RANK 2 — an uninstalled box keeps a live WireGuard tunnel into Felhom's off-site endpoint, and the teardown says nothing.** After `--uninstall` on demo-hp, `wg-quick@wg-felhom` is **enabled and active**, `/etc/wireguard/wg-felhom.conf` present, handshake to `167.233.158.164:443` **52 s old**, counters 5.86 GiB in / 2.48 GiB sent. It appears in **neither** the WIPED nor the KEPT list, though the BYO install disclosure names it prominently on the way in (*"an OUTBOUND WireGuard tunnel to the Felhom hub"*). A host told to leave Felhom retains a live network path into Felhom infrastructure, its hub-side peer registration intact, and nobody is told | **READY (S) — NEW 2026-08-09** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `85de3f9b`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). Left besides the tag: the hub deletes a removed host's WireGuard peer and pushes the peer list to the off-site endpoint (a hub change). **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. **Hub half in hub v0.138.0** (felhom.eu `0a460bcd`): host delete pushes the peer list at once. **2026-10-06: hub half DEPLOYED** (hub v0.138.0). Left: the installer half ships with `installer-v1.32.0`. | — | Tear the tunnel down and deregister the peer, or list it under KEPT with the reason and the removal command | CC | | **R-338** | Security & access | P3 | **`demo-hp` is not on the R-50 island at all, and `operations/nodes.md` states that it is.** The page records both fleet boxes as island-migrated 2026-07-25. True of `felhom-pve`; **false of `demo-hp`**, whose `agent.json` has `listen_addr: 192.168.0.87:8443` — the customer LAN address — and **no `island_bridge`/`island_guest_addr` keys at all**, whose guest 9201 has `net0` only (no `eth1`), and whose `vmbr9` exists with **zero members**. The controller's `controller.yaml` points at the LAN address, so the box works; this is inventory drift, not breakage. **Two costs.** A session trusting the page addresses the wrong endpoint — that happened on 2026-08-18 and the resulting timeout was briefly read as a fault. And the agent's local API is **bound to the customer LAN on this box** rather than to a point-to-point island, which is the exposure R-50 was built to remove — so a documented security property is claimed for a box that does not have it **Checked from source 2026-10-05 (burn-down round 2):** nodes.md:86-88 still claims demo-hp is on the R-50 island (`local_api` on 169.254.253.1:8443/vmbr9, guest eth1). git blame: that claim dates from e6b5fa1e (2026-07-30); the 2026-09-21 edit bcdd5b20 re-read addresses but only reworded the lan_resolver clause -- the island claim was NOT re-verified after the reprovision. Agent config path /etc/felhom-agent/agent.json (felhom-agent cmd/felhom-agent/main.go:171), island keys island_bridge (internal/config/config.go:246). | **READY (S) — NEW 2026-08-18** | — | Decide which is true: migrate `demo-hp` to the island, or correct `nodes.md`. Leaving both is the one option that keeps the doc lying | Viktor decides; CC executes | | **R-616** | Security & access | P3 | **[P3-LOW] The catalog credentials are stored in PLAINTEXT in the box's catalog clone and are printed by an ordinary `git remote -v`.** FOUND 2026-09-21 on guest 9202 while pointing it at a private drill catalog. `Syncer.buildRepoURL` injects `username:token` into the HTTPS URL, and `git clone` persists that URL as the clone's `origin`, so `/catalog-cache/.git/config` holds the token in the clear and **any** diagnostic that prints the remote leaks it — which is what happened in this session's own transcript, and is the same shape as R-580 (`curl -w '%{redirect_url}'`). `maskRepoURL` exists and is used for the LOG lines, so the masking intent is already there; the stored remote is the half that was missed. **INERT ON THE FLEET TODAY** — the live catalog is public and `git.token` is empty on every real box — which is exactly why it should be fixed before it is not: the day the catalog goes private, every box carries a readable credential and every support session that runs `git remote -v` prints it. **Fix shape:** store the remote WITHOUT credentials and supply them per-fetch (a credential helper, `http.extraHeader`, or `GIT_ASKPASS`), and a test asserting the clone's stored `origin` contains no `@`. **Operator action from tonight, unrelated to the fix:** the Gitea `admin` token used for the drill repo was printed by that command and must be rotated. Evidence: `audits/update-night-2026-09-21/05-9202-follows-drill.txt` (redacted). | **READY — rank P3-LOW; owner: CC (controller); one operator action (rotate the Gitea admin token)** **2026-10-05 (burn-down night): FIXED on controller `main`** (`28a5203`; the catalog clone stores no credentials; the token is supplied per fetch; R-615's repo comparison ignores credentials on both sides, so a token never re-clones (`TestR616_TokenSetSameRepoNoRecloneOriginClean`) and a credentialed origin is cleaned at the next pull. The operator's Gitea admin token rotation (the row's second half) is still owed). Ships with the next controller release; close after delivery. **2026-10-06: DELIVERED** in controller v0.298.0 (the clone stores no credentials). Left: the operator's Gitea admin token rotation. | — | — | CC + operator | | **R-717** | Security & access | P3 | **[P3-LOW] opengist and wishlist keep their sign-up switch only in their own database — the box closes them with the address block alone.** MEASURED 2026-09-29: opengist `disable-signup` is an admin-panel setting (no env, no CLI); wishlist `system_config.enableSignup` (Prisma). Their blocks are case-insensitive and refused every trick shape (`audits/signup-lock-2026-09-29/B/`). **Fix direction:** an `after_setup` command that sets the database value (wishlist: a Node/Prisma one-liner; opengist: needs its sqlite with the app stopped). | **OPEN — P3; owner: CC** | — | — | CC | @@ -246,7 +238,6 @@ stopping line that lies. | **R-35** | Box system & updates | P3 | **Config-apply should not end the customer's session.** The offsite config push bumped `config_version` 10→11 at 16:54:58 and the controller self-restarted (container `StartedAt` 16:54:59Z, back up 16:55:02); in-memory sessions died with it and **customer zero was force-logged-out mid-flow**. **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-21, size S, roadmap state `idea`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** **2026-10-05 (burn-down night): NEEDS A DESIGN.** Hot-apply needs a per-setting ruling; persisting sessions puts login tokens on disk (and into the whole-guest archive, `07` §5). Next: the operator's pick between them. | — | Direction: **hot-apply the offbox target** (no restart for a config the running process can adopt), or **persist sessions** across restart. The restart itself is by design — the collateral is not. Evidence `controller-log-full.txt` | CC | | **R-78** | Box system & updates | P3 | **`local_api` authority ruling — auto-reconcile vs detect-only** **MIGRATED FROM `ROADMAP.md` 2026-08-22 (R-369) — originally filed 2026-07-25, size M, roadmap state `idea (deferred OUT of R-77 on purpose)`.** Moved verbatim; nothing added or reinterpreted. The roadmap keeps its copy as history, marked moved. **An OWED OPERATOR DECISION, not a defect — moved because an owed ruling hidden among feature ideas is the shape this session exists to remove.** | **OPEN — migrated from ROADMAP 2026-08-22, rank unchanged** | — | R-77 ships detection because the fix is genuinely undecided, and **both directions can lose customer-visible function**. **Direction 1 (today):** `controller.yaml` wins and drift is silent → the 2026-07-25 island migration blinded the whole fleet's control plane for 17.5 h (drive gate, guest-reboot recovery, quiesce/backup all degrade). R-77 makes that loud but does not stop it recurring. **Direction 2 (`bootstrap.json` wins, auto-reconcile on boot):** a guest whose `controller.yaml` is CORRECT and whose `bootstrap.json` is stale — a half-completed re-provision, a hand-repaired guest, a setup-wizard box — gets a **working channel clobbered on the next restart**, fleet-wide and silently, during a routine deploy. That is not obviously better than the bug. Needs a spike: which writer is authoritative per field (endpoint vs fingerprint vs token — `mergeLocalAPI` replaces the whole block, so they cannot be reconciled independently today), whether the agent side should stamp a generation/mtime so 'newer wins' is even expressible, and whether reconcile should require an operator ack. Until then the drift alert plus a manual edit is the supported path. | CC | | **R-242** | Box system & updates | P3 | **A controller release that changes customer-visible behaviour is not delivered until a golden carries it — and nothing enforces that.** R-239 is the symptom; this is the mechanism, recorded **2026-08-07** and deliberately **NOT built** (the task that found it scoped it as a record-only item). **Two releases went out without a golden and the gap was invisible until a walk measured it from the customer's side**: v0.204.0 (R-237) and v0.205.0 (R-234) were written, tested, pushed, and CHANGELOG'd, and every one of those steps passed while a machine installed that night received neither. The register said CLOSED; the fleet said otherwise. **Nothing in the release path knows a golden exists.** The version bump, the image push, the CHANGELOG entry and the register closure are all repo-local; the manifest's `golden_version` is edited by a separate operator act, in a different repo, with no link back. **Proposed shapes, cheapest first — the choice is the operator's and is not taken here.** (a) **A release-path checklist step** — one line in the controller's end-of-session checklist: *a release that changes customer-visible behaviour is not finished until a golden carries it or a register row says why not.* Costs nothing, catches nothing mechanically. (b) **A gate in `repo_gates.py`** comparing the manifest's `golden_version` against the newest released controller and FAILING (or warning) past a tolerance of one minor. Mechanical, runs on every push, and would have fired the morning after v0.204.0. (c) **A hub-side checker** — the hub already knows every box's running controller version from `/hosts` and the vouched golden from the manifest; a periodic comparison against the newest published image would catch drift the repo cannot see, including a vouch that was made and then rolled back. **Earliest catch: (b).** It fires on the push that creates the gap, before any box is installed, and it needs no live fleet. **(c) catches strictly more but only after boxes exist.** (a) is worth doing regardless because it is free. **Not built. No gate was written this session.** **⚠ IT RECURRED WITHIN A DAY, WHICH IS THE ARGUMENT FOR BUILDING IT.** Controller **v0.206.0** shipped the R-241 fixes on 2026-08-07 while the vouched golden still carried **0.205.0** — so a machine installed on the morning of 2026-08-08 would have received neither. Third occurrence of the shape in three days (R-111/R-115/R-120 are the older family). **✅ SHAPE (b) BUILT 2026-08-08 — `scripts/golden_currency_gate.py`, registered in `repo_gates.py` as gate 7.** **It was shown FAILING against that exact state before anything was baked**, which is its red-proof and the reason its own introducing push needed `--no-verify` (stated in the session report rather than worked around): `newest released controller : 0.206.0 / newest golden baked : 0.205.0 → CONVICTED`. **IT IS `--fast`, AND THAT FORCED ITS DESIGN:** both `.githooks/pre-push` AND CI run `repo_gates.py --fast`, so a non-fast gate would run in NEITHER — the R-29 census failure this runner exists to end. **THEREFORE IT CHECKS THE BAKE, NOT THE VOUCH**, because the vouched version lives only in the hub's `hub_settings` with no copy in git, and putting a copy there would create a second source of truth that can drift — a green gate over a false claim being the worst outcome available. **A bake without a vouch still passes: that half is NOT closed and stays on this row.** It also compares versions rather than behaviour, so a release changing nothing customer-visible trips it too — accepted deliberately, because judging that by hand is what failed three times and the cost of a false trip is one bake; a waiver belongs here, never in a habit of bypassing. **VOUCHED 2026-08-08 with the operator's approval** — golden `0.206.0` / sha `c85230b4…108e`; `agent_version` and `min_agent` both stayed `0.127.0`, and `wrapper_sha256` was carried through explicitly because the handler clears it when omitted. **The gate was CONVICTED before the bake and OK after it** — red→green on the same command, which is its proof that it measures something real. **⚠ THE GATE FIRED FOR REAL, 2026-08-08 — and it was right.** Controller **v0.207.0** (R-249/R-252/R-253) is released, tested and pushed, and **no golden carries it** — the newest bake is 0.206.0 — so `golden_currency_gate.py` FAILED, saying exactly the true thing: *a machine installed right now would receive v0.206.0*. **The `felhom.eu` push therefore used `git push --no-verify`, declared here, in the commit message and in the session report.** A bypass and NOT a waiver, deliberately: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.207.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; the vouch is a three-field change). **This row's own remaining half is unchanged — nothing gates the VOUCH itself.** **✅ THE OWED BAKE IS DONE, SAME DAY — golden 0.207.0 baked, published, round-trip verified and VOUCHED (2026-08-08).** The gate went from red to **green**, and the `--no-verify` bypass declared above is now historical rather than standing. **Round trip is the evidence, not the build log:** the published bytes were downloaded back — 656 879 192 B, sha256 `20ec9602…22995`, both identical to what the bake reported — and **`./etc/felhom-controller-image` read OUT of the downloaded archive says `felhom-controller:0.207.0`**, which is the delivered artifact naming the controller it will start. **The vouch was a three-field change with all three checked deliberately** (`MinAgent 0.127.0` read from the golden's controller CHANGELOG header, not assumed; `agent_version` already ≥ it; `min_agent` not above `agent_version`, so not the R-216 shape) and verified by **re-reading the manifest rather than trusting the flash**. **This row's remaining half is UNCHANGED and is the whole of what is still open: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. Evidence: `tests/golden-0.207.0-2026-08-08/`. **⚠ RED AGAIN, 2026-08-08 (second time in two days) — controller v0.208.0 (R-254) is released and the vouched golden is 0.207.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.207.0 and none of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit, the CHANGELOG and the session report** — **a bypass, not a waiver**, on the same reasoning as yesterday: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **Owed: bake golden 0.208.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change, `MinAgent 0.127.0` unchanged). **Note the cadence this is establishing: two releases, two bakes owed within 24 h.** That is the argument for this row's OTHER half — nothing gates the vouch, so the only thing standing between a release and an undelivered fleet is somebody remembering. **⚠ RED AGAIN, 2026-08-30 — controller v0.224.0 (R-330) and v0.225.0 (R-331) are released and the newest golden carries 0.223.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.223.0 and neither of today's fixes. **The `felhom.eu` push used `git push --no-verify`, declared in the commit message, in `hub/CHANGELOG.md` and in `REPORT.md` — a BYPASS, not a waiver**, on the same reasoning as the two 2026-08-08 entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and these need one. **The operator was asked and ruled bypass-now-bake-later on 2026-08-30**, on the stated ground that neither fix bites a DAY-0 box — R-330 is a nightly false alarm about apps a new box has not installed yet, and R-331 is a hub-side display over backups a new box has not taken yet — and both arrive by self-update afterwards. **That ground is recorded because it is the thing to re-check, not a general licence: the next release that changes first-boot behaviour cannot reuse it.** **OWED: bake a golden carrying 0.225.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change — `golden_version` + `agent_version` + `min_agent`; MinAgent is 0.129.0 per both CHANGELOG headers). **Cadence note, unchanged and now worse: this is the fourth bypass of this gate, and the gap it names is now two releases wide rather than one.** **⚠ WIDENED TO THREE THE SAME DAY — v0.226.0 (R-353/R-357/R-358/R-360) shipped 2026-08-30 and the golden still carries 0.223.0.** The `felhom.eu` push carrying that release's documentation used `git push --no-verify` on the operator's standing ruling from earlier the same day, declared in the commit and in `REPORT.md`. **The day-0 ground still holds for all three and was re-checked rather than assumed:** R-330 alarms about apps a new box has not installed; R-331 is a hub display over backups a new box has not taken; **R-353/357/358/360 are restore-surface fixes, and a day-0 box has nothing to restore.** **The ground expires the moment a release changes first-boot behaviour — that is the thing to re-check, not a licence.** **Owed: ONE bake carrying 0.226.0 covers all three** (`RUNBOOK-manual-build.md` §4.1; three-field vouch, MinAgent 0.129.0), then raise the floor. **✅ PAID THE SAME DAY — golden `0.226.1` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-30).** Evidence: `documentation/tests/golden-0.226.1-2026-08-30/`. `golden_currency_gate.py` went **red → green** on the same command, which is its proof that it measures something real. **The three declared bypasses above are now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back — **657 197 592 B, sha256 `70ed8e93…baefe69`**, both identical to what the bake reported — and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.226.1`, i.e. the delivered artifact naming the controller it will start. **The three-field vouch was checked deliberately, not assumed:** `MinAgent 0.129.0` read from the golden's controller CHANGELOG header, `agent_version 0.130.0 ≥ min_agent 0.129.0` (so NOT the R-216 shape), and the result verified by **re-reading the manifest** rather than trusting the flash — golden option `0.226.1 SELECTED`, all four shas matching. **The floor is proven ACTING, not merely set:** `demo-felhom` self-updated within 30 s, logging `[selfupdate] Post-update startup: update successful (0.225.0 → 0.226.1)`. **⚠ AND IT HAPPENED AGAIN THE SAME DAY, AND WAS PAID AGAIN.** v0.227.0/v0.227.1 (R-359/R-397) shipped after the 0.226.1 bake, the gate convicted a fifth time, that `felhom.eu` push used `--no-verify` and declared it, and golden **0.227.1** was baked, published, round-trip verified, **VOUCHED** and the floor **RAISED to 0.227.1** within the hour. Evidence: `documentation/tests/golden-0.227.1-2026-08-30/`. **THE CADENCE IS NOW MEASURED RATHER THAN ASSERTED: five convictions and two full bakes in one day.** Every bypass was declared and every debt was paid — but the pattern this row exists to name is exactly that a release and its delivery are separate acts, performed hours apart, by whoever remembers. **The floor was proven ACTING both times:** `demo-felhom` self-updated 0.225.0→0.226.1, then 0.226.1→0.227.1 — and the second time it also registered the new `offsite-integrity` job **by itself, on a box nobody deployed to**, which is the strongest evidence this row has ever carried that a floor delivers rather than merely records. **This row's OTHER half is still open and untouched: nothing gates the VOUCH itself** — the currency gate's own docstring says it checks the bake, so a baked-but-unvouched golden still passes it silently. **2026-08-31, the SEVENTH debt and it was paid the same day — twice in one day.** v0.230.0 shipped in the morning with the newest golden at 0.229.0, **which is the build R-403 says deletes a good copy**, so the gate was red across four commits (`dddcc80`, `6e550ae`, `130f7a6`, `32a4c35`). Golden **0.230.0** baked, published, round-trip verified, vouched, and the fleet floor raised 0.229.0 → 0.230.0; `demo-felhom` moved itself off the defective build unattended (`controller-swap: new controller healthy`, 16:21:40 CEST). Evidence: `documentation/tests/golden-0.230.0-2026-08-31/`. **The gate did its job and its own weakness surfaced doing it — R-410.** | **READY — the vouch half only** — owner Viktor **⚠ SIXTH CONVICTION, 2026-08-31 — controller v0.229.0 (R-102/R-103) is released and the newest golden carries 0.228.0.** `golden_currency_gate.py` FAILS, correctly: a machine installed right now receives 0.228.0 and neither of today's fixes. The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - **a BYPASS, not a waiver**, on the same reasoning as the five entries above: the gate offers a waiver only for a release that *deliberately needs no golden*, and this one needs one. **The day-0 ground was RE-CHECKED rather than reused:** R-102 and R-103 are restore-surface changes on the Tier-2 card, and a day-0 box has taken no Tier-2 copy and has nothing to restore from one; no first-boot behaviour changed, and `MinAgent` is unchanged at 0.129.0. **The ground still expires the moment a release changes first-boot behaviour.** **OWED: bake a golden carrying 0.229.0 and vouch it** (`RUNBOOK-manual-build.md` §4.1; three-field change - `golden_version` + `agent_version` + `min_agent`, MinAgent 0.129.0), then raise the floor. Fleet floor and golden are 0.228.0 today. **Golden and fleet delivery are the operator's (this row).** **✅ PAID THE SAME DAY — golden `0.229.0` baked, published, round-trip verified, VOUCHED, and the floor RAISED (2026-08-31).** Evidence: `documentation/tests/golden-0.229.0-2026-08-31/`. `golden_currency_gate.py` went **red to green** on the same command, which is its proof that it measures something real. **The `--no-verify` bypass declared above is now HISTORICAL rather than standing.** **The round trip is the evidence, not the build log:** the published bytes were downloaded back - **656 864 331 B, sha256 `39aa886d…d7bdae87`**, both identical to what the bake reported - and `./etc/felhom-controller-image` read **out of the downloaded archive** says `felhom-controller:0.229.0`. **A THIRD independent reader agreed before anything was vouched:** the hub's own Day-0 dropdown read the same sha straight from Gitea, a different code path from the round trip. **The three-field vouch was checked deliberately, not assumed** (`MinAgent 0.129.0` read from the golden's controller CHANGELOG header; `agent_version 0.130.0` >= `min_agent 0.129.0`, so NOT the R-216 shape; `agent_sha256` and `wrapper_sha256` carried through explicitly because the handler clears a field it is not sent), and verified by **re-reading the manifest** rather than trusting the flash. The **R-120 gate passed rather than being bypassed** - fleet newest 0.229.0, golden 0.229.0. **The floor is proven ACTING:** `demo-felhom` self-updated `0.228.0 -> 0.229.0` and logged `settle-gate: GO - at/above floor 0.229.0 (we are 0.229.0)` - **nobody deployed to that box.** **Cadence note: this is the SECOND bake in one day (0.228.0 then 0.229.0) and the sixth conviction, and both debts were paid within the hour.** This row's OTHER half is still open and untouched: **nothing gates the VOUCH itself** - the currency gate checks the bake, so a baked-but-unvouched golden still passes it silently. **⚠ SEVENTH CONVICTION, 2026-08-31 — controller v0.230.0 (R-403) is released and the newest golden carries 0.229.0. AND THIS ONE IS NOT LIKE THE OTHERS: the day-0 ground does NOT apply and must not be reused.** Every previous bypass rested on 'a day-0 box has nothing to restore / nothing to alarm about yet'. R-403 is a defect in the NIGHTLY TIER-2 COPY, which a day-0 box starts running on its first night: a machine installed on 0.229.0 can have a complete recovery package on its second drive replaced by an empty one, and that is measured, not suspected (120 082 104 B -> 7 036 B on demo-hp). **The row's own standing sentence - 'the ground expires the moment a release changes first-boot behaviour' - is what expires it here.** The `felhom.eu` push carrying this release's documentation used `git push --no-verify`, declared in the commit message and in `felhom-controller/REPORT.md` - a BYPASS, not a waiver. **OWED, and more urgent than the previous six: bake a golden carrying 0.230.0, vouch it (three fields, MinAgent 0.129.0 unchanged), and raise the floor.** `demo-hp` was updated by hand; `demo-felhom` is still on 0.229.0 and still carries the defect. **See also R-404, filed today: this is the seventh bypass and the habit is now the thing being reported.** **2026-09-01 (R-404): THE BAKE HALF IS UNCHANGED AND THE VOUCH HALF IS STILL OPEN.** R-404 moved WHO the bake check refuses and added a notice in the controller repo; it did NOT touch what is checked. **Nothing gates the VOUCH.** A baked-but-unvouched golden still passes both the gate and the new notice, and the reason is unchanged and forced: the vouched version lives only in the hub's `hub_settings` table, there is no copy in git, and a hub-reading gate could not be `--fast` so it would run in neither the hook nor CI. **Do not read R-404's closure as closing this.** **2026-09-13 — NARROWED: the WAIVER half is BUILT (R-468).** The docstring's *"honest fix is a recorded waiver in the register, never a habit of bypassing"* is now a mechanism: `golden_currency_gate.py` reads `documentation/tests/golden-waiver.yml` (dated, ≤ 14 days, row-bound), turns a BEHIND conviction into a loud advisory while valid, and is red again when it expires — the difference from this row's original rule, which recurred the next day, is that a dated waiver cannot be forgotten. It never covers an UNRECORDED golden (R-385). Operator ruling the same day: goldens weekly and before any install, not per release. **What stays open on THIS row is exactly one thing: nothing gates the VOUCH.** The waiver does not touch it, and the reason it is unbuilt is unchanged (the vouched version lives only in the hub). | — | — | operator | -| **R-274** | Box system & updates | P3 | **A local golden is adopted with NO version and NO checksum check, so a reinstall can silently come up releases behind.** `felhom-host-install.sh` step 7: `if [[ -n "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then log_skip "using local golden"; return 0; fi` — **the hub manifest's `golden.sha256`, whose entire purpose is to vouch from a different trust root than Gitea, is consulted only on the FETCH path.** A locally-present archive bypasses the vouch: no version compare, no digest, no warning. On demo-hp 2026-08-09 the preflight selected `local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst`, whose baked marker reads `felhom-controller:`**`0.192.0`**, against a vouched golden of **0.210.0** — 18 releases stale. **The sharp consequence:** 0.192.0 is **below 0.200.0, where R-193's off-site recovery SCREEN shipped**, so a customer reinstalled today returns on a controller that cannot run the recovery ceremony their data depends on; it is also born below the managed-update floor (0.200.0), and the updater's auto-target is the floor, never the newest. **It compounds with the teardown**, which deliberately keeps the old golden (*"golden vzdump left in place"*). This is the R-111/R-115/R-120 drift family one layer down: the R-120 gate guards what may be VOUCHED, nothing guards what an install TAKES. **OBSERVED 2026-08-09, AND THE RESULT NARROWS THE ROW — recorded because it partly refutes what was written above.** On the RESUME path step 7 **fetched the vouched 0.210.0 correctly** (`fetching golden v0.210.0 from Gitea`), because `--resume` skips preflight and preflight is where local auto-discovery sets `GOLDEN_VOLID` (the GL6-F4 comment says so). **So the fresh-install and resume paths disagree on golden selection, and the resume path is the safe one.** Discovery is `… | sort | tail -1`, i.e. the NEWEST local archive by filename — a sensible heuristic, **and still no comparison against the manifest's version or sha**. The defect therefore stands as: *a box whose newest local golden predates the vouched one installs stale, silently* — which is exactly the state demo-hp was in before this run (newest local 0.192.0 vs vouched 0.210.0). It is now masked on this box because the freshly fetched 0.210.0 is the newest — **correct by recency, not by verification**. There are now **three** goldens on `local` (07-21, 08-03, 08-09), because the teardown keeps them. **Still not observed: a FRESH (non-resume) install taking a stale local golden.** | **VERIFY** (2026-10-03 triage: Local golden is now checked against the hub manifest's version and sha before use (GOLDEN_CHECK_WHY, R-297) — felhom.eu/scripts/felhom-host-install.sh:2855-2905) — **READY (S) — NEW 2026-08-09, NARROWED same day** **2026-10-05 (burn-down night): FIXED on `main`** (felhom.eu `f790734e`, installer 1.32.0, test + red-proof in `scripts/test_hostinstall.py`). The version/checksum check itself shipped with R-297; tonight added the disclosure line and the tests that pin the check. **Ships with the tag `installer-v1.32.0`** (not cut tonight — not in the night's release list); close at the tag with the live check. | — | Compare the local golden's version/sha against the manifest and refuse or re-fetch on mismatch; say so in the BYO disclosure, which today lists only what the install CREATES, never what it REUSES | CC | | **R-444** | Box system & updates | P3 | **[P3-LOW] Nothing runs `pct fstrim` on the fleet, and demo-hp's thin pool was carrying ~23.8 GB of blocks the guest had already freed.** MEASURED 2026-09-01 during this spike's teardown: the run itself added ~1.05 GiB that `local-lvm` did not reclaim on delete (68.97% -> 70.91%); `fstrim` INSIDE the unprivileged container is refused (`FITRIM ioctl failed: Operation not permitted`, all three mounts); `pct fstrim 9201` from the PVE host then trimmed **30.2 GiB + 57 GiB** and took `local-lvm` to **26.78%** — **23.8 GB BELOW this run's own starting point**, i.e. the surplus was long-standing, not ours. **Why it is not merely housekeeping:** a thin pool that only ever grows can reach 100% from DELETED data alone, and a full thin pool takes every guest on the host read-only. demo-hp had 16.4 GB free before the trim. **Not urgent, and the row says so** — but the appliance has no periodic trim and no operator surface reports the gap between guest-free and pool-used. Owner: **CC.** `audits/SPIKE-app-update-2026-09-01.md` | **OPEN — rank P3-LOW; owner: CC** **2026-10-05 (burn-down night): NEEDS THE OPERATOR.** A scheduled trim needs a new root grant (`pct fstrim `, widening what R-861 narrowed) and a new scheduled action on every customer guest's disks. Next: an operator yes on the grant and a cadence, then one measured run on demo-hp. | — | — | CC | | **R-468** | Box system & updates | P3 | **[P3-LOW] THE GOLDEN WAIVER — goldens on a cadence, not per release (operator ruling 2026-09-13).** 25 goldens in 26 days in August, almost one per release, because `golden_currency_gate.py` trips on every release by design and the only honest ways past it were a bake or a declared `--no-verify` (thirteen by 2026-09-01, R-404/R-417). **The ruling: bake WEEKLY, and always before any drill or fresh install.** Every release still raises the FLOOR, so both demo boxes keep getting each release in ~20 s; only the golden — which protects a fresh install and nothing else — moves to a cadence. **The mechanism (built 2026-09-13):** `documentation/tests/golden-waiver. **⚠ CORRECTED THE SAME DAY (R-472): between bakes the floor does NOT carry a release — the hub holds any floor above the vouched golden (publish-train rule 1), so releases between bakes reach the demo boxes only by hand-deploy.**yml`, four lines (`issued`, `expires`, `reason`, `register_row: R-468`), read by the gate. While valid, a golden BEHIND the record makes the gate print a loud ADVISORY and exit 0; when it expires the gate is red again until someone bakes or renews. **The 14-day cap is enforced by the gate, not the runbook** — a longer, undated, unparseable, reason-less or row-less waiver is INCONCLUSIVE (exit 2), never 0 and never silently ignored. **It never covers a golden that is UNRECORDED (R-385)** — that is not a cadence choice. **A dated waiver cannot be forgotten; it just expires** — the difference from R-242's original rule, which recurred the day after it was written. Tests: `scripts/test_golden_currency_gate.py` cases 5–15 (E/F/G/H, a 15-day, absent, unparseable, bad-row and empty-reason waiver each 2; the R-421 decoy — a file saying only `expires` — 2). **This is a PRE-CUSTOMER arrangement: the first external install retires it** (delete the file in that commit). Cadence written into `RUNBOOK-manual-build.md` §4.2 and the `felhom.eu` end-of-session checklist. **Does NOT touch R-242's open half (nothing gates the VOUCH).** | **WATCHING — rank P3-LOW; owner: CC (renew ≤ 14 days or bake); retire at the first external install** | — | — | CC | | **R-836** | Box system & updates | P3 | **A new host kernel that hangs before userspace stays the GRUB default: `--next-boot` is not a one-shot on these hosts.** MEASURED 2026-10-04 on demo-hp (operator's word, 2 reboots): both demo hosts boot UEFI + GRUB without proxmox-boot-tool ESPs; installing a kernel makes it the default at once; `kernel pin --next-boot` writes an ordinary `GRUB_DEFAULT`, and `proxmox-boot-cleanup.service` clears it only after a boot reaches userspace. With the old kernel pinned FIRST, the fallback after a good boot worked (new 60 s, old 76 s). READ FROM THE CODE, not measured: a hang leaves the new kernel default on every power cycle. Only `softdog` runs (useless before userspace); demo-hp's `sp5100_tco` ships unloaded, untested. Fix direction for the slow lane: GRUB's own one-shot (`GRUB_DEFAULT=saved` + `grub-reboot`) with the old kernel saved — to be measured, including Secure Boot (ON on demo-hp). `audits/os-updates-spike-2026-10-04/partH/` | **NARROWED 2026-10-04 (os-host-lane Part E, operator's word before each of 2 reboots) — GRUB's own one-shot is NOT a one-shot here either.** `GRUB_DEFAULT=saved` (old kernel saved) + `grub-reboot `: boot 1 → new kernel, **Secure Boot ON and fine**; but GRUB could not clear `next_entry` (`grub-reboot` itself warns: *environment block on lvm device … will remain the default until manually cleared*; `/boot` is ext4 on LVM `pve-root`), so boot 2 (no command) → **the new kernel again**. `kernel.panic = 0`: a panic leaves the host stopped (R-851). `sp5100_tco` LOADS and answers (`SP5100 TCO timer`, 60 s, inactive, nowayout 0; read from sysfs, never opened, unloaded) — a hardware watchdog exists on demo-hp, but nothing arms it before userspace. demo-hp left on 7.0.14-20 with that as the saved default. **LEFT (fix direction):** a GRUB env block GRUB can write (on the ESP, vfat) or a userspace "boot good" step that rewrites the default, plus arming `sp5100_tco`; to be measured before the kernel slow lane. `audits/os-host-lane-2026-10-04/partE/` **READY — owner: CC + operator (reboots).** | — | — | CC | diff --git a/documentation/tests/golden-0.299.0-2026-10-06/02-round-trip.txt b/documentation/tests/golden-0.299.0-2026-10-06/02-round-trip.txt new file mode 100644 index 00000000..229caac9 --- /dev/null +++ b/documentation/tests/golden-0.299.0-2026-10-06/02-round-trip.txt @@ -0,0 +1,3 @@ +== round trip 2026-10-06T06:13:10Z: downloaded golden.tar.zst 0.299.0 from the registry +sha256 9948aa1682ada23afe5aa9fdb056c5b335f264823083b4f020bf4eb34dcd7de0 +bake 9948aa1682ada23afe5aa9fdb056c5b335f264823083b4f020bf4eb34dcd7de0 diff --git a/documentation/tests/golden-0.299.0-2026-10-06/README.md b/documentation/tests/golden-0.299.0-2026-10-06/README.md new file mode 100644 index 00000000..1a017dfc --- /dev/null +++ b/documentation/tests/golden-0.299.0-2026-10-06/README.md @@ -0,0 +1,29 @@ +# Golden 0.299.0 — bake + publish, 2026-10-06 (morning; R-889 release) + +Procedure: `documentation/runbooks/RUNBOOK-manual-build.md` §4.0 and §4.1 steps 1–4, in the drill VM on DooPlex. + +| | Previous (`../golden-0.298.0-2026-10-06/`) | This bake | +|---|---|---| +| `build-golden.sh` | v3.2.0 | same file, unchanged | +| Controller | `felhom-controller:0.298.0` | **`felhom-controller:0.299.0`** (MinAgent 0.131.0, unchanged) | +| Docker engine | approved set `os-docker-20261004-142842` | same — `GOLDEN_DOCKER_PKGS` set in the runner from the start (the hub's System page still names `os-docker-20261004-142842`, read 2026-10-06 08:05) | + +## Pass markers (from `bake.log`, this folder) + +``` +[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie … docker-ce=5:29.8.2-1~debian.13~trixie … + docker OK (overlay2; data-root /var/lib/docker) +INFO: including mount point rootfs ('/') in backup +INFO: including mount point mp0 ('/var/lib/felhom') in backup +[golden] upload OK (HTTP 201) +GOLDEN_VERSION=0.299.0 +GOLDEN_SHA256=9948aa1682ada23afe5aa9fdb056c5b335f264823083b4f020bf4eb34dcd7de0 +``` + +No `excluding`, no `FATAL`, no `GOLDEN_DOCKER_PKGS not set` (a grep for `not set` hits only the six `locale: Cannot set LC_…` +noise lines, as in the 0.298.0 log). Round trip: the registry's file hashed = the bake's sha (`02-round-trip.txt`). +Token: unit properties grep = 0; saved log grep = 0, with a working control (= 1 on a copy with the token appended). + +## Teardown + +Build guest 9100 destroyed (`pct list` empty); token, runner and log shredded in the VM; VM off (no qemu process); disk on `virgin`. diff --git a/documentation/tests/golden-0.299.0-2026-10-06/bake.log b/documentation/tests/golden-0.299.0-2026-10-06/bake.log new file mode 100644 index 00000000..fd672e17 --- /dev/null +++ b/documentation/tests/golden-0.299.0-2026-10-06/bake.log @@ -0,0 +1,340 @@ +[golden] build-golden.sh v3.2.0 — baking controller gitea.dooplex.hu/admin/felhom-controller:0.299.0 +[golden] creating build LXC 9100 (nesting=1,keyctl=1, unprivileged; rootfs 32G + ONE data volume 24G @ /var/lib/felhom, backup=1) … + Logical volume "vm-9100-disk-0" created. + Logical volume pve/vm-9100-disk-0 changed. +Creating filesystem with 8388608 4k blocks and 2097152 inodes +Filesystem UUID: b3f18ec6-d5be-4136-90f5-0f8e20d67518 +Superblock backups stored on blocks: + 32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208, + 4096000, 7962624 + Logical volume "vm-9100-disk-1" created. + Logical volume pve/vm-9100-disk-1 changed. +Creating filesystem with 6291456 4k blocks and 1572864 inodes +Filesystem UUID: 071ecdfa-1642-4374-b6a0-c44215fef694 +Superblock backups stored on blocks: + 32768, 98304, 163840, 229376, 294912, 819200, 884736, 1605632, 2654208, +extracting archive '/var/lib/vz/template/cache/debian-13-standard_13.6-1_amd64.tar.zst' +Total bytes read: 553512960 (528MiB, 82MiB/s) +Detected container architecture: amd64 +Creating SSH host key 'ssh_host_ed25519_key' - this may take some time ... +done: SHA256:8BAPhX5ZWBEb92OwBw1TF6248FKko9pn0GLtVE9wCYc root@felhom-golden +Creating SSH host key 'ssh_host_ecdsa_key' - this may take some time ... +done: SHA256:T3S9x9q7kIFLlOQ8aDjRMrBUNVLYimG9XJCPhlpCHy4 root@felhom-golden +Creating SSH host key 'ssh_host_rsa_key' - this may take some time ... +done: SHA256:/EhZt/Vrz0iHqEQK4ej+FKYcvHRw2VXFHZWlbV2wwR8 root@felhom-golden +[golden] starting + installing Docker (official repo, trixie channel) … +[golden] Docker engine set PINNED to the approved release: containerd.io=2.3.6-1~debian.13~trixie docker-buildx-plugin=0.37.1-1~debian.13~trixie docker-ce=5:29.8.2-1~debian.13~trixie docker-ce-cli=5:29.8.2-1~debian.13~trixie docker-ce-rootless-extras=5:29.8.2-1~debian.13~trixie docker-compose-plugin=5.6.0-1~debian.13~trixie +apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct! +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = (unset), + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to the standard locale ("C"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_MESSAGES to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory +apt-listchanges: Can't set locale; make sure $LC_* and $LANG are correct! +perl: warning: Setting locale failed. +perl: warning: Please check that your locale settings: + LANGUAGE = (unset), + LC_ALL = (unset), + LC_CTYPE = (unset), + LC_NUMERIC = (unset), + LC_COLLATE = (unset), + LC_TIME = (unset), + LC_MESSAGES = (unset), + LC_MONETARY = (unset), + LC_ADDRESS = (unset), + LC_IDENTIFICATION = (unset), + LC_MEASUREMENT = (unset), + LC_PAPER = (unset), + LC_TELEPHONE = (unset), + LC_NAME = (unset), + LANG = "en_US.UTF-8" + are supported and installed on your system. +perl: warning: Falling back to the standard locale ("C"). +locale: Cannot set LC_CTYPE to default locale: No such file or directory +locale: Cannot set LC_MESSAGES to default locale: No such file or directory +locale: Cannot set LC_ALL to default locale: No such file or directory + installed: containerd.io 2.3.6-1~debian.13~trixie + installed: docker-buildx-plugin 0.37.1-1~debian.13~trixie + installed: docker-ce 5:29.8.2-1~debian.13~trixie + installed: docker-ce-cli 5:29.8.2-1~debian.13~trixie + installed: docker-ce-rootless-extras 5:29.8.2-1~debian.13~trixie + installed: docker-compose-plugin 5.6.0-1~debian.13~trixie +[golden] no approved guest release given - the template versions stay; first-night count vs an approved release: n/a +[golden] pending Debian upgrades in the baked guest (what a FUTURE approval may bring): 49 +[golden] baking daemon.json: classic overlay2 driver (containerd-snapshotter OFF) + log rotation … +[golden] wiring the single data volume (R-165 variant V-c): /var/lib/felhom/{docker,sys_drive} -> binds … +[golden] verifying Docker works in the build guest (storage driver should be overlay2 on the ext4 data volume) … +Unable to find image 'hello-world:latest' locally +latest: Pulling from library/hello-world +4f55086f7dd0: Pulling fs layer +4f55086f7dd0: Verifying Checksum +4f55086f7dd0: Download complete +4f55086f7dd0: Pull complete +Digest: sha256:5e23090353324d887c48ad5e5c56d294eab81588df9605b07d1afe895f9cc8f8 +Status: Downloaded newer image for hello-world:latest + docker OK (overlay2; data-root /var/lib/docker) + live-restore: on + /var/lib/docker is a real mount: /dev/mapper/pve-vm--9100--disk--1[/docker] ext4 + /mnt/sys_drive is a real mount: /dev/mapper/pve-vm--9100--disk--1[/sys_drive] ext4 + both paths are ONE filesystem: /dev/mapper/pve-vm--9100--disk--1 23317576 +[golden] baking the in-guest controller image gitea.dooplex.hu/admin/felhom-controller:0.299.0 (no registry cred at deploy) … + +WARNING! Your credentials are stored unencrypted in '/root/.docker/config.json'. +Configure a credential helper to remove this warning. See +https://docs.docker.com/go/credential-store/ + +0.299.0: Pulling from admin/felhom-controller +774043ccc8cc: Pulling fs layer +ab6b448d4be9: Pulling fs layer +23a5bfa58353: Pulling fs layer +862a57157567: Pulling fs layer +41c8e5018d25: Pulling fs layer +f01974600e7b: Pulling fs layer +862a57157567: Waiting +41c8e5018d25: Waiting +f01974600e7b: Waiting +774043ccc8cc: Download complete +862a57157567: Verifying Checksum +862a57157567: Download complete +23a5bfa58353: Verifying Checksum +23a5bfa58353: Download complete +41c8e5018d25: Verifying Checksum +41c8e5018d25: Download complete +f01974600e7b: Verifying Checksum +f01974600e7b: Download complete +ab6b448d4be9: Verifying Checksum +ab6b448d4be9: Download complete +774043ccc8cc: Pull complete +ab6b448d4be9: Pull complete +23a5bfa58353: Pull complete +862a57157567: Pull complete +41c8e5018d25: Pull complete +f01974600e7b: Pull complete +Digest: sha256:9806543112bea30967153610dc68c1c69db91d4718a94b7aa70e929113ab10f1 +Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-controller:0.299.0 +gitea.dooplex.hu/admin/felhom-controller:0.299.0 +[golden] asking the controller which infra images it manages … +[golden] baking infra images (4): traefik:v3.7.13 cloudflare/cloudflared:2026.9.3 gtstef/filebrowser:1.5.6-stable gitea.dooplex.hu/admin/felhom-samba:1.1.0 … +v3.7.13: Pulling from library/traefik +e2de96513ba9: Pulling fs layer +b686a4f73445: Pulling fs layer +78cb21c375ca: Pulling fs layer +acb2f33459b1: Pulling fs layer +acb2f33459b1: Waiting +e2de96513ba9: Verifying Checksum +e2de96513ba9: Download complete +acb2f33459b1: Verifying Checksum +acb2f33459b1: Download complete +b686a4f73445: Verifying Checksum +b686a4f73445: Download complete +e2de96513ba9: Pull complete +78cb21c375ca: Verifying Checksum +78cb21c375ca: Download complete +b686a4f73445: Pull complete +78cb21c375ca: Pull complete +acb2f33459b1: Pull complete +Digest: sha256:24841fe2de7304c149343d877d2923b4c8800a38ba015dea9174c23b20e344a0 +Status: Downloaded newer image for traefik:v3.7.13 +docker.io/library/traefik:v3.7.13 +2026.9.3: Pulling from cloudflare/cloudflared +2cc7ee286bf3: Pulling fs layer +c172f21841df: Pulling fs layer +218cf840d0d9: Pulling fs layer +f6069939f718: Pulling fs layer +d6b1b89eccac: Pulling fs layer +2780920e5dbf: Pulling fs layer +7c12895b777b: Pulling fs layer +3214acf345c0: Pulling fs layer +52630fc75a18: Pulling fs layer +dd64bf2dd177: Pulling fs layer +b839dfae01f6: Pulling fs layer +ebddc55facdc: Pulling fs layer +c4bc6f35ff5e: Pulling fs layer +b96fe2995f90: Pulling fs layer +58c0c263dc73: Pulling fs layer +bd8962e29291: Pulling fs layer +cac2ae0193cb: Pulling fs layer +f0383d5ebc47: Pulling fs layer +dd64bf2dd177: Waiting +b839dfae01f6: Waiting +ebddc55facdc: Waiting +c4bc6f35ff5e: Waiting +b96fe2995f90: Waiting +58c0c263dc73: Waiting +bd8962e29291: Waiting +f6069939f718: Waiting +d6b1b89eccac: Waiting +2780920e5dbf: Waiting +7c12895b777b: Waiting +3214acf345c0: Waiting +52630fc75a18: Waiting +cac2ae0193cb: Waiting +f0383d5ebc47: Waiting +2cc7ee286bf3: Verifying Checksum +2cc7ee286bf3: Download complete +c172f21841df: Download complete +218cf840d0d9: Verifying Checksum +218cf840d0d9: Download complete +f6069939f718: Verifying Checksum +f6069939f718: Download complete +d6b1b89eccac: Download complete +2cc7ee286bf3: Pull complete +2780920e5dbf: Verifying Checksum +2780920e5dbf: Download complete +7c12895b777b: Verifying Checksum +7c12895b777b: Download complete +3214acf345c0: Verifying Checksum +3214acf345c0: Download complete +52630fc75a18: Verifying Checksum +52630fc75a18: Download complete +dd64bf2dd177: Verifying Checksum +dd64bf2dd177: Download complete +b839dfae01f6: Verifying Checksum +b839dfae01f6: Download complete +ebddc55facdc: Verifying Checksum +ebddc55facdc: Download complete +c4bc6f35ff5e: Download complete +c172f21841df: Pull complete +bd8962e29291: Verifying Checksum +bd8962e29291: Download complete +b96fe2995f90: Verifying Checksum +b96fe2995f90: Download complete +58c0c263dc73: Verifying Checksum +58c0c263dc73: Download complete +cac2ae0193cb: Verifying Checksum +cac2ae0193cb: Download complete +218cf840d0d9: Pull complete +f0383d5ebc47: Verifying Checksum +f0383d5ebc47: Download complete +f6069939f718: Pull complete +d6b1b89eccac: Pull complete +2780920e5dbf: Pull complete +7c12895b777b: Pull complete +3214acf345c0: Pull complete +52630fc75a18: Pull complete +dd64bf2dd177: Pull complete +b839dfae01f6: Pull complete +ebddc55facdc: Pull complete +c4bc6f35ff5e: Pull complete +b96fe2995f90: Pull complete +58c0c263dc73: Pull complete +bd8962e29291: Pull complete +cac2ae0193cb: Pull complete +f0383d5ebc47: Pull complete +Digest: sha256:072c067d25ccbe61d46e18f0d0723255f2bb5304f7317caa95b27031520ff92c +Status: Downloaded newer image for cloudflare/cloudflared:2026.9.3 +docker.io/cloudflare/cloudflared:2026.9.3 +1.5.6-stable: Pulling from gtstef/filebrowser +55afa1ecc21d: Pulling fs layer +8ed8f35f8d4f: Pulling fs layer +989b226a579c: Pulling fs layer +660aeead31d5: Pulling fs layer +4f4fb700ef54: Pulling fs layer +adce24567e4c: Pulling fs layer +f17ea56b313b: Pulling fs layer +6b6f3b3efe88: Pulling fs layer +4ed1ca4f3fce: Pulling fs layer +e6fc9c6a5757: Pulling fs layer +d47782d1182a: Pulling fs layer +6b6f3b3efe88: Waiting +4ed1ca4f3fce: Waiting +e6fc9c6a5757: Waiting +d47782d1182a: Waiting +4f4fb700ef54: Waiting +adce24567e4c: Waiting +f17ea56b313b: Waiting +660aeead31d5: Waiting +55afa1ecc21d: Verifying Checksum +55afa1ecc21d: Download complete +660aeead31d5: Verifying Checksum +660aeead31d5: Download complete +4f4fb700ef54: Verifying Checksum +4f4fb700ef54: Download complete +989b226a579c: Verifying Checksum +989b226a579c: Download complete +f17ea56b313b: Verifying Checksum +f17ea56b313b: Download complete +8ed8f35f8d4f: Verifying Checksum +8ed8f35f8d4f: Download complete +6b6f3b3efe88: Verifying Checksum +6b6f3b3efe88: Download complete +adce24567e4c: Verifying Checksum +adce24567e4c: Download complete +55afa1ecc21d: Pull complete +4ed1ca4f3fce: Verifying Checksum +4ed1ca4f3fce: Download complete +d47782d1182a: Verifying Checksum +d47782d1182a: Download complete +e6fc9c6a5757: Verifying Checksum +e6fc9c6a5757: Download complete +8ed8f35f8d4f: Pull complete +989b226a579c: Pull complete +660aeead31d5: Pull complete +4f4fb700ef54: Pull complete +adce24567e4c: Pull complete +f17ea56b313b: Pull complete +6b6f3b3efe88: Pull complete +4ed1ca4f3fce: Pull complete +e6fc9c6a5757: Pull complete +d47782d1182a: Pull complete +Digest: sha256:7c5d7ac8ffda31294d278063cf9d2e04303b39e6dce1f4c691342240ca7703b8 +Status: Downloaded newer image for gtstef/filebrowser:1.5.6-stable +docker.io/gtstef/filebrowser:1.5.6-stable +1.1.0: Pulling from admin/felhom-samba +897d797d2723: Pulling fs layer +3051591aa250: Pulling fs layer +ce57a3f93416: Pulling fs layer +fb94eeec2fe1: Pulling fs layer +fb94eeec2fe1: Waiting +897d797d2723: Verifying Checksum +897d797d2723: Download complete +ce57a3f93416: Verifying Checksum +ce57a3f93416: Download complete +fb94eeec2fe1: Verifying Checksum +fb94eeec2fe1: Download complete +3051591aa250: Verifying Checksum +3051591aa250: Download complete +897d797d2723: Pull complete +3051591aa250: Pull complete +ce57a3f93416: Pull complete +fb94eeec2fe1: Pull complete +Digest: sha256:1c17c09422bec0366d7cf0e0fcfc1486ba6c90334a0a5d5c851073a9342f8f10 +Status: Downloaded newer image for gitea.dooplex.hu/admin/felhom-samba:1.1.0 +gitea.dooplex.hu/admin/felhom-samba:1.1.0 +[golden] baking the controller-bootstrap unit (deploys the BAKED controller from the config mount) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.service' → '/etc/systemd/system/felhom-controller-bootstrap.service'. +[golden] baking the controller-bootstrap PATH unit (starts the service on bootstrap-mount hot-plug — B1) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-controller-bootstrap.path' → '/etc/systemd/system/felhom-controller-bootstrap.path'. +[golden] baking the first-boot SSH host-key regeneration unit (F3) … +Created symlink '/etc/systemd/system/multi-user.target.wants/felhom-regen-hostkeys.service' → '/etc/systemd/system/felhom-regen-hostkeys.service'. +[golden] identity-clean + minimize … +[golden] stop + archive … +INFO: including mount point rootfs ('/') in backup +INFO: including mount point mp0 ('/var/lib/felhom') in backup +INFO: archive file size: 618MB +INFO: Finished Backup of VM 9100 (00:00:30) +[golden] DONE. golden archive volid: local:backup/vzdump-lxc-9100-2026_10_06-08_10_42.tar.zst (rootfs 32G + ONE data volume 24G @ /var/lib/felhom, all in the archive) +[golden] publishing golden (648498767 bytes, sha256 9948aa1682ada23a…) → https://gitea.dooplex.hu/api/packages/admin/generic/felhom-golden/0.299.0/golden.tar.zst +[golden] pre-delete existing: HTTP 404 (404/204 expected) +[golden] upload OK (HTTP 201) +GOLDEN_VERSION=0.299.0 +GOLDEN_SHA256=9948aa1682ada23afe5aa9fdb056c5b335f264823083b4f020bf4eb34dcd7de0 +[golden] Record in the hub operator UI (Configs → Day-0 artifacts): golden 0.299.0 / 9948aa1682ada23afe5aa9fdb056c5b335f264823083b4f020bf4eb34dcd7de0 +[golden] (the build guest 9100 is stopped; destroy it with: pct destroy 9100 --purge) diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index c0e93acd..7a6bf347 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -20,7 +20,7 @@ The test now uses UTC; red-proved (back to the local date: CASE 7 fails, `EXPIRED on 2026-10-06` wanted, rc=0 got). - `RUNBOOK-manual-build.md` §4.1: the bake's runner script names `GOLDEN_DOCKER_PKGS` (the 0.298.0 bake missed it once). -## felhom-host-install.sh 1.32.0 — the uninstall leaves nothing that holds a secret or a tunnel; pre-flight tells the truth (burn-down night, 2026-10-05; NOT published — ships with the tag `installer-v1.32.0`) +## felhom-host-install.sh 1.32.0 — the uninstall leaves nothing that holds a secret or a tunnel; pre-flight tells the truth (burn-down night, 2026-10-05; PUBLISHED 2026-10-06 05:52Z as `installer-v1.32.0`, operator ruling — `09` §3 decision 137) - **R-275:** the uninstall removes the agent's config directory with every copy of the config in it (the `.bak*` glob missed `agent.json.campaign9-before` & co.), removes the sudoers' dotted copies, and names the vmbr9 stanza and the ISO