From 7406ac7bbfcc243ecdf660d911a81056ec047ea0 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 06:34:26 +0200 Subject: [PATCH 01/13] =?UTF-8?q?audits:=20R-165=20Phase=200=20=E2=80=94?= =?UTF-8?q?=20P1=20and=20P2=20measured,=20nothing=20changed?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit P1 PASS: a pre-merge archive (mp0+mp1, confirmed from its own vzdump log) restore-tests clean on demo-hp with mount_parity ok in 84s. mountParity was not touched. Limit stated: run with the current agent because the merged one does not exist until after the STOP; the comparison is archive-vs-its-own- restore and never consults the host layout, so it carries provided Part 2 honours its constraint not to touch the restore path. Re-run after Part 2. P2: all three probed variants are mechanically clean — both paths writable, ONE df figure, dockerd 3/3 reboots, /mnt propagation, and a container's statfs(/) reporting the merged volume. The task's flagged ordering risk for V-b did not materialise. They are separated by SCOPING instead: V-a container sees /mnt = 8.0K but customer data sits inside Docker's data-root, so clearing /var/lib/docker destroys every local unit V-b container sees Docker's ENTIRE data-root under /mnt (17.9M on an empty box), making the bootstrap's own scoping comment false V-c neutral mount at /var/lib/felhom, both paths binds — breaks neither V-c was probed because the measurements showed each named variant violates a different documented invariant. It is offered as a measured option for the operator, NOT adopted. Teardown all three layers: 9401/9402/9403 destroyed, 5.19 GB returned, and the hub registers verified unchanged (5 customers, 4 hosts). --- .../audits/SPIKE-r165-phase0-2026-08-03.md | 183 ++++++++++++++++++ 1 file changed, 183 insertions(+) create mode 100644 documentation/audits/SPIKE-r165-phase0-2026-08-03.md diff --git a/documentation/audits/SPIKE-r165-phase0-2026-08-03.md b/documentation/audits/SPIKE-r165-phase0-2026-08-03.md new file mode 100644 index 00000000..b1a224f1 --- /dev/null +++ b/documentation/audits/SPIKE-r165-phase0-2026-08-03.md @@ -0,0 +1,183 @@ +# SPIKE R-165 Phase 0 — the two probes the merge spike left unmeasured + +**Date:** 2026-08-03 · **Author:** Claude Code · **Status:** MEASURED — **no layout changed** + +`SPIKE-r165-mp1-merge-2026-08-02.md` named two things as unmeasured and both are load-bearing. This +document measures them. **It changes nothing**: no golden rebuilt, no box reinstalled, no guest config +edited outside the throwaway probe guests, which are destroyed at the end. + +**Host: `felhom-pve` (N100) — Tier 0.** demo-hp would have been the default per the 2026-07-25 ruling, +but `target-selection.md` records its `local-lvm` as an **over-subscribed thin pool backing live guest +9201** (~144 GiB allocated over ~54 GiB) where filling it corrupts every guest, and it holds **no +container template**. felhom-pve has the template and 258 GiB free at 29% pool usage. Both are Tier 0; +this picked the one where a probe cannot damage a live guest. **Neither DooPlex, `ep0`, nor the +colleague's box was contacted at any point.** + +--- + +## P1 — does a PRE-merge archive restore cleanly into the merged world? + +**Verdict: PASS.** + +### Method + +The spike reasoned from `mountParity` that it should pass and said plainly that this had never been +executed. It is now executed, on real hardware, with the real restore-test path — not a hand-assembled +restore. + +- **Archive:** `local:backup/vzdump-lxc-9201-2026_07_28-17_43_05.tar.zst` from **demo-hp**, 1.68 GB. + Confirmed pre-merge by its own vzdump log rather than by assumption: + + ``` + including mount point rootfs ('/') in backup + including mount point mp0 ('/var/lib/docker') in backup + including mount point mp1 ('/mnt/sys_drive') in backup + ``` + +- **Command:** `felhom-agent --selftest=restore-test -archive ` on demo-hp. + +### Measured result + +``` +"pass": true, +"verified": "boot+running", +"mount_parity": "ok", +"duration_seconds": 84.2, +"mount_inventory": [ + "mp0=/var/lib/docker (50G)", + "mp1=/mnt/sys_drive (20G)", + "mp8=/mnt/felhom-drives (throwaway for the archived bind)", + "mp9=/etc/felhom-bootstrap (throwaway for the archived bind)" +] +=== selftest=restore-test OK (scratch 990000 restored+booted+verified+torn-down in 1m24s) === +``` + +`mountParity` was **not** relaxed, weakened or touched in any way. + +### The limit of this result, stated rather than glossed + +**It was run with the CURRENT agent (v0.119.0), because the merged agent does not exist yet** — Part 2 +sits after this session's STOP. What it proves is that `mountParity` compares the **archive** against +**its own restore**, so a pre-merge archive recreates its own `mp0 + mp1` in the scratch guest and the +two agree. That comparison never consults the *host's* golden layout, which is why the merge cannot +invalidate it — **provided Part 2 honours its own constraint not to touch `mountParity` or the restore +path** (§5, §12 of the task). It is an 84-second command and **should be re-run once the merged agent +exists**, which is cheap and turns a sound inference into an observation. + +--- + +## P2 — which S1 variant actually works on this platform? + +**Verdict: all three probed variants are mechanically clean. They are separated by SCOPING, not by +mechanics — and the deciding fact was not in the task's table.** + +### Method + +A throwaway unprivileged LXC per variant (`nesting=1,keyctl=1`, rootfs 8 G + one 10 G volume, +`backup=1`), Docker installed from the same repo with the **same `daemon.json` the golden bakes** +(`containerd-snapshotter: false`, overlay2, log caps). Per variant, measured at first boot and after +**each of three reboots**: + +1. both `/var/lib/docker` and `/mnt/sys_drive` present and **writable** (write → read back → delete, + a positive observable rather than an `ls`); +2. **one** filesystem — same source device **and** the same free-space figure for both paths; +3. `dockerd` active and `docker run hello-world` succeeding; +4. then once: the **real bootstrap propagation sequence** (`mount --rbind /mnt /mnt`, + `mount --make-rshared /mnt`, `docker run -v /mnt:/mnt:rslave`); +5. and: what a container mounting `/mnt:rslave` **actually sees** — the scoping check. + +No pipe hides a non-zero; every check reports its own rc and the script aborts with `MEASURED-FAIL`. + +### The variants + +| | volume mounted at | then | +|---|---|---| +| **V-a** | `/var/lib/docker` | `/mnt/sys_drive` = bind of `/var/lib/docker/sys_drive` | +| **V-b** | `/mnt/sys_drive` | `/var/lib/docker` = bind of `/mnt/sys_drive/docker` | +| **V-c** | `/var/lib/felhom` *(neutral)* | **both** consumer paths are binds of subdirectories | + +**V-c was not in the task's table.** It was probed *because* the measurements below showed V-a and V-b +each violate a different documented invariant, and V-c is the shape that violates neither. It is +offered as a **measured option for the operator at the STOP**, not adopted — the task is explicit that +a variant is not chosen mid-session. + +### Measured results + +| check | V-a | V-b | V-c | +|---|---|---|---| +| both paths present + writable | **yes** | **yes** | **yes** | +| ONE filesystem, ONE free-space figure | **yes** | **yes** | **yes** | +| dockerd active + `docker run` — initial | **yes** | **yes** | **yes** | +| dockerd active + `docker run` — reboots **1/2/3** | **3/3** | **3/3** | **3/3** | +| `/mnt` propagation `shared`; container sees `/mnt` | **yes** | **yes** | **yes** | +| both paths still real mountpoints (`findmnt` non-empty — the form the golden's assertions use) | **yes** | **yes** | **yes** | +| container `statfs("/")` reports the merged volume | **yes** (10218772 KiB) | **yes** | **yes** | +| **what a container mounting `/mnt:rslave` SEES** | `sys_drive` only — **8.0K** | `sys_drive` **+ `sys_drive/docker`** — **17.9M** | `sys_drive` only — **8.0K** | +| customer data inside Docker's data-root | **YES** | no | no | + +**The mechanical worry was misplaced.** The task flagged V-b's ordering risk — `/var/lib/docker` must +be bound before dockerd starts. An `/etc/fstab` bind is ordered by `local-fs.target`, which precedes +`basic.target` and therefore `docker.service`, and it held **3 reboots out of 3**. Ordering is not what +separates these variants. + +### What actually separates them + +**V-b breaks a documented scoping invariant, and the measurement is the proof.** The controller +container is started with `-v /mnt:/mnt:rslave`, and the bootstrap script's own comment states the +scope it relies on: + +> *"scoped to /mnt, which (Model A) holds only Felhom's felhom-data-namespace mounts, never the +> customer's other on-drive data"* + +Under V-b the container sees `/mnt/sys_drive/docker` — **Docker's entire data-root**, 17.9 MB on an +empty probe box and growing with every image and every app volume. That sentence becomes false. The +controller already holds the Docker socket, so this is **not a capability escalation** — but it puts +Docker's internal tree inside the one path the controller's own scanners, the FileBrowser surface and +the data-migration engine (which works *"in-process over the controller's `/mnt:/mnt:rslave` RW +mount"*) treat as Felhom-only. + +**V-a breaks the other one:** customer backups live inside Docker's data-root, so `du` on the data-root +stops meaning what it says, and the ordinary operator reflex for a sick Docker — clear `/var/lib/docker` +— destroys every local recovery unit on the box. + +**V-c breaks neither**, at the cost of one new mount path and two fstab lines instead of one. + +--- + +## P3 — the golden's four assertions + +**Not yet run — it belongs to Part 1, which is after the STOP.** Recorded here so it is not lost: +`build-golden.sh` fails closed on the split in **four** places (`:126`, `:130` separate-mount asserts; +`:315`, `:319` vzdump-exclusion guards), and each retargeted guard must be shown to **abort** against a +deliberately wrong shape before the golden is trusted. + +**One measurement already de-risks it:** under all three variants both `/var/lib/docker` and +`/mnt/sys_drive` remain **real mountpoints**, so `findmnt -no SOURCE,FSTYPE | grep -q .` — the +exact form the two existing assertions use — still returns non-empty. The assertions can be +**retargeted with a changed message and an added guard for the single volume**, rather than rewritten +from scratch. They must not be deleted. + +--- + +## Teardown + +| layer | action | evidence | +|---|---|---| +| **the machine** | probe guests `9401`, `9402`, `9403` destroyed; P1's scratch `990000` was torn down by the restore-test itself | `pct list` shows only `9201` | +| **the host** | `local-lvm` **112398205 → 107204406 KiB** used (30.73% → 29.31%) — **5.19 GB actually returned**, not merely deallocated | `pvesm status` before/after | +| **the hub** | **none created, and verified rather than assumed.** No probe claimed a box, minted a customer or registered an appliance — none ever ran a controller. The registers hold the same **5 customers** (`david`, `demo-felhom`, `demo-hp`, `peti-felhom`, `sess-f`) and same **4 hosts** (`demo-felhom-8363b5`, `demo-hp-bb76ea`, `drill-r50-0a4f9a`, `sess-f-2670b5`) as before Phase 0 | hub `/` + `/hosts` | + +Scratch scripts and logs removed from `felhom-pve` (`/root/p2-probe.sh`, `/tmp/p2-*.log`, +`/tmp/p1-restoretest.log`). `pct list` on both demo hosts shows only their own `9201`. + +--- + +## What this changes about the merge + +1. **P1 removes the restore risk from the decision.** A pre-merge archive restores clean with parity + ok; it should be re-confirmed against the merged agent, which is one command. +2. **The variant question is not "will it boot" — it is "which invariant do we break".** Both named + variants work perfectly and each violates one documented guarantee. That is the operator's call and + is the subject of this session's STOP. +3. **V-c exists and is measured.** It costs one extra mount path and one extra fstab line, and is the + only probed shape that keeps both guarantees. From e3525e62ac639b57eb19378d50695914885504ec Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 06:43:58 +0200 Subject: [PATCH 02/13] host-install: one data volume, derived from the disk (R-165) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit felhom-agent v0.120.0 merges the two data volumes into one, and step_grows computed two numbers while the install call passed both — so this had to change with the agent or every install would have provisioned a half-sized box. The 80/20 split is summed (226 = 184+42), so a standard appliance keeps exactly the 250 G it had, no longer split by a wall. The size still comes from the physical disk: step_grows already read the thin pool's free space, and the merge only collapsed its two outputs into one. --sysdata-grow is deprecated but still honoured, because the agent folds a hand-passed value in rather than dropping it. --- scripts/CHANGELOG.md | 16 ++++++++++++++++ scripts/felhom-host-install.sh | 29 ++++++++++++++++++++--------- 2 files changed, 36 insertions(+), 9 deletions(-) diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index 8932ef5c..b6163872 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,3 +1,19 @@ +## host-install: one data volume, derived from the disk (2026-08-03, R-165) + +**Forced by a census, not planned.** `felhom-agent` v0.120.0 merges the appliance's two data volumes +into one (decision D-a). `step_grows` computed **two** numbers and the install call passed both, so +this script had to change with the agent or every install would have provisioned a half-sized box. + +- **`step_grows` computes ONE total.** The old 80/20 docker-vs-sysdata split is summed: `226` where it + was `184 + 42`, `106` where it was `84 + 22`, `46` where it was `34 + 12`. **A standard appliance + keeps exactly the capacity it had — 250 G — it is simply no longer split by a wall.** +- **The size still comes from the physical disk.** `step_grows` already read the thin pool's real free + space (`lvs /dev/pve/data`); the merge only collapsed its two outputs into one. This is what makes + the merge safe to ship: an unflagged install does **not** get the golden's 24 G base. +- **`--sysdata-grow` is DEPRECATED but still honoured.** It is no longer auto-computed (set to 0), and + a hand-passed value still counts because the agent **folds** it into the single volume's grow rather + than dropping it — so an operator reproducing an old command line gets the same total. + ## CI — a Gitea Actions runner, and a red run that reaches a person (2026-08-02, R-168) **No version bump anywhere: nothing in the product repos is compiled, built or deployed by this.** diff --git a/scripts/felhom-host-install.sh b/scripts/felhom-host-install.sh index f7a78b59..3aed3872 100644 --- a/scripts/felhom-host-install.sh +++ b/scripts/felhom-host-install.sh @@ -115,8 +115,8 @@ # (default: appliance → island 169.254.253.1:8443; byo → vmbr0 IP:8443) # --no-island appliance only: keep the historical LAN bind instead of the R-50 island # --rootfs-grow N grow OS rootfs by N GiB (default: auto-compute) -# --datavol-grow N grow Docker-data vol by N GiB (default: auto-compute) -# --sysdata-grow N grow user-data vol by N GiB (default: auto-compute) +# --datavol-grow N grow the single data volume by N GiB (default: auto-compute from the pool) +# --sysdata-grow N DEPRECATED (R-165): added to --datavol-grow; there is one volume now # # Guest cap (appliance: optional — protect a SHARED host's other guests; byo: BOTH REQUIRED — # the only noisy-neighbor protection on a host you do not own; needs agent >= v0.52.0): @@ -1772,23 +1772,34 @@ step_token() { #------------------------------------------------------------------------------- step_grows() { log_step "3/8 compute volume grows" - # Golden base: rootfs 32G + Docker-data 16G + user-data 8G (build-golden.sh). + # Golden base since build-golden.sh v3.0.0 (R-165): rootfs 32G + ONE data volume 24G. The separate + # 8G user-data volume was MERGED AWAY — one volume, one free-space figure, no ceiling — so there is + # one number to compute here instead of two. + # + # THE SIZE IS DERIVED FROM THE PHYSICAL DISK, which is what makes the merge safe to ship: an + # unflagged install does NOT get the golden's 24G, it gets a share of the thin pool's real free + # space. (Before R-165 this same block already did the deriving; the merge only collapsed its + # 80/20 docker-vs-sysdata split into a single total.) if [[ -z "$ROOTFS_GROW$DATAVOL_GROW$SYSDATA_GROW" ]]; then local free_gib free_gib=$(lvs --noheadings --units g -o lv_size,data_percent /dev/pve/data 2>/dev/null | awk '{gsub(/[^0-9.]/,"",$1); used=$2; print int($1*(100-used)/100)}' 2>/dev/null || echo 0) - # Reserve headroom; split the rest ~ docker 80% / sysdata 20%; rootfs stays golden. + # Reserve headroom; the totals below are the pre-merge pair SUMMED, so an appliance gets the + # same capacity it did before — it is simply no longer split by a wall. ROOTFS_GROW=0 if [[ "${free_gib:-0}" -ge 300 ]]; then - DATAVOL_GROW=184; SYSDATA_GROW=42 # reproduces the standard 200G/50G appliance + DATAVOL_GROW=226 # 184+42 -> the standard 250G appliance (was 200G+50G) elif [[ "${free_gib:-0}" -ge 150 ]]; then - DATAVOL_GROW=84; SYSDATA_GROW=22 + DATAVOL_GROW=106 # 84+22 else - DATAVOL_GROW=34; SYSDATA_GROW=12 # minimal floors + DATAVOL_GROW=46 # 34+12 — minimal floor fi - log_info " auto-computed from ~${free_gib} GiB free" + SYSDATA_GROW=0 + log_info " auto-computed from ~${free_gib} GiB free (ONE volume since R-165)" fi ROOTFS_GROW="${ROOTFS_GROW:-0}"; DATAVOL_GROW="${DATAVOL_GROW:-0}"; SYSDATA_GROW="${SYSDATA_GROW:-0}" - log_info " grows: rootfs +${ROOTFS_GROW}G (->$((32+ROOTFS_GROW))G), docker +${DATAVOL_GROW}G (->$((16+DATAVOL_GROW))G), sys_drive +${SYSDATA_GROW}G (->$((8+SYSDATA_GROW))G)" + # A hand-passed --sysdata-grow is still ACCEPTED and still counts: the agent folds it into the one + # volume (bringup.go 4b), so an operator reproducing an old command line gets the same total. + log_info " grows: rootfs +${ROOTFS_GROW}G (->$((32+ROOTFS_GROW))G), data +$((DATAVOL_GROW+SYSDATA_GROW))G (->$((24+DATAVOL_GROW+SYSDATA_GROW))G, ONE volume)" _state_mark grows } From 14d8c0078183845c66e266b365c3a4264fddf05a Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 07:16:15 +0200 Subject: [PATCH 03/13] docs: R-165 merge built and proven at the bake; R-163 + R-175 closed, R-178 filed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit 07-backup-architecture.md gains §7.5.1 (S-1: the contract changed in the same session): the ceiling §7.5 describes no longer exists for a box built from golden >= 0.192.0, the bulkhead's replacement is recorded, and R-175 is FIXED here rather than left standing — the bound is restated as a function of mp1 and scoped to split-layout boxes, naming all three real shapes. Capability map: new row as IMPLEMENTED, deliberately NOT proven-live, with the missing leg named — no box has been reinstalled from the golden, and "the golden baked" is not "a box built from it works". R-163 CLOSED: the ceiling it recorded stops existing. R-176(a) answered by P1; (b) WITHDRAWN, since every node is reinstalled rather than migrated. R-178 filed for the reinstalls, which were not done this session. CONTEXT S-13 (the variant chosen on measurement; pruning rejected with its reason) and S-14 (prove first, then vouch — the golden is published but deliberately unvouched, because vouching is what makes a fresh install pick up a layout no box has been proven from). STATUS: plain-language section; both operator questions now answered, so the waiting-on-you item is cleared. Two older entries trimmed so the page did not grow. --- CONTEXT.md | 29 +++ REPORT.md | 186 ++++++------------ STATUS.md | 62 +++--- .../architecture/00-capability-map.md | 1 + .../architecture/07-backup-architecture.md | 37 +++- documentation/backlog/OPEN-ITEMS.md | 9 +- 6 files changed, 163 insertions(+), 161 deletions(-) diff --git a/CONTEXT.md b/CONTEXT.md index ddeef93b..b915b46a 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -17,6 +17,35 @@ ## Standing rulings +**S-13 — the `mp1` merge landed, and the variant was chosen on measurement (2026-08-03, R-165 / D-a).** +The appliance's two data volumes are one. **Variant V-c**: the volume mounts at the NEUTRAL path +`/var/lib/felhom`, and both `/var/lib/docker` and `/mnt/sys_drive` are binds of subdirectories of it. +**Three shapes were built and rebooted before choosing** (`audits/SPIKE-r165-phase0-2026-08-03.md`) — +all three boot, reboot 3/3, give ONE `df` figure and keep a container's `statfs("/")` on the merged +volume, so **the ordering risk that motivated the probe was not what mattered.** They differ only in +which documented guarantee they break: volume-at-`/var/lib/docker` puts customer backups inside +Docker's data-root, so the ordinary "clear `/var/lib/docker`" reflex destroys every local unit; +volume-at-`/mnt/sys_drive` puts Docker's entire data-root under `/mnt`, which the controller container +mounts wholesale — **measured: it then sees `/mnt/sys_drive/docker`**, falsifying the bootstrap's own +comment that `/mnt` holds only Felhom's namespace mounts. **V-c breaks neither**, for one extra path. + +**B2 is the bulkhead replacement, and "or prune the oldest" is REJECTED with its reason**, because the +question will be asked again: nothing on that filesystem is generational — a unit is ONE fixed path per +app (`backups/primary/`) refreshed in place, and a DB dump is `-.sql`, also fixed — +so pruning could only mean deleting a **different** app's only local recovery unit. +`pruneStalePrimaryDirs` is an ORPHAN sweep with no notion of age and must never be repurposed. + +**No migration exists, and that is a ruling not an omission:** every node is REINSTALLED. Both demo +boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed shortly. +So R-176's in-place migration rehearsal is **withdrawn**, not deferred. + +**S-14 — prove first, then vouch (2026-08-03).** Golden **0.192.0** is baked, published and verified in +the registry, and is **deliberately UNVOUCHED**. Vouching is what makes a fresh install pick a golden +up, so vouching one that no box has been proven from would put an unproven disk layout in front of the +next install anywhere. The bake script already treats the hub record as a separate deliberate step; this +makes the ordering a rule. **The remaining work is R-178** — reinstall each demo box from the merged +golden, prove claim → deploy → back up → restore, **one box at a time**, and only then vouch. + **S-11 — D-c's routing, and why R-158's own proposal was overruled (2026-08-02, R-167 SHIPPED).** Decision D-c splits two signals by AUDIENCE, and the split is the ruling: **a fill warning is the CUSTOMER's** (they can free space, delete files, add a drive) and **a per-app backup capture failure diff --git a/REPORT.md b/REPORT.md index 36a86e3b..2545840f 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,147 +1,83 @@ -# REPORT — hub v0.89.0: the two halves of decision D-c, plus the R-165 merge spike (2026-08-02) +# REPORT — R-165: the `mp1` merge, built and proven at the bake (2026-08-03) -**Overwritten** per the standing rule. The prior contents (R-168, the CI runner, same day) have their -durable record in `scripts/CHANGELOG.md` and `CONTEXT.md` S-8/S-9/S-10. +**Overwritten** per the standing rule. The prior contents (hub v0.89.0 / R-167, 2026-08-02) have their +durable record in `hub/CHANGELOG.md` and `CONTEXT.md` S-11/S-12. -**Companion report:** `felhom-controller/REPORT.md` holds the controller side (v0.191.0/.1/.2), the -full red-proof table, the Hungarian copy, and the live evidence for all three flows. This file covers -the hub change, the documentation coupling, and **Part 3's spike**. +**Companion:** `felhom-agent/REPORT.md` holds the full session detail — probes, variant evidence, the +bake transcript, red-proofs and teardown. **This file covers what changed in THIS repo, and the CI +verification for all three.** + +> **Scope, stated first.** The merge is **built and green through Phase 5**. **Phases 6–7 — reinstalling +> the two demo boxes from the merged golden and proving one end to end — were NOT done, and nothing +> was wiped.** The golden is therefore deliberately **unvouched**. Remaining work: **R-178**. --- -## 1. Baseline drift — recorded, because the task's §1 was wrong +## 1. What changed in `felhom.eu` -The task targeted hub **v0.87.0 → v0.88.0**. On arrival `main` was at `8ef92a3f` with hub **v0.88.0 -already shipped** (R-172, the WAL fix), not `d5774d318941`/v0.87.0. Target corrected to **v0.89.0**. -Highest register ID in use was **R-173**, not R-171. +**No hub change; no hub version bump.** This repo changed in two places, one of them forced. -## 2. Hub change (v0.89.0) +**`scripts/felhom-host-install.sh` — forced by a census, not planned.** `step_grows` computed **two** +volume sizes and the install call passed both, so it had to change with the agent or every install +would have provisioned a half-sized box. It now computes ONE total, summing the old 80/20 split +(`226` = the previous `184 + 42`), so **a standard appliance keeps exactly the 250 G it had** — it is +simply no longer split by a wall. **The size still comes from the physical disk**: `step_grows` +already read the thin pool's real free space, and the merge only collapsed its two outputs into one. +That is the answer to the task's §8.2 — no row needed filing. `--sysdata-grow` is deprecated but still +honoured, because the agent **folds** a hand-passed value in rather than dropping it. -**One new event type, not two.** The task called for a new customer-facing type *and* a new -operator-only one. Reconnaissance found `disk_warning`/`disk_critical` already allowlisted here, with -Hungarian `customerMessages`, in the controller's `DefaultEnabledEvents` and behind a UI checkbox — -**and with no producer in any repo.** The operator chose to wire that inert pair rather than mint a -near-duplicate, so only the operator type is new. +**Documentation**, per the coupling rule — see §3. -| Change | File | Why | -|---|---|---| -| `+ "recovery_unit_capture_failed"` | `internal/api/handler.go` (`allowedEventTypes`) | without it the controller's POST 400s and the event vanishes | -| `+ "recovery_unit_capture_failed"` | `internal/notify/dispatcher.go` (`operatorOnlyEvents`) | **this** is what makes it operator-only; the allowlist does not, and v0.78.0 claimed otherwise and shipped the defect | -| `- customerMessages["disk_warning"]`, `- ["disk_critical"]` | `internal/notify/templates.go` | `FormatCustomerEmail` PREFERS the entry over the message, so a static template would discard the drive label and the free-space figures the controller now sends. Same reason `offbox_enlarge_blocked` and `disk_health_degraded` have none | -| `+ func IsOperatorOnly` | `internal/notify/dispatcher.go` | lets the `api` package pin BOTH registers in ONE test; checked separately, allowlisted-but-not-operator-only is invisible. Read-only — the register stays unexported so nothing can widen it at runtime | -| `REUSE.md` §5 "new event type" rewritten | `REUSE.md` | it told readers to always add a `customerMessages` entry, which is **wrong** for operator-only types and **harmful** for dynamic-message ones | +## 2. The bake, as this repo's audit trail records it -**Tests 574 → 579**, full suite green (`go build ./... && go vet ./... && go test ./...`), all five -`repo_gates.py` gates OK. +`documentation/audits/SPIKE-r165-phase0-2026-08-03.md` (new) holds P1, P2 and P3 with method, +measurement and ruling, including the teardown of every probe artefact at all three layers. The +headline measurements: -**Red-proof (Scenario G), demonstrated not argued:** removing `recovery_unit_capture_failed` from -`operatorOnlyEvents` fails two tests, one reading *"a customer was emailed the OPERATOR-ONLY -recovery_unit_capture_failed (customer@example.com)"*. The dispatch test runs under the **breaking** -configuration — the customer has the event enabled and an email set — because that is the only -configuration in which the missing entry is visible. +- **P1 PASS** — a real pre-merge archive restore-tests clean, `mount_parity: ok`, 84 s, `mountParity` + untouched. +- **P2** — all three variants boot and reboot 3/3; they are separated by **scoping**, not mechanics. + The container's view of `/mnt` is 8.0K under V-a and V-c, and **17.9M — Docker's entire data-root — + under V-b**. The operator chose **V-c**. +- **P3** — the four retargeted golden assertions, run against a deliberately wrong shape: **8/8**. -**Live (guest 9201 → hub):** both event types accepted and stored; `operator | sent`; and the positive -observable `customer | recovery_unit_capture_failed | skipped | operator_only` read from -`notification_log`. The customer half: `customer | disk_warning | sent` and `customer | disk_critical -| sent` with the dynamic Hungarian intact. - -**Deploy:** GitOps only — `manifests/hub.yaml` bumped 0.88.0 → 0.89.0 (`6d359a5`), pushed, then a -deliberate ArgoCD hard-refresh + sync. Never `kubectl set image`. App `felhom` **Synced / Healthy**, -`deploy/hub` rolled out, running `gitea.dooplex.hu/admin/felhom-hub:0.89.0`, startup log clean. - -## 3. Part 3 — the R-165 spike. **M1-M5 each answered; nothing was changed.** - -Full document: `documentation/audits/SPIKE-r165-mp1-merge-2026-08-02.md`. No partition was created, -resized, moved or deleted; no golden rebuilt; no guest config edited. `ep0` and Peti's box were not -contacted (D-d, `runbooks/target-selection.md`). - -**M1 — what is actually there. ANSWERED, and it contradicts the architecture doc.** - -| | demo-felhom | demo-hp | golden default | -|---|---|---|---| -| `mp0` `/var/lib/docker` | **200 G** (13 G used) | **50 G** (5.4 G used) | 16 G | -| `mp1` `/mnt/sys_drive` | **50 G** (2.0 G used, 5%) | **20 G** (92 M used, 1%) | 8 G | - -§7.5 documents the appliance as `mp0 50G / mp1 20G` — that is demo-hp exactly and **not** demo-felhom. -Any merge plan expressed as a fixed pair is already wrong for one of the two boxes that exist. §7.5's -headline bound (*"≈ 19 GB … ≈ 10 GB"*) is derived from `mp1 = 20 G` and is therefore one box's, not -the fleet's → **R-175**, filed and §7.5 annotated in this session. - -**M2 — what lives on `mp1`. ANSWERED, and it is not only backups.** Four things would move: -Tier-1 units of driveless apps (269 M, ~30 apps on demo-felhom), **Tier-2 mirrors (1.7 G — i.e. the -MAJORITY is Tier 2, not Tier 1)**, the `userdata/import` drop zone which lives on the system drive by -**contract** (R-75), and the system-data userdata namespace. Observed fill is 5% / 1%: the constraint -is a **ceiling** problem, not a current-fill one. - -**M3 — which merge shapes exist. ANSWERED for three shapes, with ONE item explicitly unmeasured.** -The golden **fails closed on the split in four places**, not one (`build-golden.sh:126,130` -separate-mount asserts + `:315,319` vzdump-exclusion guards). The archive scope `rootfs+mp0+mp1` stays -complete after a merge (the data moves onto `mp0`). `mountParity` holds for new archives. **Unmeasured -and reported as such:** whether a *pre-merge* archive restore-tests into a *merged* guest — reading -`mountParity` says it should, but that is reasoning from source about an unvalidated mechanism, which -this project has got wrong four times → **R-176**. - -**M4 — the bulkhead. ANSWERED, and it is the important one.** `mp1` is not only a ceiling: today an -overflow is refused per app with the last good unit byte-identical **and cannot reach -`/var/lib/docker`**. After the merge it can, and a full Docker data-root is a stopped box, not a slow -one. Four replacements costed — a reserved block percentage, **a refusal threshold in the capture -path**, a project quota, or deeming R-167's warnings sufficient — with the trade-off of each. -**Deliberately not chosen: this is the operator's ruling.** - -**M5 — existing boxes. ANSWERED for the measurable population; one part honestly UNMEASURED.** The -hub's `/hosts` register holds four hosts, **two ONLINE**, both demo boxes — and both are **Tier 0, -therefore reinstallable rather than migratable (D-d)**, so migration cost for the measurable population -is **zero**. **D-a's condition (1) — "before any external install" — is currently SATISFIED**, which -makes this the cheapest this decision will ever be. **`peti-felhom` exists as a customer with NO host -in the register**, so its layout is not knowable from the hub and the box was not contacted; whether it -needs converting or reinstalling is the operator's information. The in-place migration procedure has -**never been rehearsed**, so "is the box restorable at every point of it?" is currently unknown → also -**R-176**. - -**Ranked options and recommendation:** (1) **S1 — one volume with the two paths as directories — plus -B2, a refusal threshold in the capture path**, shipped as a fresh-install shape with the demo boxes -reinstalled; (2) S1 + warnings only; (3) S3, grow `mp1` and keep the split (D-a's rejected baseline, -measured for comparison); (4) S2, two mounts on one pool — **not recommended at all**, it satisfies -every assertion while delivering none of the benefit and converts a clean per-app refusal into a -shared-pool exhaustion neither `df` can see coming. - -**STOPPED at the operator's question**, per the task. The merge is next session's supervised work. - -## 4. Documentation coupling +## 3. Documentation coupling | File | Change | |---|---| -| `documentation/backlog/OPEN-ITEMS.md` | **R-158** closed (by R-167 — *no second row for the same wire*); **R-167** closed; **R-165** updated with M1-M5 + the operator question, stays open; **4 new rows** R-174/175/176/177 | -| `documentation/backlog/ROADMAP.md` | R-158 collapsed to a shipped one-liner; R-167 added as shipped; R-165 added as spiked/waiting-on-operator | -| `documentation/architecture/00-capability-map.md` | **two new rows**, both **PROVEN-LIVE** with live citations | -| `documentation/architecture/07-backup-architecture.md` §7.5 | **S-1: the contract changed in the same session.** The section's closing claim *"nothing warns when an app crosses the line"* is now false; the alerting is written in, and the one-box-vs-fleet caveat added | -| `CONTEXT.md` | **S-11** (D-c's routing, and why R-158's own `backup_failed` proposal was overruled) and **S-12** (the monitoring landed *before* the merge, not with it) | -| `STATUS.md` | new plain-language section; the merge decision added to *Waiting on you*; **two older entries trimmed** so the page did not grow — one screen, per its own rule | -| `REUSE.md` | the "new event type" extension point rewritten (see §2) | +| `documentation/architecture/07-backup-architecture.md` | **S-1: the contract changed in the same session.** New **§7.5.1** — the ceiling §7.5 describes no longer exists for a box built from golden ≥ 0.192.0, the bulkhead's replacement (B2) is recorded, and **R-175 is FIXED here**: the bound is restated as a function of `mp1` and scoped to split-layout boxes, naming all three real shapes | +| `documentation/architecture/00-capability-map.md` | new row — **IMPLEMENTED, not PROVEN-LIVE**, with the missing leg named (no box reinstalled, R-178) and the bake cited as the evidence it is | +| `documentation/backlog/OPEN-ITEMS.md` | **R-165** → shipped-not-yet-proven-live; **R-163 CLOSED** (its ceiling no longer exists); **R-175 CLOSED**; **R-176** (a) answered by P1, (b) **withdrawn** — every node is reinstalled, not migrated; **R-178 NEW** | +| `CONTEXT.md` | **S-13** (the variant, chosen on measurement; B2's floor; pruning rejected with its reason; no migration exists) and **S-14** (prove first, then vouch) | +| `STATUS.md` | rewritten section in plain language; the *Waiting on you* item cleared — both questions are answered; **two older entries trimmed** so the page did not grow | +| `scripts/CHANGELOG.md` | the host-install change, with why it was forced | -## 5. Register IDs +## 4. CI — run ids and conclusions -**Opened:** R-174, R-175, R-176, R-177. Each established free by -`grep -ro "R-17n\b" documentation/ *.md` → **0 hits**, run before minting. -**Closed:** R-158, R-167, R-174. **Updated, still open:** R-165, R-163 (unchanged — it stays the -record of the constraint until the merge lands). - -## 6. CI — run ids and conclusions - -Checked by PULL from `…/actions/tasks`, matching `head_sha` to each commit — CI emails only on -failure, so a green that was never looked at is an assumption, not an observation. **Every commit this -session, both repos, is green.** +Checked by **PULL**, matching `head_sha` to each commit — CI mails only on failure, so an unchecked +green is an assumption. | Repo | Commit | Task id | Run # | Conclusion | |---|---|---|---|---| -| `felhom-controller` | `cf48214` (v0.191.0) | 31 | 11 | **success** | -| `felhom-controller` | `5adae4d` (v0.191.1) | 34 | 12 | **success** | -| `felhom-controller` | `9a3c485` (v0.191.2) | 35 | 13 | **success** | -| `felhom.eu` | `179dd79` (hub v0.89.0) | 32 | 17 | **success** | -| `felhom.eu` | `6d359a5` (manifest 0.89.0) | 33 | 18 | **success** | -| `felhom.eu` | `41dbecb` (docs) | 36 | 19 | **success** | +| `felhom-controller` | `4be6467` (v0.192.0) | 39 | 15 | **success** | +| `felhom-agent` | `cd6e267` (v0.120.0) | — | — | *see below* | +| `felhom-agent` | `4bb84fc` (REPORT) | — | — | *see below* | +| `felhom.eu` | `7406ac7` (phase-0 audit) | 40 | 21 | **success** | +| `felhom.eu` | `e3525e6` (host-install) | — | — | *see below* | -## 7. `--no-verify` +*(The remaining rows are filled in from `…/actions/tasks` after the final push; any that had not +finished at write time are listed with their status rather than assumed green.)* -**Not used anywhere.** Every push in this session ran `.githooks/pre-push` (`repo_gates.py --fast` / -`controller_gates.py --fast`) and passed. +## 5. `--no-verify` + +**Not used anywhere.** Every push in this session ran its repo's `.githooks/pre-push` and passed. + +## 6. What remains — R-178 + +1. Reinstall **demo-hp** from golden 0.192.0 through the real installer path; prove claim → deploy an + app → back up → restore; show `df` proving one filesystem and a recovery unit landing on it. +2. Only then reinstall **demo-felhom** (it carries the PBS-DR/offsite tier, so it is the box whose + backup chain a reinstall actually disturbs). +3. **Then** vouch golden 0.192.0 in the hub, and flip the capability-map row to PROVEN-LIVE. +4. Re-run P1's restore-test against agent v0.120.0 — one command, turns a sound inference into an + observation. diff --git a/STATUS.md b/STATUS.md index b2669430..4c97d561 100644 --- a/STATUS.md +++ b/STATUS.md @@ -65,12 +65,32 @@ barrier against a runaway backup filling the space the machine needs to run; put first means that when it comes down, the thing watching is already working and already tested. *(R-167, R-158)* +**The backup partition is gone from the base image.** A machine built from now on has one storage area +instead of two, so a backup can use whatever space the machine actually has free rather than a fixed +slice decided when it was built. The wall does not move; it stops existing. + +**What replaced the wall.** It was quietly doing a second job — keeping a runaway backup from eating +the space the machine needs to keep running. That job is now explicit: if a backup would push the disk +below a safe reserve, **that one app's backup is refused, its last good copy is left exactly as it +was, and you are told.** Nothing is ever deleted to make room; every app has only one local copy, so +"delete the oldest" would always mean destroying some other app's only copy. + +**How the shape was chosen — worth one line, because it was not the obvious one.** Three ways of doing +it were built and rebooted rather than argued about. All three worked. They differed in what they +quietly broke: one put your backups inside Docker's own storage, where the normal way of fixing a sick +Docker would wipe them; another exposed all of Docker's internals to the part of the system that +manages your drives. The third does neither, and costs one extra line of configuration. + +**Nothing has changed on any existing machine.** They keep their current layout and go on working +exactly as before; they get the new shape only when they are reinstalled. The new base image is +deliberately **not switched on yet** — nothing will pick it up until a machine has been rebuilt from +it and checked, which is the next step. *(R-165, R-163)* + ## What we're working on - **Now:** the last app whose data was never saved; today's decisions written down. -- **Next:** merging the small backup partition into the large one — **the warnings for it are already - done and working**, so this step is now only the partition change. It needs one decision from you - first (below). +- **Next:** finishing the partition merge — **the build is done and the decision is made**; what is + left is to reinstall the two demo machines from the new base image and check one end to end. - **After:** rebuilding how the machine records whether an app is meant to be running. ## Waiting on you @@ -82,35 +102,23 @@ first means that when it comes down, the thing watching is already working and a settle both. *(R-110, R-115)* - **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a session log; nothing suggests anyone else saw it. *(R-132)* -- **The partition merge: one decision, now measured.** Removing the backup partition also removes a - barrier — today a runaway backup is refused on its own and cannot touch the space the machine needs - to run; afterwards it can, and a machine out of that space is stopped, not slow. So: do we add a - hard stop that refuses a backup before it eats the last of the room, or do we rely on the new - warnings? The recommendation is the hard stop, because it keeps exactly what the barrier gave us. - **Second question, which only you can answer:** does the tester's box need converting in place, or - can it be reinstalled? It does not report to the hub, so nothing here can tell. *(R-165, R-176)* +- **Nothing — both partition-merge questions are answered.** You chose the storage shape and the hard + stop; both are built. The tester's box needs no conversion: it will simply be reinstalled. *(R-165)* ## Changed since last update -- **2026-08-02** — The false "host offline" warning is fixed, and the cause was not what it looked - like. The hub's database was supposed to be in a mode where reading a page cannot block a machine's - status update — the code said so, but a one-word syntax difference meant the setting had **never - taken effect**, for the hub's whole life. So opening an operator page could make a machine's report - fail; two failures in a row crossed the half-hour threshold and sent you an alert about a machine - that was up and healthy. It had already done that twice that day. Now genuinely fixed and verified - live. **Also found while checking it: the hub's own database is not in any automatic backup** — it - holds every machine's emergency password and the escrow records. Filed, not yet fixed. +- **2026-08-02** — The false "host offline" warning is fixed. The hub's database was supposed to be in + a mode where reading a page cannot block a machine's status update; a one-word difference meant that + setting had **never taken effect**, for the hub's whole life. Fixed and verified live. **Also found: + the hub's own database is in no automatic backup** — it holds every machine's emergency password. + Filed, not yet fixed. -- **2026-08-02** — Boot recovery finished: the machine records what the customer asked for, and waits - for the system to finish starting before deciding what is missing. Six hard resets, everything back - every time. A hole the previous day's change had opened — starting an app whose external drive was - missing — was found by reading the code, reproduced on the demo box first, and fixed the same day. - **A second instance of the same hole was found and fixed today**, on the path that restarts an app - after an interrupted backup. +- **2026-08-02** — Boot recovery finished; six hard resets, everything back every time. Two instances + of the same hole — starting an app whose external drive was missing — were found by reading the code + and fixed the same day. -- **2026-08-02** — Thirteen mechanical checks had built up across the four repositories and nothing - ran most of them; two were failing quietly, one since 14 July. Both fixed, and the arrangement that - replaced them is described above. +- **2026-08-02** — Thirteen mechanical checks had built up and nothing ran most of them; two were + failing quietly. Fixed, and the arrangement that replaced them is described above. - **2026-08-02** — Decided: the 20 GB backup partition goes away and shares space with app data. That changes the disk layout, so it happens before any machine is installed outside the house. **Measured since:** no machine outside the house is registered yet, so this is as cheap now as it will ever be; diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 78f6eef1..6d74a858 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -86,6 +86,7 @@ | Box survives a **site/network change** (relocation, different subnet, DHCP re-lease) with the control plane intact | agent v0.96.0 (island NIC), host-install v1.19.0, controller (unchanged), bootstrap | **PROVEN-LIVE (2026-07-25)** | **R-50 SHIPPED and deployed to the whole fleet.** The control plane now rides a host-internal, portless island bridge (`vmbr9`, `169.254.253.1/30`↔`.2/30`) with a fixed private address that no LAN/DHCP/site move can invalidate. Proven end-to-end: the spike's F1 replay (renumber the LAN → agent stays bound on the island, control plane HTTP 200; the LAN-literal contrast reproduces the original `bind: cannot assign requested address` daemon-death) + cold-reboot survival (`SPIKE-island-bridge-2026-07-25.md`), the migration runbook run verbatim (`RUNBOOK-island-migration.md`), a fresh provision auto-attaching the island `net1` (A4), and the live migration of **both demo boxes** (demo-hp + demo-felhom, 2026-07-25) — island `/storage` HTTP 200, LAN DNS pinned to the LAN IP (Finding-1), **apps served throughout (0 container restarts)**, hub reporting 0.96.0. **Origin:** `audits/AUDIT-vacation-remote-ops-2026-07-20.md` — the real relocation where the agent's LAN-literal bind took storage/PBS/quiesce/restore-test/DR down silently; that is now structurally impossible on a migrated box | **Fleet: DONE.** Remaining: **R-74** — bring the island to Peti's 2-node cluster (SDN vnet / bridge parity), its own supervised runbook. Related historical: R-51 (dead-primary alerting), R-52 (boot desired-state reconciliation), both shipped | | **The customer is warned BEFORE a filesystem fills** — per filesystem, in Hungarian, naming the drive and the free space, edge-triggered | controller **v0.191.0/.1/.2**, hub **v0.89.0** (R-167, decision D-c) | **PROVEN-LIVE (2026-08-02)** | `audits/SPIKE-r165-mp1-merge-2026-08-02.md` (context) + `felhom-controller/REPORT.md`. Exercised on guest 9201 against a REAL filesystem (`/mnt/sys_drive` filled with `fallocate`): **`disk_warning` at 90% used / 4.7 GB free** → hub `notification_log` `customer | disk_warning | sent` with the dynamic Hungarian rendered; grown to 1.7 GB free → **`disk_critical`** → `customer | sent`; file removed → `critical → ok … cleared silently, re-armed` and the persisted state emptied. **Exactly two events across three boots** — the boot in between produced none, which is the edge trigger holding | **Nothing warned before this.** The only prior signal was the healthcheck's generic `health_degraded` at 90%, for REGISTERED STORAGE PATHS ONLY — it never looked at the docker area or the system-data area, never gave a free-byte figure and never named a drive. **The two event types already existed with NO PRODUCER** (`disk_warning`/`disk_critical`: allowlisted, copy'd, in `DefaultEnabledEvents`, checkbox'd) — the **sixth** *built-but-never-wired* instance here; this ships their producer rather than a seventh near-duplicate type. **Two threshold terms, whichever trips first, and the live proof vindicated the design:** the critical crossing fired on the FREE-BYTE term (1.7 GB) at only **91%** used — a percentage-only rule would have missed it. The hub's generic `customerMessages` entries were REMOVED, because `FormatCustomerEmail` prefers the entry over the message and would discard the label and figures. **Known gap → R-177:** there is no operator-triggerable run-now path; the check is daily 03:30 + once at startup, so confirming a cleared warning on a support call needs a controller restart or a wait | | **A failed per-app Tier-1 backup reaches the OPERATOR** (app, error, and the target filesystem's used/free bytes at the moment of failure) | controller **v0.191.0**, hub **v0.89.0** (R-158, closed by R-167) | **PROVEN-LIVE (2026-08-02)** | `felhom-controller/REPORT.md`. Two real capture failures on guest 9201 (`mkdir …/backups: permission denied`) → both accepted and stored by the hub, `operator | recovery_unit_capture_failed | sent`, and the positive observable **`customer | recovery_unit_capture_failed | skipped | operator_only`** read from the hub's `notification_log`. One event per app, loop continuing | **Before this the failure was a `[WARN]` line and nothing else** — the manager carried three notify seams and none for the unit capture, so `/backups/apps`, the page you open to ask whether ONE app is backed up, was the one page that never said. **Deliberately NOT `backup_failed`:** that type is customer-enabled by default and carries Hungarian copy, so reusing it — which R-158's own proposal said — would email the customer about a failure they cannot act on. **D-c routes it to the operator and overrides the proposal.** Operator-only is enforced by `notify.operatorOnlyEvents`, NOT by the absence of a `customerMessages` entry (the v0.78.0 defect); a red-proof removing the register entry shows the customer receiving it | +| **A local backup is bounded by the box's FREE SPACE, not by a partition set at build time** — the appliance ships ONE data volume, and a capture that would exhaust it is refused per app rather than allowed to stop the container runtime | golden `build-golden.sh` **v3.0.0**, agent **v0.120.0**, controller **v0.192.0** (R-165 / D-a / B2) | **IMPLEMENTED — NOT PROVEN-LIVE** | `audits/SPIKE-r165-phase0-2026-08-03.md` (P1/P2/P3) + the bake transcript. **The golden bake is real evidence and is cited as such:** `build-golden.sh v3.0.0` produced `including mount point mp0 ('/var/lib/felhom')` with **no `mp1` line at all**, and its own guards printed `/var/lib/docker is a real mount`, `/mnt/sys_drive is a real mount` and `both paths are ONE filesystem`. Archive published (registry HTTP 200, sha `54e2a4c4…`). The B2 floor is unit-proven with 3 red-proofs and live on 9201 | **NOT PROVEN-LIVE, and the missing leg is named: no box has been reinstalled from this golden (R-178).** Per this map's own rule a PROVEN-LIVE claim needs an end-to-end citation, and "the golden baked" is not "a box built from it works". **The golden is deliberately UNVOUCHED** so no fresh install picks up an unproven layout — prove first, then vouch (`CONTEXT.md` S-14). Every box in the field is still on the SPLIT layout and is unaffected: nothing assumes the merged shape at runtime, the controller's system_data_path is a path rather than a volume, and agent v0.120.0 FOLDS the retired `-sysdata-grow` into the single grow so an older `felhom-host-install.sh` still provisions the same total capacity | | Soft-quota: usage bar, pre-push enlargement block, customer notification | controller v0.109/134, hub v0.41/55 | **PROVEN-LIVE** | 6D/6E; hub OffsiteChecker | | | **A customer (not the operator) performs a restore via UI alone** | all | **MISSING** (as evidence) | — | Alpha will produce this; script it into R-3. **2026-07-19:** the C6 evidence attempt ran and found a **product gap instead of evidence** — `audits/DIAG-immich-restore-2026-07-19.md`. A customer-driven UI restore of a DB-indexed app cannot currently succeed (R-43 file-only restore, R-44 stale dump), so this row cannot flip until those close. Row stays MISSING **by finding, not by absence of attempt** — the rehearsal system working, not failing. **2026-07-19: the blocking product gaps are CLOSED in controller v0.148.0** (R-43 + R-44 shipped), so this row is now blocked only on the evidence run itself, not on missing capability. It flips the moment the §9 acceptance produces screenshots + the outcome flash + a snapshot ID. **2026-07-19 round 2 — PARTIAL EVIDENCE ONLY, row NOT flipped** (`audits/DIAG-immich-restore-round2-2026-07-19.md`): a deliberate run from snapshot `49e7cb46` did recover all 11 assets (`status=active`, files resolve), but the operation **reported failure** and left immich reporting schema drift, because the replay aborted against the running app (H4). Photos back ≠ clean acceptance. **2026-07-20: H4 closed in controller v0.153.0 (R-47) on BOTH paths, AND THE EVIDENCE RUN HAPPENED.** *(The "closing in v0.149" wording above was wrong — v0.149.0 was the F3 dashboard fix; R-47 shipped in v0.153.0.)* The C6 drill ran end-to-end **through the UI**: photos deleted, **trash emptied**, the full files+database restore pressed on `/backups/restore`, 40 files placed + 1 DB dump replayed rc-0, 11 assets back, no drift, timeline visually confirmed. The method note below is now DEMONSTRATED, not merely written down. Evidence: `felhom-controller/REPORT.md` 4e. **Residual: the run was performed by the OPERATOR, not by a customer** — for this row literal wording the alpha still owes one genuinely customer-driven pass, but no product gap blocks it. Method note for R-3's script: deleting in an app's own UI usually means *trash*, not deletion, so a drill written that way merges 0 files, flashes success and proves nothing — a real drill must empty the trash **and** verify the app's *content*, not the file count **Lane split → `07-backup-architecture.md` §3**: this row is Lane 1 (customer, unassisted). §8 rows 1–5 are the routes it would exercise | diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 33b24696..1d56e22e 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -545,11 +545,38 @@ decision D-c; controller v0.191.x + hub v0.89.0).** The last sentence of this se **operator-tier** (`notify.operatorOnlyEvents`) and deliberately not `backup_failed`: a customer can take no action on a capture failure. -**A caveat this section must carry, because the bound below depends on it.** The `mp0 50G / mp1 20G` -table above is **demo-hp's** shape, not the fleet's — demo-felhom ships `mp0 200G / mp1 50G`, where -the same bound is ≈ 49 GB / ≈ 24 GB, and the golden's own defaults are `16 G / 8 G` before provision -grows them. **The bound below is a FUNCTION of `mp1`, not a constant.** Measured 2026-08-02, -`audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1; correcting the numbers throughout is **R-175**. +### 7.5.1 — THE CEILING THIS SECTION DESCRIBES HAS BEEN REMOVED (2026-08-03, R-165 / decision D-a) + +**Everything above describes the SPLIT layout, which is now the legacy shape.** A golden built by +`build-golden.sh` **v3.0.0** ships **one** data volume; `mp1` does not exist. Both consumer paths are +binds of subdirectories of it (variant **V-c**): + +``` +mp0 -> /var/lib/felhom ├─ docker/ --bind--> /var/lib/docker + └─ sys_drive/ --bind--> /mnt/sys_drive +``` + +**So the size bound below no longer applies to a box built from that golden.** A driveless app's +recovery unit is limited by the box's actual free space, not by a partition set at build time. The +mismatch table above (`mp0` 50 G vs `mp1` 20 G) describes what a merged box no longer has. + +**R-175, fixed here rather than left standing.** The bound below was stated as the fleet's and was +**one box's**: it is derived from `mp1 = 20 G`, which is demo-hp exactly and never was demo-felhom +(`mp0 200G / mp1 50G`, where the same arithmetic gives ≈ 49 GB / ≈ 24 GB), nor the golden (`16 G / 8 G` +before provision grew them). **Read it as a function of `mp1`, and only for a box still on the split +layout.** Measured: `audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1. + +**What replaced the partition's second job.** `mp1` was also a BULKHEAD: an overflow was refused per +app with the last good unit byte-identical, and it **could not reach `/var/lib/docker`**, because that +was a different filesystem. On a merged box it can. Decision **B2**, shipped in controller +**v0.192.0**, is that bulkhead made deliberate — a two-term capture floor (97% used or 1 GiB free) in +`internal/fillwatch`'s shape, sitting beyond its critical band so the customer is always warned first. +It **refuses per app and never deletes**: nothing on this filesystem is generational, so pruning could +only destroy a different app's only local copy. + +**Status caveat, deliberately explicit:** as of 2026-08-03 **no box has been reinstalled from the +merged golden** (R-178), so every box in the field is still on the split layout and everything above +still describes them exactly. This subsection describes what a box built from golden ≥ 0.192.0 gets. Two things are deliberately **not** recorded here. **The sizing ratio is the operator's ruling** (**R-163**) — this section states the constraint, not a number. And **the same-device placement is diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index dadda18f..d6efc59c 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -90,8 +90,9 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha | **R-171** | **The boot sweep started apps whose data drive was ABSENT — a regression introduced by v0.189.0, now FIXED.** Replacing `isBootOrphan`'s container-count term with recorded intent made a drive-gate-stopped app (`compose down` ⇒ zero containers, and the gate never touches `desired_state` because it is not the customer) read as a boot orphan | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.190.0, 2026-08-02) | — | **Reasoned from the diff, then CONFIRMED on hardware before any fix was written** (`audits/DIAG-bootrecon-drive-absent-2026-08-02.md`). The sweep found and started calibre-web with its drive unmounted, burned both attempts and handed it to the dead-app alarm — **a false alarm about an app the drive gate is deliberately holding**. The *write* hazard did NOT materialise: compose failed `mkdir …/userdata: permission denied` because the unbound mountpoint is host-root-owned and the guest is unprivileged — **an accidental protection no code owns, no test pins, and one `chown` or one privileged guest away from gone**. Fix: new consumer-side seam `bootrecon.StartGate`, **fail-safe (cannot determine ⇒ do not start)**, wired in `main.go`; `Manager.DriveLive` reuses the userdata belt's own `isMountPoint` seam so the two cannot drift. **The rule is not new** — the API's `startGatedByMissingDrive` already refused this to the customer; the sweep bypassed it by calling `Manager.StartStack` directly. Widening the window (R-157 A) made two more holders reachable, so the same seam also refuses an app held by a **quiesce** or an **in-flight app-data operation** (§8.2), reusing `quiesce.SuppressedStacks()` and a new read-only `AppStopGuard.HeldStacks()`. Held apps report as `HeldByDrive`, never `StillDown` — that is the alarm's bucket. **ID established free:** `grep -ro "R-171\b" documentation/ *.md` → 0 hits before minting | — | | **R-172** | ~~**A false `host_stale` alarm fires when the hub's SQLite refuses two consecutive host reports.**~~ | **CLOSED — SHIPPED + PROVEN-LIVE** (hub v0.88.0, 2026-08-02) | — | **ROOT CAUSE WAS NOT TUNING — THE PRAGMAS WERE NEVER APPLIED.** `store.New` used `?_journal_mode=WAL&_busy_timeout=5000`, which is **mattn/go-sqlite3** syntax; the driver is **modernc.org/sqlite**, whose `applyQueryParams` reads only `_pragma`/`_time_format`/`_time_integer_format`/`_txlock`/`_inttotime` and **ignores the rest without an error**. The hub ran in rollback-journal mode with `busy_timeout=0` for its entire life while its own source said WAL — a configuration asserting an invariant the code did not provide. Proof: a 128 MB open `/data/hub.db` with **no `-wal`/`-shm` beside it**. **Fix:** `?_pragma=journal_mode(WAL)&_pragma=busy_timeout(5000)&_txlock=immediate`. **`_txlock=immediate` is not optional** — `database/sql`'s `Begin()` is DEFERRED, so a read-then-write tx must upgrade its lock and a failed upgrade is `SQLITE_BUSY_SNAPSHOT`, which **`busy_timeout` does not retry**; this store has 10+ `db.Begin()` sites, all write paths. **Retry options (b) and (c) were deliberately NOT taken** — with readers no longer blocking writers a surviving `SQLITE_BUSY` would be a real signal, and a retry would hide it; revisit only on evidence. **Live:** `-wal`+`-shm` now present, **zero `SQLITE_BUSY` since rollout**, host back to `ok`, and `PRAGMA integrity_check` = `ok` with `journal_mode=wal` after three unrelated OOM restarts. **Operational consequence handled:** a WAL DB cannot be copied by taking `hub.db` alone — the break-glass retrieval in `operations/nodes.md` did exactly that and is now WAL-aware (the live `-wal` was 729 KB, i.e. a bare `cat` would have silently omitted it) | — | | **R-174** | ~~**The app-stop guard's crash recovery started apps onto MISSING drives — a regression in v0.189.0 code.**~~ | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.191.0, 2026-08-02) | — | **Found by REVIEW on 2026-08-02, in code shipped 2026-08-01, and closed the same session — R-171 one path over.** `appStopGuard.SetStarter(stackMgr)` handed `Recover` the RAW stack manager, whose `StartStack` has no drive gate, and `Recover` runs **at startup** — exactly when an external drive may not have come back. So: a backup stops an app, the box loses power, the drive does not remount, and the app is started on a missing drive. The rule was not new — the API's own `startGatedByMissingDrive` already refused this to the customer; the guard bypassed it. **`bootDriveGate` could NOT be reused whole**, and the reason is recorded in the code: its holder #2 reads `bootAppStopGuard.HeldStacks()`, which during `Recover` is **the guard's own marker** — it would refuse every recovery it was meant to perform — and holders #1/#2 read package-level vars assigned AFTER `Recover()` runs, so a whole-gate reuse would be correct only by accident of nil-safety. Holder #3 is extracted into a shared `driveStartGate` with **two callers, one implementation**, and `TestBootDriveGateAndAppStopShareTheDrivePredicate` pins the delegation. **A REFUSAL IS NOT A FAILURE:** new `ErrStartRefused` + a `Refused` bucket — both keep the marker, only `Failed` alarms, because routing a deliberate hold into `NotifyBackupFailed` (customer-enabled by default) is the very R-171 false alarm this fixes. `main.go` guards on `Alarming()`, not `!= nil`, and the pre-existing seam test was TIGHTENED to require it. **Live on 9201, both directions:** drive held unmounted → `refusing to restart "calibre-web" … drive /mnt/felhom-drives/hdd_1 is not a live mountpoint`, marker retained byte-identical, zero containers started, `not alarming`; drive returned → `restarted calibre-web`, marker CLEARED. **ID established free:** `grep -ro "R-174\b" documentation/ *.md` → 0 hits | — | -| **R-175** | **`07-backup-architecture.md` §7.5 states ONE box's size bound as if it were the fleet's.** Its headline — *"an app can be restored by Lane 1 only while its recovery unit still fits the retained space — ≈ 19 GB of app data for a file-only app, ≈ 10 GB for a DB-backed one"* — is derived from `mp1 = 20 G`. **That is demo-hp exactly, and it is not demo-felhom**, which ships `mp0 200G / mp1 50G`, where the real bound is ≈ 49 GB / ≈ 24 GB. The golden's own defaults are a third pair (`GOLDEN_DOCKER_GB=16` / `GOLDEN_SYSDATA_GB=8`), grown at provision time | **READY (S) — NEW 2026-08-02** | — | **Measured, not inferred** (`audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1: `pct config 9201` on both hosts). Independent of the merge — the sentence is wrong today and will be wrong differently after R-165. **The fix is to state the bound as a FUNCTION of `mp1`, not a constant**, and to say which box any quoted figure came from. Same class as the comment-asserting-an-invariant rule: a doc stating a fleet-wide number that only one machine satisfies reads as settled and is not. **ID established free:** `grep -ro "R-175\b" documentation/ *.md` → 0 hits | CC | -| **R-176** | **Two prerequisites for the R-165 merge are UNMEASURED, and both are cheap.** (a) Whether a **pre-merge archive** (carrying `mp1`) restore-tests cleanly into a **merged-layout** guest — reading `mountParity` (`felhom-agent/internal/reconcile/restoretest.go:347`) says it should, because the restore recreates `mp1` from the archive so archive and restored guest agree; **that was reasoned from source and never executed.** (b) The in-place per-box migration (move `/felhom-data` onto `mp0`, drop the slot, verify) has **never been rehearsed even once**, so "is the box restorable at every point of it?" is currently unknown | **READY (S) — NEW 2026-08-02** | blocks R-165 landing safely | **Filed because this project's own record is that FOUR production designs specced against unvalidated mechanisms were all wrong** — which is exactly why R-165's own spike refused to design. Both are one command on a **Tier-0** box (D-d: both demo boxes are disposable). (b) is only required work if Peti's box turns out to need migrating rather than reinstalling — the hub cannot answer that (M5: `peti-felhom` exists as a customer with **no host in the register**), so it is the operator's input. **ID established free:** `grep -ro "R-176\b" documentation/ *.md` → 0 hits | CC | +| **R-175** | ~~**`07-backup-architecture.md` §7.5 states ONE box's size bound as if it were the fleet's.**~~ | **CLOSED — FIXED 2026-08-03** (same pass as R-165) | — | **Measured, not inferred** (`audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1: `pct config 9201` on both hosts). Independent of the merge — the sentence is wrong today and will be wrong differently after R-165. **The fix is to state the bound as a FUNCTION of `mp1`, not a constant**, and to say which box any quoted figure came from. Same class as the comment-asserting-an-invariant rule: a doc stating a fleet-wide number that only one machine satisfies reads as settled and is not. **ID established free:** `grep -ro "R-175\b" documentation/ *.md` → 0 hits **FIXED.** §7.5 gained a **7.5.1** which (a) states plainly that the bound is a FUNCTION of `mp1` and applies only to a box still on the split layout, naming all three real shapes, and (b) records that the ceiling itself has been removed by R-165 for boxes built from golden ≥ 0.192.0. Fixed in the same pass as the merge rather than filed and forgotten, because the section would otherwise have been wrong in two ways at once | CC | +| **R-176** | **Two prerequisites for the R-165 merge are UNMEASURED, and both are cheap.** (a) Whether a **pre-merge archive** (carrying `mp1`) restore-tests cleanly into a **merged-layout** guest — reading `mountParity` (`felhom-agent/internal/reconcile/restoretest.go:347`) says it should, because the restore recreates `mp1` from the archive so archive and restored guest agree; **that was reasoned from source and never executed.** (b) The in-place per-box migration (move `/felhom-data` onto `mp0`, drop the slot, verify) has **never been rehearsed even once**, so "is the box restorable at every point of it?" is currently unknown | **(a) ANSWERED 2026-08-03 (P1: PASS). (b) NOT REQUIRED — operator ruling: every node is reinstalled, none migrated** | blocks R-165 landing safely | **Filed because this project's own record is that FOUR production designs specced against unvalidated mechanisms were all wrong** — which is exactly why R-165's own spike refused to design. Both are one command on a **Tier-0** box (D-d: both demo boxes are disposable). (b) is only required work if Peti's box turns out to need migrating rather than reinstalling — the hub cannot answer that (M5: `peti-felhom` exists as a customer with **no host in the register**), so it is the operator's input. **ID established free:** `grep -ro "R-176\b" documentation/ *.md` → 0 hits **UPDATE 2026-08-03.** **(a) is measured and passed** — `audits/SPIKE-r165-phase0-2026-08-03.md` P1: a real pre-merge archive (`mp0+mp1`, confirmed from its own vzdump log) restore-tested on demo-hp, `pass: true`, `mount_parity: ok`, 84 s, with `mountParity` untouched. One limit stated rather than glossed: it ran with the pre-merge agent because the merged one did not exist yet, and the comparison is archive-vs-its-own-restore which never consults the host layout — **re-run it once against agent v0.120.0**, which is one command. **(b) is withdrawn, not deferred:** the operator ruled that every node is REINSTALLED rather than migrated in place (both demo boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed in a few weeks), so the in-place migration rehearsal has no consumer. Recorded explicitly rather than silently skipped | CC | +| **R-178** | **The merged golden (0.192.0) is built and published but NO BOX HAS BEEN REINSTALLED FROM IT, and it is deliberately UNVOUCHED.** `build-golden.sh` v3.0.0 baked it with variant V-c and every retargeted assertion passed on the real bake (`including mount point mp0 ('/var/lib/felhom')`, no `mp1` line, `both paths are ONE filesystem`); it is in the registry (HTTP 200, sha `54e2a4c431daf580…`). What has NOT happened is Part 4: reinstall each demo box from it and prove claim → deploy an app → back up → restore | **READY (M) — NEW 2026-08-03** | blocks R-165 reaching PROVEN-LIVE; blocks the capability-map row | **The golden is UNVOUCHED ON PURPOSE and that is the safe state**, not an oversight: vouching is what makes a fresh install pick it up, so vouching a golden no box has been proven from would put an unproven disk layout in front of the next install anywhere. **Prove first, then vouch** — the bake script's own output treats the hub record as a separate deliberate step for this reason. **Everything else for the merge is shipped and green:** controller v0.192.0 (the B2 floor) is live on 9201, agent v0.120.0 is live on BOTH hosts, and `felhom-host-install.sh` computes the single grow from the thin pool. So a reinstall is now a self-contained piece of work with no code left to write. **Order matters: ONE box at a time**, demo-hp first, proven end to end, and only then demo-felhom — two in parallel leaves no working reference to compare against. Note demo-felhom carries the PBS-DR/offsite tier, so it is the one whose backup chain a reinstall actually disturbs. **ID established free:** `grep -ro "R-178\b" documentation/ *.md` → 0 hits | CC | | **R-177** | **There is no operator-triggerable "run the fill check now" path.** `fill-watch` is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller | **READY (S) — NEW 2026-08-02** | — | **Noticed while live-validating R-167 on 9201, not by a failure.** It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. **Partially mitigated already** — v0.191.2 makes every run log a positive observable (`checked N filesystem(s), M unreadable/skipped, K notification(s)`), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has `GetJobs` but no run-now, so this is a general affordance, not a fill-watch one — **scope it as "run a named scheduler job now", operator-gated.** **ID established free:** `grep -ro "R-177\b" documentation/ *.md` → 0 hits | CC | | **R-173** | **The hub's SQLite PVC is excluded from every Longhorn backup job.** `pvc/hub-data` carries `recurring-job-group.longhorn.io/default: disabled`, and `backup-daily` + `backup-weekly` (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the `default` group — so the 128 MB `/data/hub.db` has **no volume-level backup**. That database holds `host_recovery` (every managed box's break-glass root password), `host_escrow` + `host_escrow_superseded` (escrow custody), `host_pbs_secrets`, `customer_configs`, `dr_recipe` and the wg endpoints/peers — i.e. the material several documented recovery routes depend on | **READY (M) — NEW 2026-08-02** | — | **Noticed while checking the blast radius of the R-172 WAL change, not by a failure** — the WAL work needed to know who copies this file, and the answer turned out to be nobody on a schedule. **Establish before designing:** (a) whether the exclusion is deliberate (a 1 Gi RWO Longhorn volume snapshotting a 128 MB SQLite file is cheap, so the label looks like a leftover rather than a decision) and by whom; (b) whether anything else backs it up out-of-band that this census missed — the `_recovery-inventory-2026-07-28.md` records a MANUAL hot copy, which is not a backup. **When it is designed, it must be WAL-aware** (R-172): a volume snapshot of a live WAL database is crash-consistent and replays on open, which is fine, but any file-level copy must take `hub.db-wal` too or it silently loses the newest writes. **Grep establishing the ID was free:** `grep -ro "R-173\b" documentation/ *.md` → 0 hits | CC | | **R-158** | ~~**A local Tier-1 app-data backup failure reaches no hub channel.**~~ | **CLOSED BY R-167 — SHIPPED + PROVEN-LIVE** (controller v0.191.0 + hub v0.89.0, 2026-08-02) | — | **Closed by the wire it named; no second row was filed for it** (R-167 subsumes and widens it). New `unitNotify` seam + `SetUnitNotify` beside the manager's existing three, called from `captureAllRecoveryUnits` **per app with the loop continuing**, carrying the target filesystem's used/free bytes at the moment of failure — the cause is usually a full filesystem and those numbers answer *why* without an operator logging in. **ROUTED TO THE OPERATOR, NOT `backup_failed`, AND THAT OVERRIDES THIS ROW'S OWN PROPOSAL.** The proposal above said *"emitting the existing `backup_failed`"*; that type carries a `customerMessages` entry AND sits in `settings.DefaultEnabledEvents`, so it would email the customer in Hungarian about a failure they cannot act on — precisely the mistake R-97a avoided by minting `whole_guest_backup_failed`. **Decision D-c routes it to the operator and D-c wins.** New `recovery_unit_capture_failed` in `allowedEventTypes` **and** `notify.operatorOnlyEvents`; `notify.IsOperatorOnly` added so ONE test pins both registers (allowlisted-but-not-operator-only is invisible when they are checked separately — the v0.78.0 defect). **Red-proof:** removing the register entry shows the customer being emailed. **Live on 9201:** two events accepted and stored, `operator | sent`, and the positive observable `customer | recovery_unit_capture_failed | skipped | operator_only` read from the hub's `notification_log` | — | @@ -99,9 +100,9 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha | **R-160** | **gramps-web persisted three paths and wrote to none of them.** `/app/data` appears nowhere in the image's environment; the accounts DB (`GRAMPSWEB_USER_DB_URI`) and **the family tree** (`GRAMPS_DATABASE_PATH=/root/.gramps/grampsdb`) both landed in the writable layer. Upstream persists **eight** paths; the template persisted three, one a phantom. | **SHIPPED** (`templates/gramps-web/docker-compose.yml`, 2026-08-02) | — | **Severity above papra's, and worth keeping visible:** papra loses documents the customer may hold elsewhere; gramps-web loses **the family tree — the artefact built inside the app, of which no other copy exists by construction.** Evidence: `app-catalog-felhom.eu/audits/persistence-sweep-2026-08-02/` | CC | | **R-161** | **The volume-persistence gate is enforced by CONVENTION, not automatically.** The catalog repo has no CI of any kind (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none exists). | **REDUCED SCOPE — open** (operator ruling 2026-08-02) | a second person touching templates | **RULED. Both obvious enforcement points were rejected for measured reasons.** *Controller-side at template load:* rejected because such a check can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — **it would pass on the exact defect it exists to catch**; the property is decidable only at runtime. *CI:* rejected for now — neither repo has any, and there are no users yet. **SHIPPED instead** (`app-catalog-felhom.eu` `fd7747d`): `scripts/catalog_gates.py`, ONE entry point running all three gates, non-zero exit on any failure, **mandated in the catalog's `CLAUDE.md`** the way `site_gates.py` is. Rationale for the record: of this project's gates, the only ones that ever get run are those with a single entry point named in a CLAUDE.md — `site_gates.py` is run, R-29's three orphans are named nowhere and have stopped nothing. **What remains open is only the automatic half:** this is convention, run by a person, and that is sufficient while one person touches templates. Revisit when a second does **UPDATE 2026-08-02:** `catalog_gates.py` gained `--fast` (gate 1 only — the network and runtime gates are deliberately NOT in a hook: a push that pulls images and starts containers gets bypassed within a week and the bypass becomes the habit) and `.githooks/pre-push` now runs it. **The automatic half now has a designated successor row: R-168** (Gitea Actions runner). This row stays open at its reduced scope — the runtime gate remains a deliberate periodic run **UPDATE 2026-08-02 (second):** the automatic half now EXISTS — R-168's runner executes `catalog_gates.py --fast` on every push to this repo (measured: run #1, `image-pin gate OK — 53 templates`, with the two runtime gates announced as skipped and their own output absent from the log). This row's *original* scope — the RUNTIME volume-persistence gate — is deliberately still NOT automatic and should stay that way: CI that pulls 53 images on every push gets disabled. It remains a periodic run | operator | | **R-162** | **`docker diff` is the gate's only witness, and its failure mode is quiet.** The gate's power comes from `docker diff` excluding mounted paths, which makes "in the writable layer" mechanically decidable — an implementation detail of the overlay driver. On a driver where `docker diff` is unsupported or lies, the gate degrades to the mount-occupancy and writability legs **and would not say so**. | **WATCHING** — a limitation, not a defect | — | It **fails closed**: the canary self-test would stop reporting BROKEN and the gate would then refuse to report at all. What is wrong is the message — it would blame the prober rather than the driver. Revisit only if a non-overlay storage driver ever ships | CC | -| **R-163** | **`mp1` is RETENTION, not staging — and it is sized as if it were neither.** A recovery unit is the KEPT copy on the app's **own** drive (`GetAppDrivePath`, `internal/backup/backup.go:245-255`); for an app with no `HDD_PATH` the namespace falls back to the system SSD — *"the SSD-only system-data fallback"* (`internal/appbackup/paths.go:26-27`). There is **no post-copy deletion**: the only prune is F5 (`backup.go:1053-1112`), residue on OLD drives when an app MOVES. So `mp1` (**20 G**) retains the units of every driveless app, while `mp0` permits **50 G** of volumes — and a DB app's unit is up to **~2×** its data (volume tar **plus** SQL dump; measured 21.1 GB → 40.2 GB). `--sysdata-grow` defaults to **0** (`felhom-agent/cmd/felhom-agent/main.go:178`) and is **not** derived from the physical drive; demo-hp's real guest 9201 ships `mp0 50G / mp1 20G`. | **RE-FRAMED 2026-08-02 — open, no longer waiting on a ratio** | — (the sizing question is answered; the work is **R-165**) | **RE-FRAMED, NOT CLOSED (operator decision D-a, 2026-08-02 — `CONTEXT.md` S-5).** The row asked *what ratio should `mp1` be?* and that question is **withdrawn rather than answered**: `mp1` is merged into `mp0` so local recovery units share the app-data area and the ceiling stops existing — a bigger number is the same wall further away. **This row stays open as the record of the constraint** (what `mp1` is for, what it gates, and the measured 2× DB-app unit size) **until the merge lands**, because until then every consequence below is still live on every box. **The work is R-165; the warning that must ship with it is R-167.** Original finding, unchanged, follows. **No number is proposed here deliberately.** What is recorded is the constraint and its blast radius: **`mp1` gates the whole app-data chain**, because Tier-2 mirrors the unit *"(always)"* from `RecoveryUnitPath` (`internal/backup/tier2.go:302,368`) and Tier-3 carries it too — a unit that cannot be written has nothing for either to copy. Bounded on the other side: a unit holds **volume tars + DB dumps only, never `mp8` userdata** (`internal/backup/recovery_unit.go:20-25`), so a 1 TB photo library is never in one. **This bounds D5's Lane-1 independence** — see `architecture/07-backup-architecture.md` §7.5. Overflow itself is SAFE (R-158's measurement: refuses per app, last good unit preserved byte-identical) — what is missing is the warning, which is R-158 (widened to R-167) | CC | +| **R-163** | ~~**`mp1` is RETENTION, not staging — and it is sized as if it were neither.**~~ | **CLOSED by R-165 — the ceiling it describes no longer exists** (golden v3.0.0, 2026-08-03) | — (the sizing question is answered; the work is **R-165**) | **Closed, not merely re-framed.** This row was the record of a constraint that was to stay open *"until the merge lands"*. It has landed: the golden ships ONE data volume, so there is no separate 20 G area for a driveless app's recovery unit to outgrow, and the free space an app can use is the box's actual free space. **What replaced the constraint is recorded on R-165**: the bulkhead the partition also provided is now B2's explicit capture floor (controller v0.192.0), and the measured 2× DB-app unit size this row documented is what justifies the floor's reserve being a reserve rather than a working budget. **Caveat carried forward, deliberately:** no box has been reinstalled from the merged golden yet (R-178), so every box in the field still has the split layout and this row's consequences remain live ON THOSE BOXES until they are reinstalled. Original finding unchanged below | CC | | **R-164** | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | -| **R-165** | **Merge `mp1` into `mp0` — the dedicated 20 G backup partition stops existing.** Operator decision **D-a**, 2026-08-02 (`CONTEXT.md` S-5). Local recovery units share the app-data area instead of holding their own fixed ceiling, so the wall R-163 describes is removed rather than moved further away. Guest 9201 on demo-hp ships `mp0 50G / mp1 20G` today | **READY (M) — NEW 2026-08-02** | — | **Two conditions travel WITH the decision and are not optional.** **(1) Before any external install.** It changes the **disk layout**, so it is a fresh-install shape while there are no external boxes and a per-box migration after — and the decision's cheapness is entirely a function of that ordering. **(2) It removes a wall that currently fails safely**, so **R-167** (D-c: fill warning + failure alert) lands in the same step, never after: today an app that outgrows `mp1` is refused per app with the last good unit preserved byte-identical (R-158's measurement), and after the merge the same overflow consumes the space the app itself is using. Touches the installer/agent guest shape (`--sysdata-grow` defaults to **0** and is not derived from the physical drive, `felhom-agent/cmd/felhom-agent/main.go:178`) and the golden. **Does NOT close R-163** — that row is the record of the constraint and stays open until this lands **MEASURED 2026-08-02 — `audits/SPIKE-r165-mp1-merge-2026-08-02.md`; no layout was touched.** **M1: "the layout" is not one thing** — demo-felhom ships `mp0 200G / mp1 50G`, demo-hp `50G / 20G`, the golden `16G / 8G`; any plan expressed as a fixed pair is already wrong for one of the two boxes (this also invalidates §7.5's fleet-wide bound → **R-175**). **M2: the majority of `mp1` is NOT Tier-1** — on demo-felhom 1.7 G of 2.0 G is Tier-2 mirrors, plus the `userdata/import` drop zone which is on the system drive by CONTRACT (R-75); a plan accounting only for the units is wrong. Observed fill is 5% / 1% — the constraint is a ceiling problem, not a current-fill one. **M3: three assertions break, and the golden fails closed on the split in FOUR places** (`build-golden.sh:126,130` separate-mount + `:315,319` vzdump-exclusion guards), not one; the archive scope `rootfs+mp0+mp1` stays complete after the merge; `mountParity` holds for new archives. **M4 — THE BULKHEAD IS THE REAL COST:** today an overflow is refused per app with the last good unit byte-identical AND CANNOT REACH `/var/lib/docker`; after the merge it can, and a full Docker data-root is a stopped box, not a degraded one. Four candidate replacements costed (reserve / capture-path refusal / project quota / warnings-only); **not chosen — operator's ruling.** **M5: D-a's condition (1) is currently SATISFIED** — no external box is in the hub's host register (only the two ONLINE demo boxes, both **Tier 0 and reinstallable**, so migration cost for the measurable population is ZERO). **`peti-felhom` exists as a customer with NO host in the register, so its layout is UNMEASURED** and was not contacted (D-d). **Recommendation: S1 (one volume, two directories) + B2 (a refusal threshold in the capture path), as a fresh-install shape with the demo boxes REINSTALLED.** **Two prerequisites are unmeasured → R-176.** **STOPPED for the operator's ruling; the merge is a supervised session.** | CC | +| **R-165** | ~~**Merge `mp1` into `mp0` — the dedicated 20 G backup partition stops existing.**~~ | **SHIPPED — golden `build-golden.sh` v3.0.0 + agent v0.120.0 + controller v0.192.0 (B2), 2026-08-03. NOT YET PROVEN-LIVE: no box has been reinstalled from the merged golden** | — | **Variant V-c chosen by the operator on MEASURED evidence, not by reading** (`audits/SPIKE-r165-phase0-2026-08-03.md`): one volume at the NEUTRAL path `/var/lib/felhom`, with `/var/lib/docker` and `/mnt/sys_drive` both binds of subdirectories. Three shapes were built and rebooted; **all three boot, reboot 3/3, give ONE `df` figure and keep a container's `statfs("/")` on the merged volume — the ordering worry that motivated the probe did not materialise.** They differ only in which documented guarantee they break: volume-at-`/var/lib/docker` puts customer backups INSIDE Docker's data-root (so the ordinary "clear /var/lib/docker" reflex destroys every local unit); volume-at-`/mnt/sys_drive` puts Docker's ENTIRE data-root under `/mnt`, which the controller container mounts wholesale — **measured: it then sees `/mnt/sys_drive/docker`**, falsifying the bootstrap's own scoping claim. V-c breaks neither. **P1 answered R-176(a):** a pre-merge archive (`mp0+mp1`) restore-tests clean with `mount_parity: ok` in 84 s; `mountParity` was not weakened. **P3: the four golden assertions were RETARGETED, never deleted, and each was RUN against a deliberately wrong shape — 8 checks, 8 passed**, including a NEW 2b asserting both paths are ONE filesystem (which catches the S2 shape the spike ranked worse than the split) and a new guard for a leftover `mp1` (the old "was mp1 excluded?" pattern could no longer match — a guard that cannot match has silently stopped guarding). **B2 shipped first, in controller v0.192.0**: a two-term capture floor (97% / 1 GiB) in `fillwatch`'s shape, deliberately beyond its critical band so the customer is always warned before a refusal; it refuses per app and **never deletes**, because nothing here is generational. **Golden 0.192.0 is published (registry HTTP 200, sha `54e2a4c4…`) but DELIBERATELY NOT VOUCHED** — vouching is what makes fresh installs pick it up, and the right order is prove-then-vouch. **Remaining: reinstall both demo boxes from it, prove end to end, then vouch → the work is R-178** | CC | | **R-166** | ~~**App state gets a desired/observed model with its own store.**~~ Operator decision **D-b**, 2026-08-02 (`CONTEXT.md` S-5) | **SHIPPED + PROVEN-LIVE** (controller v0.189.0, 2026-08-02) | — | **Both blocking facts were established at source before any code was written, and the answers changed the shape.** **(a) Does a crash-safe journal already exist for the in-flight case?** YES, twice — `internal/quiesce/quiesce.go` (marker + `Recover`, proven on live hardware by Campaign 8 fault 10) and `internal/stacks/migrate.go` (`migration.json` + `RecoverMigration`) — but **neither covers the app-data path**: `DumpAppVolumesSafe` stopped and restarted an app with **no marker, no journal and not even a `defer`**. So the pattern existed and the coverage did not; `backup.AppStopGuard` copies the proven shape into its **own** file (one file, one writer). **(b) Is the SQLite store reachable?** Irrelevant, and deliberately unused: `metrics.db` is optional by design (the controller runs with it absent), and operational state must not live in a store designed to be droppable. **Shipped:** tri-state `desired_state` in `app.yaml` written ONLY by the customer's action (API action switch, `DeployStack`, `UpdateOptionalConfig`'s redeploy branch, `.fab` import — a 14-caller census established that `StartStack`/`StopStack` must NOT be writers); `isBootOrphan` reads intent instead of `len(Containers) > 0`; **absent means UNKNOWN, never running**, so a legacy `app.yaml` keeps byte-identical pre-v0.189.0 behaviour; running-only backfill. D-b's every-container requirement was already met by `aggregateState` and was NOT re-implemented. **Live on 9201:** all three flows (stop survives a restart; a zero-container `running` app is recovered by name; a legacy app.yaml is skipped and never inferred as stopped). **Also fixed en route:** `SaveAppConfig` rebuilt `AppConfig` field-by-field (the R-100 shape) and would have dropped the new field on every save across nine call sites | — | | **R-167** | ~~**Storage monitoring and backup alerts.**~~ | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.191.0/.1/.2 + hub v0.89.0, 2026-08-02) | — | **Operator decision D-c. It shipped BEFORE the R-165 merge, not with it** — D-a's condition (2) says the monitoring lands in the same step and never after, and landing it first is strictly better and costs nothing. **Customer half:** new `internal/fillwatch`, per FILESYSTEM (never per app — one full disk holding ten apps would fire ten times). **It emits the PRE-EXISTING `disk_warning`/`disk_critical` pair, which was allowlisted, copy'd, in `DefaultEnabledEvents` and checkbox'd with NO PRODUCER IN ANY REPO** — a complete customer pipeline with no producer, the **sixth** *built-but-never-wired* instance here; minting a new near-duplicate type would have left it inert forever. **Two threshold terms, whichever trips first** (85% / 5 GiB; critical 95% / 2 GiB) because a percentage alone lies at both ends of this fleet's size range — **proven live: the critical crossing fired on the FREE-BYTE term (1.7 GB) at only 91% used.** Edge-triggered on escalation, state persisted, hysteresis dead zone at 75% / 7 GiB pinned by a test; a nil usage read never warns and never clears one (§8.4). The hub's two generic `customerMessages` entries were **removed** — `FormatCustomerEmail` prefers the entry over the message, so keeping them would discard the drive label and the byte figures. **Operator half: see R-158.** **Live on 9201, all three flows:** `disk_warning` then `disk_critical` both `customer | sent` with the Hungarian rendered, exactly two events across three boots (the edge trigger held on the one between), then a silent clear that re-armed. **v0.191.1** added the once-at-startup run (Daily/Every both wait for their first tick, so a box BOOTING over the line would have stayed silent up to 24 h — the R-100 shape); **v0.191.2** added a per-run positive observable, earned when a quiet run during this session's own validation proved unreadable as evidence. Follow-ups: **R-177** (no run-now path) | — | | **R-168** | ~~CI: no runner exists, and with trunk-based pushes CI can DETECT but not BLOCK~~ | **SHIPPED — and the alarm is DEMONSTRATED** (2026-08-02) | — | **Runner live**: `homelab-manifests/gitea-system/act-runner.yaml`, an unprivileged host-mode `act_runner` in `gitea-system`, one owner-scoped registration serving all four repos (measured: tasks 7-10 all claimed by `felhom-gates-runner`). `.gitea/workflows/gates.yml` in each repo runs that repo's entry point with `--fast` and nothing else; no `uses:` step anywhere. **Six probes, all answered, none STOPped** — `audits/SPIKE-ci-runner-2026-08-02.md`. The two that changed the design: **P2** (stock image has git but NO python3 → custom image `felhom-act-runner:0.1.0`, base pinned, python3 and nothing else) and **P6** (a runner that loses `/data/.runner` re-registers and leaves a dead record behind → the PVC is load-bearing, measured both ways). **P5 is the one that mattered**: a failed run produced NO mail, NO notification row and NO log line from Gitea, so the run now sends its own alarm via Resend and prints the provider's accepted id. **Proven end to end, not asserted**: a deliberately broken commit pushed with `--no-verify` → run #6 `failure` → `RESEND-ACCEPTED id=5ff34766-c5f8-4588-8104-08296aeb45ab`. Posture shown from the live pod spec: `privileged: false`, all caps dropped, no docker socket, no hostPath, `automountServiceAccountToken: false`, sized at half Gitea's limits so it cannot crowd out the service holding every repository on the same node. **The standing limit stays true and is written into the manifest and every workflow: it DETECTS, it does not BLOCK** — making it block is → R-169 | — | From bdd1a9d130ae2ebacdbebd5c113c3ed3ef00c0ce Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 07:16:44 +0200 Subject: [PATCH 04/13] REPORT: CI run ids and conclusions for all six commits (all green) --- REPORT.md | 14 +++++++------- 1 file changed, 7 insertions(+), 7 deletions(-) diff --git a/REPORT.md b/REPORT.md index 2545840f..5f495068 100644 --- a/REPORT.md +++ b/REPORT.md @@ -57,16 +57,16 @@ headline measurements: Checked by **PULL**, matching `head_sha` to each commit — CI mails only on failure, so an unchecked green is an assumption. +**Every commit in this session, across all three repos, is green.** + | Repo | Commit | Task id | Run # | Conclusion | |---|---|---|---|---| -| `felhom-controller` | `4be6467` (v0.192.0) | 39 | 15 | **success** | -| `felhom-agent` | `cd6e267` (v0.120.0) | — | — | *see below* | -| `felhom-agent` | `4bb84fc` (REPORT) | — | — | *see below* | +| `felhom-controller` | `4be6467` (v0.192.0, the B2 floor) | 39 | 15 | **success** | +| `felhom-agent` | `cd6e267` (v0.120.0 + golden 3.0.0) | 41 | 3 | **success** | +| `felhom-agent` | `4bb84fc` (REPORT) | 43 | 4 | **success** | | `felhom.eu` | `7406ac7` (phase-0 audit) | 40 | 21 | **success** | -| `felhom.eu` | `e3525e6` (host-install) | — | — | *see below* | - -*(The remaining rows are filled in from `…/actions/tasks` after the final push; any that had not -finished at write time are listed with their status rather than assumed green.)* +| `felhom.eu` | `e3525e6` (host-install one grow) | 42 | 22 | **success** | +| `felhom.eu` | `14d8c00` (docs + registers) | 44 | 23 | **success** | ## 5. `--no-verify` From aa624496949aa61b084d5fcf2b03dbf7754f4581 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 09:34:15 +0200 Subject: [PATCH 05/13] R-178 CLOSED: both demo boxes reinstalled from the merged golden and proven Two boxes, two DIFFERENT supply paths, so the session proved the disk shape and the delivery route rather than one of them twice. demo-hp (layout proof, --golden ): mp0 at /var/lib/felhom, backup=1, 70G, no mp1; /var/lib/docker and /mnt/sys_drive both real mounts of its subdirectories via fstab; one df figure and one device id (64519) on all three paths; reboots 3/3 with the binds surviving each. demo-felhom (pipeline proof, --force-gitea-golden): fetch_verify succeeding against the vouched manifest for BOTH artifacts -- 'verified sha256 54e2a4c431daf580... matches the hub manifest' for the golden, a7763d31... for the agent. 250G single volume, grep -c '^mp1:' = 0, reboots 3/3. Journey proven on both, endpoint-level: claim -> deploy -> back up -> restore, with a planted marker returning byte-identical on each box. Ceiling measured gone: 65 GiB and 233 GiB available to a recovery unit, against 19 and 45. R-165 -> IMPLEMENTED, not PROVEN-LIVE, on the operator's ruling. B2, which that row records as the bulkhead's replacement, fired live for the first time and does refuse per app, delete nothing and alert -- but it is checked only in captureAllRecoveryUnits while runVolumeDumps writes the bulk unguarded, and its 'the previous unit is untouched' claim was measured false (182,272 B dump replaced by 2,147,666,432 B under a manifest still dated 06:34:26). -> R-181. New: R-179 (uninstall leaves NAS network-storage units), R-180 (--archive-storage not cross-checked against the ACL grant; 403 at step 8/8 after root@pam is rotated), R-181. Third instance of R-115 recorded (agent 0.120.0 unpublished). No code written, no version bumps -- this was a runbook. --- CONTEXT.md | 31 +- REPORT.md | 383 +++++++++++++++--- STATUS.md | 62 ++- .../architecture/00-capability-map.md | 2 +- documentation/backlog/OPEN-ITEMS.md | 9 +- scripts/CHANGELOG.md | 29 ++ 6 files changed, 424 insertions(+), 92 deletions(-) diff --git a/CONTEXT.md b/CONTEXT.md index b915b46a..da6085e2 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -39,12 +39,31 @@ so pruning could only mean deleting a **different** app's only local recovery un boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed shortly. So R-176's in-place migration rehearsal is **withdrawn**, not deferred. -**S-14 — prove first, then vouch (2026-08-03).** Golden **0.192.0** is baked, published and verified in -the registry, and is **deliberately UNVOUCHED**. Vouching is what makes a fresh install pick a golden -up, so vouching one that no box has been proven from would put an unproven disk layout in front of the -next install anywhere. The bake script already treats the hub record as a separate deliberate step; this -makes the ordering a rule. **The remaining work is R-178** — reinstall each demo box from the merged -golden, prove claim → deploy → back up → restore, **one box at a time**, and only then vouch. +**S-14 — prove first, then vouch (2026-08-03) — SPENT, and the ordering did not survive contact.** The +rule was: golden **0.192.0** stays UNVOUCHED until a box has been proven from it, because vouching is +what makes a fresh install pick a golden up. **In the event the golden was vouched at 07:23:26 CEST on +2026-08-03, before any box was reinstalled** (hub log `Artifact manifest set: agent=0.119.0 +golden=0.192.0`), so the ordering was already spent when R-178's session opened; the operator elected +to accept it rather than revert the manifest. **Both boxes were then reinstalled and proven** (R-178, +`REPORT.md`), so the end state is the intended one and no unproven layout was ever in front of a real +install — but the rule protected nothing, because nothing enforced it. **The lesson is R-115's, one +layer up:** an ordering that lives only in a `CONTEXT.md` sentence and a runbook's §7 is a reminder, +and reminders do not hold. If prove-then-vouch is to be a rule it needs the shape R-120's gate has — +a refusal at `handleSetArtifacts`, the sole path to `SetArtifactManifest`, which runs without anyone +choosing to run it. + +**S-15 — the merged layout is proven live, by two different supply paths (2026-08-03, R-178).** Both +demo boxes were wiped and reinstalled from golden 0.192.0 and taken through claim → deploy → back up → +**restore**. **demo-hp** was installed with `--golden ` (the layout proof) and +**demo-felhom** by the normal manifest route with `--force-gitea-golden` (the pipeline proof — +`verified sha256 54e2a4c431daf580… matches the hub manifest`), deliberately different so the session +proved the disk shape *and* the delivery route rather than one of them twice. Live shape on both: +`mp0` at `/var/lib/felhom`, `backup=1`, **no `mp1`**; `/var/lib/docker` and `/mnt/sys_drive` both real +mounts of its subdirectories via `/etc/fstab`; ONE `df` figure and one device id on all three paths; +3/3 reboots each with the binds surviving every time. **What is NOT proven is B2** → **R-181**: the +floor guards `captureAllRecoveryUnits` and not `runVolumeDumps`, which is the leg that fills the +volume, and its refusal message's "the previous unit is untouched" was measured false. R-165 is +therefore **IMPLEMENTED**, not PROVEN-LIVE. **S-11 — D-c's routing, and why R-158's own proposal was overruled (2026-08-02, R-167 SHIPPED).** Decision D-c splits two signals by AUDIENCE, and the split is the ruling: **a fill warning is the diff --git a/REPORT.md b/REPORT.md index 5f495068..b1cb4cf8 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,83 +1,342 @@ -# REPORT — R-165: the `mp1` merge, built and proven at the bake (2026-08-03) +# REPORT — R-178: both demo boxes reinstalled from the merged golden and proven (2026-08-03) -**Overwritten** per the standing rule. The prior contents (hub v0.89.0 / R-167, 2026-08-02) have their -durable record in `hub/CHANGELOG.md` and `CONTEXT.md` S-11/S-12. +**Overwritten** per the standing rule. The prior contents (R-165, the bake, 2026-08-03) have their +durable record in `documentation/backlog/OPEN-ITEMS.md` R-165 and the per-repo CHANGELOGs. -**Companion:** `felhom-agent/REPORT.md` holds the full session detail — probes, variant evidence, the -bake transcript, red-proofs and teardown. **This file covers what changed in THIS repo, and the CI -verification for all three.** - -> **Scope, stated first.** The merge is **built and green through Phase 5**. **Phases 6–7 — reinstalling -> the two demo boxes from the merged golden and proving one end to end — were NOT done, and nothing -> was wiped.** The golden is therefore deliberately **unvouched**. Remaining work: **R-178**. +**Runbook, not a task.** No repo got a version bump and nothing was built. One artifact was +**published** (agent 0.120.0) on an explicit operator ruling — see §2. All code findings are filed as +register rows, not commits, per the runbook's §7. --- -## 1. What changed in `felhom.eu` +## 1. Preconditions P1–P6, each as measured -**No hub change; no hub version bump.** This repo changed in two places, one of them forced. +| # | Precondition | Measurement | +|---|---|---| +| **P1** | Golden sha256 matches R-178's record | **PASS** — computed from the artifact itself on demo-hp: `sha256sum` → `54e2a4c431daf580d2807b82d810be36a7c6a9f697b27094dc2912d26a43b3e0`, 649,547,835 B. Matches R-178's recorded prefix and the hub manifest's stored value exactly. Re-verified after copying to `local`: identical | +| **P2** | A rollback golden still exists | **PASS, three depths.** Split-layout golden **0.188.0** fetchable from Gitea (`HTTP 200`, 649,310,288 B). A **local** split-layout golden sat on each box — `local:backup/vzdump-lxc-9100-2026_07_21-18_24_49.tar.zst` (demo-hp), `…2026_07_20-17_50_57…` (demo-felhom). Best: **full guest vzdumps**, three per box, newest `vzdump-lxc-9201-2026_08_03-04_54_28.tar.zst` (2,351,557,870 B, demo-hp) and `…2026_08_03-04_44_50…` (6,323,506,165 B, demo-felhom), both ~3 h old — so each box could be put back exactly as it was | +| **P3** | Neither demo box holds anything wanted | **PASS, stated explicitly.** demo-hp: paperless-ngx (webserver/postgres/redis) + filebrowser + samba shares. demo-felhom: immich (×4), docmost (×3), calibre-web, bookstack (×2) + filebrowser + samba. Both are demo customers (`Demo HP`, `Demo Ügyfél`); no real customer data. **All of it was destroyed by the wipes and none of it was restored** — that was the point, and the operator confirmed each wipe | +| **P4** | The colleague's box untouched | **PASS** — `peti-felhom` is Tier 2 and out of scope. No command was sent to it; it appears in the hub customer list as `DOWN`, exactly as before | +| **P5** | Both hosts' agent is v0.120.0 | **PASS** — `felhom-agent 0.120.0` on both, before and after | +| **P6** | The hub's host register, BEFORE | **Captured**: `demo-felhom-8363b5 / Demo Ügyfél / 0.120.0 / ONLINE / 1-of-2 guests`; `demo-hp-bb76ea / Demo HP / 0.120.0 / ONLINE / 1-of-2`; `drill-r50-0a4f9a / drill-r50 / 0.113.0 / DOWN` (the fixture, untouched throughout) | -**`scripts/felhom-host-install.sh` — forced by a census, not planned.** `step_grows` computed **two** -volume sizes and the install call passed both, so it had to change with the agent or every install -would have provisioned a half-sized box. It now computes ONE total, summing the old 80/20 split -(`226` = the previous `184 + 42`), so **a standard appliance keeps exactly the 250 G it had** — it is -simply no longer split by a wall. **The size still comes from the physical disk**: `step_grows` -already read the thin pool's real free space, and the merge only collapsed its two outputs into one. -That is the answer to the task's §8.2 — no row needed filing. `--sysdata-grow` is deprecated but still -honoured, because the agent **folds** a hand-passed value in rather than dropping it. +--- -**Documentation**, per the coupling rule — see §3. +## 2. Two findings that contradicted the runbook's premise, both surfaced before any wipe -## 2. The bake, as this repo's audit trail records it +**(a) The golden was ALREADY vouched.** Hub log, a positive observable: +`2026/08/03 07:23:26 [INFO] Artifact manifest set: agent=0.119.0 golden=0.192.0 min_agent="0.113.0" +wrapper_sha=true` — roughly ten minutes before this session's first hub read, and this session had +POSTed nothing. The configuration page confirmed `Artifacts.GoldenVersion=0.192.0`, +`GoldenSHA256=54e2a4c4…3b3e0`. So §7's prove-then-vouch order was already spent, and Phase A's stated +safety ("nothing is official yet, so a failure reaches nobody") was void. **Operator ruling: accept it +and proceed**, keeping the two different supply paths. Recorded on `CONTEXT.md` S-14, now marked SPENT +with the reason — the rule lived only in prose and nothing enforced it. -`documentation/audits/SPIKE-r165-phase0-2026-08-03.md` (new) holds P1, P2 and P3 with method, -measurement and ruling, including the teardown of every probe artefact at all three layers. The -headline measurements: +**(b) Agent v0.120.0 had never been published — the vouched agent was 0.119.0.** +`GET …/generic/felhom-agent/0.120.0/felhom-agent` → **HTTP 404** (0.119.0 → 200). Installer step 5's +idempotent skip requires `installed == vouched` exactly, so a documented-path reinstall would have +**downgraded both boxes** to the pre-merge 0.119.0 — and would have *succeeded* while doing it, since +the current `step_grows` sets `SYSDATA_GROW=0` and 0.119.0's `mp1` resize (`bringup.go` 4c, fatal on +error) therefore never fires. The session would have proven a stack nobody ships. **Operator ruling: +publish and vouch first.** `scripts/publish-agent.sh 0.120.0` from a clean tree at `4bb84fc3` +(`git status --porcelain` empty, `HEAD == origin/main`): upload HTTP 201, **round-trip GET verified**, +`AGENT_SHA256=a7763d31b55b5ce75457b4dba7b06aa300325811834b0be78af4587b47110b9d`. Vouched via +`POST /configuration/artifacts` → `303 …flash=artifacts_set`, hub log +`Artifact manifest set: agent=0.120.0 golden=0.192.0`; the hub resolved the sha authoritatively from +Gitea rather than trusting the submitted value. Filed as the **third instance of R-115**, not a new +ID — R-115 is that row, and it has been WAITING-ON-OPERATOR since 2026-07-29. -- **P1 PASS** — a real pre-merge archive restore-tests clean, `mount_parity: ok`, 84 s, `mountParity` - untouched. -- **P2** — all three variants boot and reboot 3/3; they are separated by **scoping**, not mechanics. - The container's view of `/mnt` is 8.0K under V-a and V-c, and **17.9M — Docker's entire data-root — - under V-b**. The operator chose **V-c**. -- **P3** — the four retargeted golden assertions, run against a deliberately wrong shape: **8/8**. +--- -## 3. Documentation coupling +## 3. demo-hp — the layout proof -| File | Change | +**Install path.** Uninstall: `./felhom-host-install.sh --uninstall --vmid 9201` (typed-vmid +confirmation supplied over a pty). Install: + +``` +./felhom-host-install.sh --customer-id demo-hp --mode appliance --vmid 9201 \ + --cores 7 --memory 26906 \ + --golden local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst \ + --passphrase-file /root/.pp-demo-hp +``` + +Script fetched from `https://felhom.eu/scripts/felhom-host-install.sh`, **v1.22.0**, sha256 +`ed02acb2da46c8d2b5c486ce99d5b9a2747e8786c6eb03652cf755ed1abdd9f4` — byte-identical to the repo copy +that was read. The passphrase went file→file into a 0600 file and never onto a command line. A +`--dry-run` preceded the real run and resolved `manifest: agent v0.120.0 (sha a7763d31…), golden +v0.192.0` with `grows: rootfs +0G (->32G), data +46G (->70G, ONE volume)`. + +**One deviation, mine, and it cost a restart.** The first attempt staged the golden on +`felhom-backup` with `--archive-storage felhom-backup`. Pre-flight passed; **step 8/8** failed: +`HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)` — +the default pool-scoped ACL grants `local local-lvm felhom-pbs` only. Fixed by copying the golden to +`local` (sha re-verified after the copy) and `--resume`. Filed as **R-180**: the condition is +statically checkable in pre-flight, and the failure lands *after* step 4b has rotated and vaulted +root@pam. + +**Layout evidence.** + +``` +mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=70G # and NO mp1 line +rootfs: local-lvm:vm-9201-disk-0,size=32G + +/var/lib/felhom /dev/mapper/pve-vm--9201--disk--1 ext4 rw,relatime,stripe=16 +/var/lib/docker /dev/mapper/pve-vm--9201--disk--1[/docker] ext4 +/mnt/sys_drive /dev/mapper/pve-vm--9201--disk--1[/sys_drive] ext4 + +Filesystem Size Used Avail Use% Mounted on +/dev/mapper/pve-vm--9201--disk--1 69G 977M 65G 2% /var/lib/felhom +/dev/mapper/pve-vm--9201--disk--1 69G 977M 65G 2% /var/lib/docker +/dev/mapper/pve-vm--9201--disk--1 69G 977M 65G 2% /mnt/sys_drive + +stat -c %d → 64519 for all three paths +/etc/fstab: /var/lib/felhom/docker → /var/lib/docker ; /var/lib/felhom/sys_drive → /mnt/sys_drive +both binds writable (touch succeeded on each) +``` + +`/mnt/sys_drive` shows a second `findmnt` row — it is the controller container's +`-v /mnt:/mnt:rslave` propagation (peer group 383 vs the fstab bind's 333), the same shape the split +layout had, not a stacked bind. + +**Reboots — each individually, minimum three:** + +| # | started | controller healthy | outcome | +|---|---|---|---| +| 1 | 08:13:33 | 08:13:55, `Up 12 seconds (healthy)` | one filesystem, 69G/65G on both paths | +| 2 | 08:13:59 | 08:14:15, `Up 7 seconds (healthy)` | same | +| 3 | 08:14:19 | 08:14:35, `Up 7 seconds (healthy)` | same | + +After all three, `mountpoint -q` returns true for **all three paths** and both binds still resolve to +the single volume's subdirectories — which is what the reboots exist to test. `uptime -s` = +`2026-08-03 06:14:24 UTC`, matching reboot 3, so these were real reboots. + +**The journey — method stated: endpoint-level, not a browser.** `claude-in-chrome` does not exist on +DooPlex; every step below invoked the exact endpoint the dashboard's own JavaScript calls. + +| Leg | Endpoint | Observable | +|---|---|---| +| Claim | `POST /claim` with the pre-auth HMAC CSRF token (64 chars) **and** its `felhom_claim_csrf` cookie | `302 → /`; gate discriminator flipped `{"error":"dashboard not yet claimed"}` → `{"error":"authentication required"}`. **The code is emailed-only (R-119) — the operator supplied it**, after an operator resend rotated the generation (the previous code had been consumed 12 d earlier) | +| Deploy | `POST /api/stacks//deploy` with `{"values":{…}}`, field values scraped from the deploy form exactly as the form's own `fetch` does | `{"ok":true}`; `privatebin Up (healthy)`, then `opengist Up (healthy)` | +| Back up | `POST /api/debug/backup/dbdump` — this runs the **production** `RunDBDumps` path (DB leg → `runVolumeDumps` → `captureAllRecoveryUnits`); the debug route only starts it instead of waiting for 03:30 | `Volume dump: opengist/… → 178.0 KB`, `privatebin/… → 2.5 KB`, `App-data backup completed … 2 volume dump(s) (3.962s)`, `Recovery unit captured for …` ×2 | +| **Restore** | `POST /backup/restore` `stack_name=privatebin snapshot_id=primary` | A marker planted in the live volume was **deleted**, then restored: `{"ok":true,"message":"privatebin visszaállítva (primary)."}` in **9.2 s**, and the marker returned with an **identical sha256 `ac1faae6134871d640d5cb1bbc6b5092d1e96b7e46c6c3bcb6e8236cc33ae861`**. App `Up (healthy)` afterwards | + +**Recovery unit path, and it lands on the single volume:** +`/mnt/sys_drive/felhom-data/backups/primary//{manifest.json,compose/,volume-dumps/}`, whose `df` +is `/dev/mapper/pve-vm--9201--disk--1` — the merged volume. + +**Note on the first app chosen.** `privatebin`'s catalog entry declares no `backup:` section, so its +first per-stack capture produced `"volume_dumps": null` — correct for that declaration, not a defect, +but it means a per-stack `POST /stacks//backup` writes compose+config only; the volume leg lives in +the full pass. A second app (`opengist`) was deployed so the run had both a capture and, later, a +refusal. + +**The ceiling is gone, measured:** a recovery unit can use **65 GiB** — the whole volume — against the +**19 GiB** the pre-wipe `mp1` slice offered (`/dev/mapper/pve-vm--9201--disk--2 20G 95M 19G 1%`). + +--- + +## 4. The floor's first live firing — and it does not do what it says + +**Instrument, proven before use.** demo-hp's thin pool is **53.93 GiB** and the auto-sized volume is +70 G, so a real fill to 97 % would have exhausted the pool and corrupted every guest on the box, +including the protected `drill-r50` fixture. A 5 GiB `fallocate` probe moved guest `df` from +`977M used / 65G avail` to `6.0G / 60G` while thin-pool `data_percent` stayed **29.03 → 29.03** — +zero blocks allocated — and cleanup returned both to baseline. The floor reads `statfs`, which is +exactly what `fallocate` moves, so the condition it guards is genuinely present. + +**Setup.** A **real** 2 GiB file in opengist's data volume (so its capture writes 2 GiB), then +`fallocate` to leave `69G / 63G used / 3.0G avail / 96%` — **both floor terms deliberately still +clear**, so the run would start. Capture order was established empirically from the previous run's log +(opengist first, privatebin second), not assumed. + +**What happened, 06:40:03:** + +``` +Volume dump: opengist/opengist_opengist_data → 2.0 GB # unguarded +Volume dump: privatebin/privatebin_privatebin_data → 2.5 KB +App-data backup completed … 2 volume dump(s) (20.737s) +[WARN] Recovery unit capture REFUSED for opengist — … 1.0 GB free; the previous unit is untouched and NOTHING was deleted +[WARN] Recovery unit capture REFUSED for privatebin — … 1.0 GB free; the previous unit is untouched and NOTHING was deleted +[INFO] Event pushed: recovery_unit_capture_failed (error) — … ×2 (HTTP 200) +``` + +**What holds:** it refuses **per app** rather than aborting the run; **nothing was deleted** (both +units present afterwards); and the operator alert reached the hub — `recovery_unit_capture_failed`, +severity `error`, accepted `HTTP 200`, twice. + +**What does not hold — measured, not inferred.** The refusal message claims *"the previous unit is +untouched"*: + +| unit file | before | after | +|---|---|---| +| `privatebin/volume-dumps/privatebin_privatebin_data.tar` | `26c546c2…` | **`b538ab89…`** | +| `opengist/volume-dumps/opengist_opengist_data.tar` | 182,272 B | **2,147,666,432 B** | + +Both were rewritten by the earlier leg, while each `manifest.json` kept +`created_at: 2026-08-03T06:34:26Z` and its `checksums` block covers only the three compose files — so +a unit's payload can be swapped under a stale descriptor and nothing inside the unit can detect it. + +**Cause, in the code and not the log.** The floor is consulted in exactly one place — +`m.unitFloorBlocked(stack.Name)` at `recovery_unit.go:328`, inside `captureAllRecoveryUnits`, which +writes a manifest and a compose copy: a few KB. `runVolumeDumps` (`backup.go:535`) — the leg that +writes the bulk, and the leg that consumed the reserve — has **no floor check at all**; its gates are +protected-stack, volume-less, disconnected, decommissioned. And it runs first *by design* +(`backup.go:483`). **The floor guards the cheap leg and not the leg that fills the volume.** + +This is why R-165 is IMPLEMENTED and not PROVEN-LIVE: that row records B2 as the deliberate +replacement for the bulkhead the `mp1` partition provided, and pre-merge the unguarded leg could only +fill a dedicated 20 G partition — post-merge it can reach Docker's data-root. Filed as **R-181**. +**No code was written**, per the runbook's §7. + +**Cleanup:** fill file and payload removed; `df` back to `1.2G used / 65G avail`; a clean re-run left +both units valid (`Volume dump … 178.0 KB` / `2.5 KB`, `App-data backup completed … (2.399s)`). + +--- + +## 5. demo-felhom — the pipeline proof + +**Install path — deliberately different, and this is the reason the second box exists.** + +``` +./felhom-host-install.sh --customer-id demo-felhom --mode appliance --vmid 9201 \ + --cores 3 --memory 12288 \ + --force-gitea-golden \ + --passphrase-file /root/.pp-demo-felhom +``` + +A copy of golden 0.192.0 already sat on this box's `local` storage — it is where the golden was +**baked** at 06:58 (a `.log` beside it), sha `54e2a4c4…`. `--force-gitea-golden` is the documented +C.3 customer path and overrides local discovery in **both** pre-flight and step 7, which pre-flight +confirmed: `golden: none local — will fetch + verify from Gitea in step 7/8`. The bake artifact was +left untouched. + +**`fetch_verify` succeeding against the vouched sha — the observable this box exists to produce:** + +``` +5/8 fetching agent binary v0.120.0 from Gitea … + verified sha256 a7763d31b55b5ce7… matches the hub manifest +7/8 fetching golden v0.192.0 from Gitea → /var/lib/vz/dump/vzdump-lxc-9100-2026_08_03-09_15_34.tar.zst + verified sha256 54e2a4c431daf580… matches the hub manifest + golden imported + verified: local:backup/vzdump-lxc-9100-2026_08_03-09_15_34.tar.zst +Day-0 provision SUCCESS — vmid=9201 host_id=demo-felhom-8363b5 customer=demo-felhom +``` + +Both artifacts — agent and golden — were fetched anonymously (the normal customer shape) and +sha-verified against the manifest. Controller `0.192.0` healthy. + +**Layout:** `mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=250G`; +`pct config 9201 | grep -c '^mp1:'` → **0**. Both binds real mounts +(`…disk--1[/docker]`, `…disk--1[/sys_drive]`), `df` one figure — `246G 977M 233G 1%` on all three +paths — `stat -c %d` = `64519` on all three, fstab carries both binds. + +**Reboots:** 1 — 09:18:55 → healthy 09:19:09; 2 — 09:19:12 → 09:19:26; 3 — 09:19:30 → 09:19:45. All +three paths still mountpoints after each; `uptime -s` = `2026-08-03 07:19:34 UTC`, matching reboot 3. + +**Journey:** claim (`302 → /`, discriminator flipped to `authentication required`; code supplied by +the operator after a resend) → deploy `opengist` (`{"ok":true}`, `Up (healthy)`) → capture +(`Volume dump: opengist/… → 178.0 KB`, `Recovery unit captured`, unit on +`/dev/mapper/pve-vm--9201--disk--1`, marker present inside the tar) → **restore** +(`{"ok":true,"message":"opengist visszaállítva (primary)."}` in **9.4 s**, marker back with identical +sha256 `bc5507987f3f56dc19a9c24785826c304f1a986998d4aeb0b06d1a810b59e939`, app healthy). + +**Ceiling:** 233 GiB available to a recovery unit, against the 45 GiB the pre-wipe `mp1` offered. + +**Step 8 was NOT repeated here, deliberately** — stating it rather than leaving it ambiguous. The +floor fired on demo-hp and R-181 characterises it fully; re-firing would add no information and would +mean filling a 246 G volume. + +--- + +## 6. Vouching + +Already done before the session (§2a), at **07:23:26 CEST on 2026-08-03**, by an operator action in +the hub UI. The session's own manifest write was the **agent** half, at **07:44:31**: +`Artifact manifest set: agent=0.120.0 golden=0.192.0 min_agent="0.113.0" wrapper_sha=true`. The +manifest afterwards, read back: `agent_sha256=a7763d31b55b5ce7…10b9d`, +`golden_sha256=54e2a4c431daf580…43b3e0`, `min_agent=0.113.0`, `wrapper_sha256=104db0a4…` (preserved +verbatim). The R-120 gate did not block: golden 0.192.0 equals the newest controller the fleet +reports. + +--- + +## 7. Teardown — all three layers + +1. **The machine.** No throwaway guest was created this session, so there is none to delete. VM 300 + `drill-r50` on demo-hp — the protected drift fixture — was never touched and is still `stopped`; + it is not in the `felhom` pool (`pvesh get /pools/felhom` listed only `lxc/9201`), so the + uninstall's shared-box logic never reached it. +2. **The host.** `local-lvm`, before → after: **demo-felhom 29.31 % → 1.37 %** (the old 200 G + 50 G + volumes returned; the new 250 G volume is thin and barely allocated). **demo-hp 39.13 % → + 36.75 %** — *higher than a clean reinstall would leave it*, because the floor test's 2 GB tar + blocks cannot be reclaimed: `fstrim` inside an unprivileged LXC returns + `FITRIM ioctl failed: Operation not permitted`. No operational impact (the guest shows 65 G free of + 69 G, the host 35.7 GiB free of 53.93), but it is real residue and is stated rather than rounded + away. Each box's data volume is now the only data volume; no old guest volumes remain. + **Residue found and cleared by hand on demo-hp:** `--uninstall` left the NAS network-storage units + `mnt-felhom\x2ddrives-Felhom\x2dShare.{mount,automount}` behind (automount `failed`, parent bind + still mounted) → **R-179**. demo-felhom left none, because it had no network share configured. +3. **The hub.** **No old records exist to dispose of, and this is the honest finding, not an + omission:** both enrollments were **idempotent** — `host REUSED (idempotent — existing + credential)`, `host_id: demo-hp-bb76ea` and `demo-felhom-8363b5`, the same ids as before. The + reinstalls therefore produced **no new host records**, so nothing was orphaned and nothing needed + deleting. Final register: `demo-felhom-8363b5 ONLINE 0.120.0`, `demo-hp-bb76ea ONLINE 0.120.0`, + `drill-r50-0a4f9a DOWN 0.113.0` — the same three rows as at P6. No scratch customers were created. + `/appliances` returns 404 on hub 0.89.0 — there is no appliance-record surface to clean. + +**Secrets:** both retrieval passphrases were moved file→file into 0600 files, used via +`--passphrase-file`, and `shred -u`'d afterwards on both hosts along with the session cookie files and +helper scripts; the local scratch copies were deleted. Nothing was written to a committed file and +`curl -w '%{redirect_url}'` was never used (R-132). + +--- + +## 8. Registers changed + +| Row | Change | |---|---| -| `documentation/architecture/07-backup-architecture.md` | **S-1: the contract changed in the same session.** New **§7.5.1** — the ceiling §7.5 describes no longer exists for a box built from golden ≥ 0.192.0, the bulkhead's replacement (B2) is recorded, and **R-175 is FIXED here**: the bound is restated as a function of `mp1` and scoped to split-layout boxes, naming all three real shapes | -| `documentation/architecture/00-capability-map.md` | new row — **IMPLEMENTED, not PROVEN-LIVE**, with the missing leg named (no box reinstalled, R-178) and the bake cited as the evidence it is | -| `documentation/backlog/OPEN-ITEMS.md` | **R-165** → shipped-not-yet-proven-live; **R-163 CLOSED** (its ceiling no longer exists); **R-175 CLOSED**; **R-176** (a) answered by P1, (b) **withdrawn** — every node is reinstalled, not migrated; **R-178 NEW** | -| `CONTEXT.md` | **S-13** (the variant, chosen on measurement; B2's floor; pruning rejected with its reason; no migration exists) and **S-14** (prove first, then vouch) | -| `STATUS.md` | rewritten section in plain language; the *Waiting on you* item cleared — both questions are answered; **two older entries trimmed** so the page did not grow | -| `scripts/CHANGELOG.md` | the host-install change, with why it was forced | +| **R-178** | **CLOSED** — both boxes reinstalled and proven, by two different supply paths | +| **R-165** | **IMPLEMENTED**, not PROVEN-LIVE — operator ruling; the layout half is proven, the B2 half is not (→ R-181) | +| **R-115** | **Third instance recorded** — agent 0.120.0 built, deployed to both hosts, never published | +| **R-181** *(new)* | The capture floor guards the recovery-unit leg and not `runVolumeDumps`; its "previous unit is untouched" claim measured false | +| **R-180** *(new)* | `--archive-storage` is not cross-checked against the ACL grant; the 403 lands at step 8/8, after root@pam has been rotated | +| **R-179** *(new)* | `--uninstall` leaves NAS network-storage systemd units behind when a share was configured | -## 4. CI — run ids and conclusions +**IDs established free before minting:** `grep -ro "R-179\b\|R-180\b\|R-181\b\|R-182\b"` over +`documentation/` and `*.md` → **0 hits**, and over all four repo roots (`felhom-agent`, +`felhom-controller`, `felhom.eu`, `app-catalog-felhom.eu`) → **0 hits**. `R-182` was checked and left +unused. -Checked by **PULL**, matching `head_sha` to each commit — CI mails only on failure, so an unchecked -green is an assumption. +--- -**Every commit in this session, across all three repos, is green.** +## 9. CI -| Repo | Commit | Task id | Run # | Conclusion | -|---|---|---|---|---| -| `felhom-controller` | `4be6467` (v0.192.0, the B2 floor) | 39 | 15 | **success** | -| `felhom-agent` | `cd6e267` (v0.120.0 + golden 3.0.0) | 41 | 3 | **success** | -| `felhom-agent` | `4bb84fc` (REPORT) | 43 | 4 | **success** | -| `felhom.eu` | `7406ac7` (phase-0 audit) | 40 | 21 | **success** | -| `felhom.eu` | `e3525e6` (host-install one grow) | 42 | 22 | **success** | -| `felhom.eu` | `14d8c00` (docs + registers) | 44 | 23 | **success** | +Appended after the push — run id and conclusion, per `CLAUDE.md`'s pull-check rule. -## 5. `--no-verify` +--- -**Not used anywhere.** Every push in this session ran its repo's `.githooks/pre-push` and passed. +## 10. Observations — noticed and NOT acted on -## 6. What remains — R-178 - -1. Reinstall **demo-hp** from golden 0.192.0 through the real installer path; prove claim → deploy an - app → back up → restore; show `df` proving one filesystem and a recovery unit landing on it. -2. Only then reinstall **demo-felhom** (it carries the PBS-DR/offsite tier, so it is the box whose - backup chain a reinstall actually disturbs). -3. **Then** vouch golden 0.192.0 in the hub, and flip the capability-map row to PROVEN-LIVE. -4. Re-run P1's restore-test against agent v0.120.0 — one command, turns a sound inference into an - observation. +- **The runbook's central claim was wrong in a way that mattered.** R-178 said *"a reinstall is now a + self-contained piece of work with no code left to write."* True about code; false about the artifact + channel — the agent half of the merge was unpublished, and following the documented path without + checking would have downgraded both boxes and produced a green, meaningless result. **No code change + was needed to make any step pass** — §10 asks this loudly, and the answer is no. What was needed was + a publish. +- **Prove-then-vouch was already spent when the session opened.** Not a defect in anything, but the + rule protected nothing because nothing enforced it. R-120's gate is the shape that would. +- **A per-stack `POST /stacks//backup` does not produce volume dumps** — the volume leg lives in the + full app-data pass. Not wrong, but the endpoint's name suggests otherwise and it cost time here. +- **privatebin's recovery unit carries no user data** (its catalog entry declares no `backup:` + section), so restoring it loses every paste. That may be intentional for an expiring, E2E-encrypted + paste bin — but nothing in the app's description tells the customer so. Catalog question, not filed. +- **V-c doubles systemd's mount-unit count** — every docker overlay appears twice, + `var-lib-docker-…-merged.mount` and `var-lib-felhom-docker-…-merged.mount`, because `/var/lib/docker` + is a bind of `/var/lib/felhom/docker`. Cosmetic, inherent to the chosen variant, no action. +- **`fstrim` cannot run inside the guest** (`EPERM`, unprivileged LXC), so space freed inside the guest + is not returned to the thin pool. It did not matter here; on a box that fills and empties repeatedly + it would. +- **The floor's own message mixes units** — it reports `64.2/68.7 GB used (93%)` against a threshold + stated as `97% used or 1.0 GiB free`, while `df` showed 96 %. GB-vs-GiB, so the percentage term + fires later than an operator reading `df` would expect. Minor; noted on R-181's fix shape rather + than filed separately. diff --git a/STATUS.md b/STATUS.md index 4c97d561..bb515504 100644 --- a/STATUS.md +++ b/STATUS.md @@ -1,6 +1,6 @@ # STATUS — what works, what's broken, what's next -**Updated 2026-08-02.** +**Updated 2026-08-03.** > **A view, not a source.** `documentation/backlog/OPEN-ITEMS.md` is the authority on open work; this > page restates part of it in plain words, and **nothing may exist only here**. **Not `CONTEXT.md`**, @@ -31,9 +31,12 @@ also delete it. A daily snapshot is armed as a stopgap, and we have never restor the data would vanish on the next update. Two are fixed, the third is now clear to fix because it is installed nowhere. *(R-156)* -**Local backups get 20 GB while apps get 50 GB.** An app that outgrows the smaller space stops being -backed up locally — and the off-site copy is made from the local one, so that stops too. Nothing is -lost: the last good copy is kept intact. *(R-163)* +**The reserve does not guard the step that actually fills the disk.** The rule that refuses a backup +when space runs low is checked at the wrong moment: the big write happens first, unchecked, and only +the small write after it is refused. So the thing meant to stop a runaway backup is the thing it +runs past. Worse, when it does refuse it says *"your last good copy is untouched"* — and we measured +that copy being overwritten by the earlier step anyway. Found by deliberately filling a rebuilt demo +machine. Nothing was deleted and you were emailed, both correctly. *(R-181)* **When that happens, only one page says so** — no email, no alert. The page that answers "is this app backed up?" is the one that stays silent. *(R-158)* @@ -65,9 +68,18 @@ barrier against a runaway backup filling the space the machine needs to run; put first means that when it comes down, the thing watching is already working and already tested. *(R-167, R-158)* -**The backup partition is gone from the base image.** A machine built from now on has one storage area -instead of two, so a backup can use whatever space the machine actually has free rather than a fixed -slice decided when it was built. The wall does not move; it stops existing. +**The backup partition is gone — and both demo machines now run on the new shape.** Each was wiped +and rebuilt from the new base image on 3 August and taken through the whole journey a customer takes: +set the machine up, install an app, back it up, restore it. One storage area instead of two, and the +space a backup can use went from 19 GB to 65 GB on the small machine and from 45 GB to 233 GB on the +big one. The wall does not move; it stops existing. Each machine was rebooted three times over and +came back correctly every time. **The two machines were rebuilt deliberately differently** — the +first from a copy of the image held locally, the second by the ordinary route a real customer takes, +fetching the image and checking it against the fingerprint we publish, so both the disk shape and the +delivery route are now proven rather than one twice. + +**Their previous demo apps and data are gone.** That was the point of a wipe and you approved it; the +machines now carry a couple of small test apps instead. **What replaced the wall.** It was quietly doing a second job — keeping a runaway backup from eating the space the machine needs to keep running. That job is now explicit: if a backup would push the disk @@ -81,32 +93,42 @@ quietly broke: one put your backups inside Docker's own storage, where the norma Docker would wipe them; another exposed all of Docker's internals to the part of the system that manages your drives. The third does neither, and costs one extra line of configuration. -**Nothing has changed on any existing machine.** They keep their current layout and go on working -exactly as before; they get the new shape only when they are reinstalled. The new base image is -deliberately **not switched on yet** — nothing will pick it up until a machine has been rebuilt from -it and checked, which is the next step. *(R-165, R-163)* +**The new base image is now switched on**, so any machine installed from here on gets the new shape. +The tester's box is untouched and keeps working exactly as before; it gets the new shape whenever it +is reinstalled. *(R-165, R-178)* ## What we're working on - **Now:** the last app whose data was never saved; today's decisions written down. -- **Next:** finishing the partition merge — **the build is done and the decision is made**; what is - left is to reinstall the two demo machines from the new base image and check one end to end. +- **Next:** fixing the reserve so it guards the step that fills the disk, and stopping it from + claiming your last good copy is untouched when it is not *(R-181)*. The partition merge itself is + done and proven on both machines. - **After:** rebuilding how the machine records whether an app is meant to be running. ## Waiting on you -- **How a new version reaches a machine.** Pushing the installer publishes it — half a minute later - every new machine downloads it, with no staging and no way back but another push. And publishing is - a step we remember rather than one the release performs, forgotten twice: a fix can be live here - and still not reach a new machine. Nothing is installing today, so this is the cheapest moment to - settle both. *(R-110, R-115)* +- **How a new version reaches a machine — and it has now been forgotten a third time.** Pushing the + installer publishes it: half a minute later every new machine downloads it, with no staging and no + way back but another push. And publishing is a step we remember rather than one the release + performs. On 3 August the new agent — the half of the partition merge that runs on the machine — + turned out to have been built and installed on both demo machines but **never published**, so a + rebuild would have quietly put the *old* one back and proved a version nobody ships. Caught before + the wipe and fixed in ten minutes, but only because someone happened to look. Nothing is installing + today, so this is the cheapest moment to settle it. *(R-110, R-115)* - **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a session log; nothing suggests anyone else saw it. *(R-132)* -- **Nothing — both partition-merge questions are answered.** You chose the storage shape and the hard - stop; both are built. The tester's box needs no conversion: it will simply be reinstalled. *(R-165)* +- **Nothing on the partition merge itself** — the shape is chosen, built, and now proven on both demo + machines. The tester's box needs no conversion: it will simply be reinstalled. What is left is the + reserve defect above, which is ours to fix, not yours to decide. *(R-165, R-181)* ## Changed since last update +- **2026-08-03** — Both demo machines wiped and rebuilt from the new base image, and taken through + set-up → install an app → back it up → restore it. The backup space ceiling is gone and measured. + The hard stop that replaced the old partition was fired for the first time on real hardware — it + refused, it deleted nothing, it emailed you, **and it turned out to be watching the wrong step**; + that is now the top thing to fix. + - **2026-08-02** — The false "host offline" warning is fixed. The hub's database was supposed to be in a mode where reading a page cannot block a machine's status update; a one-word difference meant that setting had **never taken effect**, for the hub's whole life. Fixed and verified live. **Also found: diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 6d74a858..00fec6ee 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -86,7 +86,7 @@ | Box survives a **site/network change** (relocation, different subnet, DHCP re-lease) with the control plane intact | agent v0.96.0 (island NIC), host-install v1.19.0, controller (unchanged), bootstrap | **PROVEN-LIVE (2026-07-25)** | **R-50 SHIPPED and deployed to the whole fleet.** The control plane now rides a host-internal, portless island bridge (`vmbr9`, `169.254.253.1/30`↔`.2/30`) with a fixed private address that no LAN/DHCP/site move can invalidate. Proven end-to-end: the spike's F1 replay (renumber the LAN → agent stays bound on the island, control plane HTTP 200; the LAN-literal contrast reproduces the original `bind: cannot assign requested address` daemon-death) + cold-reboot survival (`SPIKE-island-bridge-2026-07-25.md`), the migration runbook run verbatim (`RUNBOOK-island-migration.md`), a fresh provision auto-attaching the island `net1` (A4), and the live migration of **both demo boxes** (demo-hp + demo-felhom, 2026-07-25) — island `/storage` HTTP 200, LAN DNS pinned to the LAN IP (Finding-1), **apps served throughout (0 container restarts)**, hub reporting 0.96.0. **Origin:** `audits/AUDIT-vacation-remote-ops-2026-07-20.md` — the real relocation where the agent's LAN-literal bind took storage/PBS/quiesce/restore-test/DR down silently; that is now structurally impossible on a migrated box | **Fleet: DONE.** Remaining: **R-74** — bring the island to Peti's 2-node cluster (SDN vnet / bridge parity), its own supervised runbook. Related historical: R-51 (dead-primary alerting), R-52 (boot desired-state reconciliation), both shipped | | **The customer is warned BEFORE a filesystem fills** — per filesystem, in Hungarian, naming the drive and the free space, edge-triggered | controller **v0.191.0/.1/.2**, hub **v0.89.0** (R-167, decision D-c) | **PROVEN-LIVE (2026-08-02)** | `audits/SPIKE-r165-mp1-merge-2026-08-02.md` (context) + `felhom-controller/REPORT.md`. Exercised on guest 9201 against a REAL filesystem (`/mnt/sys_drive` filled with `fallocate`): **`disk_warning` at 90% used / 4.7 GB free** → hub `notification_log` `customer | disk_warning | sent` with the dynamic Hungarian rendered; grown to 1.7 GB free → **`disk_critical`** → `customer | sent`; file removed → `critical → ok … cleared silently, re-armed` and the persisted state emptied. **Exactly two events across three boots** — the boot in between produced none, which is the edge trigger holding | **Nothing warned before this.** The only prior signal was the healthcheck's generic `health_degraded` at 90%, for REGISTERED STORAGE PATHS ONLY — it never looked at the docker area or the system-data area, never gave a free-byte figure and never named a drive. **The two event types already existed with NO PRODUCER** (`disk_warning`/`disk_critical`: allowlisted, copy'd, in `DefaultEnabledEvents`, checkbox'd) — the **sixth** *built-but-never-wired* instance here; this ships their producer rather than a seventh near-duplicate type. **Two threshold terms, whichever trips first, and the live proof vindicated the design:** the critical crossing fired on the FREE-BYTE term (1.7 GB) at only **91%** used — a percentage-only rule would have missed it. The hub's generic `customerMessages` entries were REMOVED, because `FormatCustomerEmail` prefers the entry over the message and would discard the label and figures. **Known gap → R-177:** there is no operator-triggerable run-now path; the check is daily 03:30 + once at startup, so confirming a cleared warning on a support call needs a controller restart or a wait | | **A failed per-app Tier-1 backup reaches the OPERATOR** (app, error, and the target filesystem's used/free bytes at the moment of failure) | controller **v0.191.0**, hub **v0.89.0** (R-158, closed by R-167) | **PROVEN-LIVE (2026-08-02)** | `felhom-controller/REPORT.md`. Two real capture failures on guest 9201 (`mkdir …/backups: permission denied`) → both accepted and stored by the hub, `operator | recovery_unit_capture_failed | sent`, and the positive observable **`customer | recovery_unit_capture_failed | skipped | operator_only`** read from the hub's `notification_log`. One event per app, loop continuing | **Before this the failure was a `[WARN]` line and nothing else** — the manager carried three notify seams and none for the unit capture, so `/backups/apps`, the page you open to ask whether ONE app is backed up, was the one page that never said. **Deliberately NOT `backup_failed`:** that type is customer-enabled by default and carries Hungarian copy, so reusing it — which R-158's own proposal said — would email the customer about a failure they cannot act on. **D-c routes it to the operator and overrides the proposal.** Operator-only is enforced by `notify.operatorOnlyEvents`, NOT by the absence of a `customerMessages` entry (the v0.78.0 defect); a red-proof removing the register entry shows the customer receiving it | -| **A local backup is bounded by the box's FREE SPACE, not by a partition set at build time** — the appliance ships ONE data volume, and a capture that would exhaust it is refused per app rather than allowed to stop the container runtime | golden `build-golden.sh` **v3.0.0**, agent **v0.120.0**, controller **v0.192.0** (R-165 / D-a / B2) | **IMPLEMENTED — NOT PROVEN-LIVE** | `audits/SPIKE-r165-phase0-2026-08-03.md` (P1/P2/P3) + the bake transcript. **The golden bake is real evidence and is cited as such:** `build-golden.sh v3.0.0` produced `including mount point mp0 ('/var/lib/felhom')` with **no `mp1` line at all**, and its own guards printed `/var/lib/docker is a real mount`, `/mnt/sys_drive is a real mount` and `both paths are ONE filesystem`. Archive published (registry HTTP 200, sha `54e2a4c4…`). The B2 floor is unit-proven with 3 red-proofs and live on 9201 | **NOT PROVEN-LIVE, and the missing leg is named: no box has been reinstalled from this golden (R-178).** Per this map's own rule a PROVEN-LIVE claim needs an end-to-end citation, and "the golden baked" is not "a box built from it works". **The golden is deliberately UNVOUCHED** so no fresh install picks up an unproven layout — prove first, then vouch (`CONTEXT.md` S-14). Every box in the field is still on the SPLIT layout and is unaffected: nothing assumes the merged shape at runtime, the controller's system_data_path is a path rather than a volume, and agent v0.120.0 FOLDS the retired `-sysdata-grow` into the single grow so an older `felhom-host-install.sh` still provisions the same total capacity | +| **A local backup is bounded by the box's FREE SPACE, not by a partition set at build time** — the appliance ships ONE data volume, and a capture that would exhaust it is refused per app rather than allowed to stop the container runtime | golden `build-golden.sh` **v3.0.0**, agent **v0.120.0**, controller **v0.192.0** (R-165 / D-a / B2) | **IMPLEMENTED — the LAYOUT half is PROVEN-LIVE (2026-08-03); the REFUSAL half is not (R-181)** | `REPORT.md` (R-178 reinstalls) + `audits/SPIKE-r165-phase0-2026-08-03.md` (P1/P2/P3) + the bake transcript. **The golden bake is real evidence and is cited as such:** `build-golden.sh v3.0.0` produced `including mount point mp0 ('/var/lib/felhom')` with **no `mp1` line at all**, and its own guards printed `/var/lib/docker is a real mount`, `/mnt/sys_drive is a real mount` and `both paths are ONE filesystem`. Archive published (registry HTTP 200, sha `54e2a4c4…`). The B2 floor is unit-proven with 3 red-proofs and live on 9201 | **The row's FIRST clause is now PROVEN-LIVE; its SECOND is not, and they are separated deliberately.** **Proven (R-178, 2026-08-03):** *"a local backup is bounded by the box's FREE SPACE, not by a partition set at build time"* — both demo boxes reinstalled from this golden, by two different supply paths (demo-hp `--golden `; demo-felhom the normal manifest route with **`verified sha256 54e2a4c431daf580… matches the hub manifest`**), each showing `mp0` at `/var/lib/felhom` with **no `mp1`**, both consumer paths real mounts on ONE filesystem (`stat -c %d` = `64519` on all three), 3/3 reboots each, and claim → deploy → backup → **restore** with a planted marker returning byte-identical. Space available to a recovery unit measured at **65 GiB / 233 GiB**, against the **19 GiB / 45 GiB** those boxes' `mp1` slices offered. **NOT proven — and measured FALSE in part:** *"a capture that would exhaust it is refused per app rather than allowed to stop the container runtime"*. The floor fired live for the first time (demo-hp 06:40:03) and does refuse per app, delete nothing, and alert — **but it is checked only in `captureAllRecoveryUnits`, while `runVolumeDumps` writes the bulk with no floor check at all**, so the leg that exhausts the volume is the unguarded one; and the refusal's claim that the previous unit is untouched was measured false (a 182,272 B dump replaced by 2,147,666,432 B under a manifest still dated 06:34:26). → **R-181**. The golden **is now VOUCHED** (2026-08-03, hub `Artifact manifest set: … golden=0.192.0`), so fresh installs pick up the merged layout. Every box in the field that has not been reinstalled is still on the SPLIT layout and is unaffected: nothing assumes the merged shape at runtime, the controller's system_data_path is a path rather than a volume, and agent v0.120.0 FOLDS the retired `-sysdata-grow` into the single grow so an older `felhom-host-install.sh` still provisions the same total capacity | | Soft-quota: usage bar, pre-push enlargement block, customer notification | controller v0.109/134, hub v0.41/55 | **PROVEN-LIVE** | 6D/6E; hub OffsiteChecker | | | **A customer (not the operator) performs a restore via UI alone** | all | **MISSING** (as evidence) | — | Alpha will produce this; script it into R-3. **2026-07-19:** the C6 evidence attempt ran and found a **product gap instead of evidence** — `audits/DIAG-immich-restore-2026-07-19.md`. A customer-driven UI restore of a DB-indexed app cannot currently succeed (R-43 file-only restore, R-44 stale dump), so this row cannot flip until those close. Row stays MISSING **by finding, not by absence of attempt** — the rehearsal system working, not failing. **2026-07-19: the blocking product gaps are CLOSED in controller v0.148.0** (R-43 + R-44 shipped), so this row is now blocked only on the evidence run itself, not on missing capability. It flips the moment the §9 acceptance produces screenshots + the outcome flash + a snapshot ID. **2026-07-19 round 2 — PARTIAL EVIDENCE ONLY, row NOT flipped** (`audits/DIAG-immich-restore-round2-2026-07-19.md`): a deliberate run from snapshot `49e7cb46` did recover all 11 assets (`status=active`, files resolve), but the operation **reported failure** and left immich reporting schema drift, because the replay aborted against the running app (H4). Photos back ≠ clean acceptance. **2026-07-20: H4 closed in controller v0.153.0 (R-47) on BOTH paths, AND THE EVIDENCE RUN HAPPENED.** *(The "closing in v0.149" wording above was wrong — v0.149.0 was the F3 dashboard fix; R-47 shipped in v0.153.0.)* The C6 drill ran end-to-end **through the UI**: photos deleted, **trash emptied**, the full files+database restore pressed on `/backups/restore`, 40 files placed + 1 DB dump replayed rc-0, 11 assets back, no drift, timeline visually confirmed. The method note below is now DEMONSTRATED, not merely written down. Evidence: `felhom-controller/REPORT.md` 4e. **Residual: the run was performed by the OPERATOR, not by a customer** — for this row literal wording the alpha still owes one genuinely customer-driven pass, but no product gap blocks it. Method note for R-3's script: deleting in an app's own UI usually means *trash*, not deletion, so a drill written that way merges 0 files, flashes success and proves nothing — a real drill must empty the trash **and** verify the app's *content*, not the file count **Lane split → `07-backup-architecture.md` §3**: this row is Lane 1 (customer, unassisted). §8 rows 1–5 are the routes it would exercise | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index d6efc59c..5767345f 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -15,7 +15,7 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha | **R-94** | ~~A hand-synced version constant drifts, and the gate that would catch it is never run~~ | **CLOSED — SHIPPED** (hub v0.87.0, 2026-08-02) | — | **All three legs closed.** **(a) closed by DELETION, not derivation** — deriving is not achievable honestly: the Setup command fetches `felhom-host-install.sh` at RUN TIME from a website that git-syncs `main` every 30 s (R-110), so no build-time value in the hub can be true, and a number that is wrong carries a version number's authority while being a guess. The const, the `pageData.ScriptVersion` field, its assignment and the rendered label are gone; a NOTE stands where the const was so it is not helpfully re-added. **(b)** `hostinstall_gates.py` gate 1 INVERTED — it now asserts the hub carries **no** host-install version literal, in six code shapes across every `.go`/`.html` under `hub/`; and the gate is now invoked, by `scripts/repo_gates.py` and the pre-push hook (→ R-29). **(c)** the tautological `render_test.go:219` assertion is deleted, not replaced — there is no version to assert. It was demonstrated PASSING with the const at `9.9.9` while the script was 1.22.0. The label had been wrong for 19 days (since 2026-07-14) | — | | **R-110** | **`main` is the installer's publish channel — there is no staging.** `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a 30 s period and nginx serves that working tree directly (`location /scripts/`, `root …/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`) receives. There is no tag, no pinned-version path, no staging copy and no rollback other than another push — for the artifact that runs as **root on a virgin box**, the single most privileged thing Felhom ships | **WAITING-ON-OPERATOR (S)** | operator ruling | **Two consequences worth stating:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29 — and the precaution recorded on the old R-94 row as "do not point every new box at an installer that has never run" **was never available to take**. **Open question for the operator, not a defect to fix blind:** whether `/scripts/` should serve a pinned release (tag-tracked path, or a versioned directory with the customer command naming a version) or whether `main`-tracking is the accepted shape for a one-operator product. Exposure today is zero — there are no boxes installing — which is exactly why it is cheap to decide now. **SECOND INSTANCE, found 2026-07-29 by the E-2d run and filed here rather than as a new ID:** `felhom-host-install.sh` fetches **nine** files from `raw/branch/main` (`:2072`–`:2206`) and the hub manifest vouches a sha for exactly **one** (`wrapper_sha256` → `felhom-pbs-apply`; re-checked this run, no drift). E-2a's `felhom-backup-target-apply` (`:2116`) is installed **0755 to `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n` — a root-executed artifact taken from `main` with no pinned integrity, which is this row's class exactly | CC | | **R-111** | ~~**The Day-0 artifact channel is 17 agent releases stale — a box installed today gets agent `0.96.0`, not `0.113.0`.**~~ `felhom-host-install.sh` does not use `main`: it reads the hub-vouched manifest (`:423-436`) and fetches Gitea generic packages (agent `:1945`, golden `:2573`). Gitea's newest are **agent 0.96.0** and **golden 0.161.0**, and the hub's manifest selects exactly those — so a fresh box lands on **agent 0.96.0 + controller 0.161.0** (global floor `v0.156.0` < the golden's 0.161.0, so no self-update) against `main`'s 0.113.0 / 0.185.1. Agent 0.113.0 reached both demo boxes by **direct deploy and was never published** | **SHIPPED 2026-07-29 — the channel now serves agent 0.113.0 + golden 0.185.1** | — | **FIXED the same day it was found.** Agent **0.113.0** built from the clean tree @ `58b598b` and published (`scripts/publish-agent.sh`), sha `5f3247f756cb658e…`, round-trip GET verified. Golden **0.185.1** baked on the nested drill VM embedding controller `0.185.1`, published, sha `dba00f3e845c415e…` — bake clean: `Result=success`, overlay2, **all 3 mounts included** (rootfs+mp0+mp1), 0 FATAL/exclusions, upload HTTP 201, token-leak grep 0; log `drill/bake-0.185.1.log`; GL-1 teardown done (guest 9100 purged, secrets shredded, disk restored to `virgin`). Hub Day-0 manifest moved **both together in one POST** so it never vouched a new agent against an old golden; `min_agent` **0.93.0 → 0.113.0**, which is what controller v0.185.0 declares (`felhom-controller/CHANGELOG.md:15`) — **zero fleet impact, verified: all three enrolled hosts already run agent 0.113.0, so no box is held.** `wrapper_sha256` preserved verbatim (re-checked against `configs/felhom-pbs-apply` — no drift). **The global controller floor was deliberately NOT raised**: the golden now bakes 0.185.1, so a fresh box needs no self-update, and raising it would have been an unnecessary fleet-wide write. Original finding follows. **Found 2026-07-29 by the E-2d Phase 0 gate, which stopped the run before a VM was created.** 17 unpublished releases (v0.97.0–v0.113.0) strand the **entire R-82 tiered-backup arc** plus **F-CRIT-2** (a failed backup looking fresh — 7 days silent) and **F-REBOOT** (a guest rebooted mid-backup never returns): a new customer's box would install without them. **Blocks E-2d's C3/C4/C5** — those test endpoints and events that do not exist in 0.96.0/0.161.0. The **controller is fine** (registry has 0.185.1, floor-driven self-update), so the gap is specific to the two Gitea-generic artifacts. **Mirror of R-110, not a duplicate:** R-110 = the installer publishes instantly with no staging; R-111 = the agent/golden publish gate exists and was never walked. Fix should decide whether publishing joins the release train rather than staying a remembered step (R-29's shape, one layer up). Evidence: `audits/E2D-fresh-vm-2026-07-29.md` **DEFERRED LEG, AND IT RECURRED → R-115.** This row's shipped half stands and is not reopened: the bump happened, was verified, and was proven end-to-end by the E-2d install. But its own closing line — *decide whether publishing joins the release train rather than staying a remembered step* — was never acted on, and agent 0.114.0 reproduced the exact condition the same afternoon. The recurrence is filed as **R-115**, not as a reopen, because the stale-channel finding is closed while the process defect that caused it is a distinct problem with a distinct fix. | CC | -| **R-115** | **Publishing is a remembered step, and it was forgotten within eight hours of being documented as forgettable.** A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published — so "deployed" and "installable" are independent states that drift silently. **Two instances, both real:** **R-111** (2026-07-29 morning) — 17 agent releases v0.97.0–v0.113.0 stranded, so a new customer would have installed without the entire R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT; found only because the E-2d Phase 0 gate happened to look. **Agent 0.114.0** (same afternoon) — the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, and **unpublished until this task**, which blocked Session C: a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix | **WAITING-ON-OPERATOR (M)** | operator ruling on the release process | **The finding is the RECURRENCE, not either instance** — both instances are fixed. R-111's own text already named this leg (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; the leg then recurred the same day, which is the evidence that a note is not a mechanism. **Class: → R-29, one layer up** — a control that exists and is never walked; deliberately NOT given its own ID. **The decision is the operator's; the options, mechanisms first:** (a) **publish as a step in the build/release path**, so deployed and installable cannot diverge; (b) **a gate that refuses to deploy a version that is not published+vouched** — the strongest, and it fails closed; (c) a session-end checklist entry; (d) accept it as manual and add a pre-Session-C verification. **(a) and (b) are mechanisms; (c) and (d) are reminders — and R-29's whole finding is that reminders do not hold.** No code this session by design | CC | +| **R-115** | **Publishing is a remembered step, and it was forgotten within eight hours of being documented as forgettable.** A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published — so "deployed" and "installable" are independent states that drift silently. **Two instances, both real:** **R-111** (2026-07-29 morning) — 17 agent releases v0.97.0–v0.113.0 stranded, so a new customer would have installed without the entire R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT; found only because the E-2d Phase 0 gate happened to look. **Agent 0.114.0** (same afternoon) — the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, and **unpublished until this task**, which blocked Session C: a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix | **WAITING-ON-OPERATOR (M)** | operator ruling on the release process | **The finding is the RECURRENCE, not either instance** — both instances are fixed. R-111's own text already named this leg (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; the leg then recurred the same day, which is the evidence that a note is not a mechanism. **Class: → R-29, one layer up** — a control that exists and is never walked; deliberately NOT given its own ID. **The decision is the operator's; the options, mechanisms first:** (a) **publish as a step in the build/release path**, so deployed and installable cannot diverge; (b) **a gate that refuses to deploy a version that is not published+vouched** — the strongest, and it fails closed; (c) a session-end checklist entry; (d) accept it as manual and add a pre-Session-C verification. **(a) and (b) are mechanisms; (c) and (d) are reminders — and R-29's whole finding is that reminders do not hold.** No code this session by design. **THIRD INSTANCE, 2026-08-03 — and it was found by a runbook that had been told there was nothing left to do.** Agent **v0.120.0** — the agent half of the R-165 merge — was built, committed at `cd6e267`, and deployed to BOTH demo hosts, and was **never published**: `GET …/generic/felhom-agent/0.120.0/felhom-agent` → **HTTP 404** (0.119.0 → 200), and the hub manifest accordingly vouched **0.119.0**. The consequence is the sharpest yet, because installer step 5's idempotent skip requires `installed == vouched` EXACTLY: a documented-path reinstall would have **downgraded both boxes** from the merge-aware 0.120.0 to the pre-merge 0.119.0 — silently, since the current `step_grows` sets `SYSDATA_GROW=0` so 0.119.0's `mp1` resize (`bringup.go` 4c, fatal on error) never fires and the install would have *succeeded* while proving a stack nobody ships. **R-178's own row asserted `agent v0.120.0 is live on BOTH hosts` and `no code left to write`; both were true and both were beside the point** — the gap was publication, which no one checks. Fixed in-session on the operator's ruling: `scripts/publish-agent.sh 0.120.0` (sha `a7763d31b55b5ce75457b4dba7b06aa300325811834b0be78af4587b47110b9d`, round-trip GET verified) then vouched, and both reinstalls then fetched and sha-verified it from Gitea. **This is the third instance of a row that has been WAITING-ON-OPERATOR since 2026-07-29; option (b) — a gate that refuses to deploy or vouch an unpublished version — would have caught all three** | CC | | **R-116** | ~~**The drive-absent alarm and its recovery were a MISMATCHED PAIR — absent fired the GENERIC `storage_disconnected`, return the SPECIFIC `backup_target_restored`; `backup_target_absent` never fired at all**~~ | **SHIPPED + PROVEN-LIVE** (agent v0.116.0, 2026-07-30) | — | **CLOSED. The full four-event sequence, on the wire, on a fresh box** (`audits/R116-v0116-2026-07-30.md`): `backup_target_absent (error)` on detach → `backup_target_restored (info)` on return for the TARGET, and `storage_disconnected (error)` → `storage_reconnected (info)` for a NON-target drive on the same box four minutes apart. **Two matched pairs, correctly discriminated — and discrimination is proven NON-trivially for the first time**, since both prior runs had the target itself emit the generic event. Gate fired in **3 s**; all four events reached the hub, so the specific alarm, its severity, its Hungarian copy and the hub routing are now exercised end-to-end. **Over-correction PASSES with a positive observable** (0 ABSENT lines / 0 drive events over 2m14s with both drives present, target `degraded:false`, while 2 `RETURNED` lines prove the gate was ticking). **NARROWED by the R-117 spike (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §12), and it stands as written:** the 2 `RETURNED` lines are a genuine positive observable, so rule 3 is satisfied — but `degraded:false` over that window was read off a drive whose bind was **dead** (R-117), so the window evidences **"the gate did not over-fire"** and **NOT** **"the drive was healthy."** No other part of this row changes: every input to the pairing fix is configuration-derived (`storage.cfg`'s `path` vs the `.mount` unit's `Where`), which R-117 does not touch. Ran on a nested PVE on **demo-hp** per `runbooks/target-selection.md` — through the **real day-0** from the v1.25.0 ISO, with the agent **installed unaided from the vouched Day-0 manifest** (published sha `b47c5c4dab641ee5…`, independent registry GET verified, manifest read back), drives enrolled through the real endpoints, device loss a real hot-detach. **THE FIX, and the ruling is the substantive part:** the mechanism was first isolated from the captured payload (`DIAG-r116-disks-payload-2026-07-30.md`) after two fixes aimed at shapes that do not occur. **Both smaller-looking options were REJECTED because they regress R-114** — `backup_target_offer.go:79` reads `BackupTarget && MountPath != ""` as *"a real drive with its own mountpoint — healthy"* and returns before its `TargetAbsent` branch, so back-filling `MountPath` on the Observe row **or** flagging the registry row (whose `MountPath` is the stale unit-file value) would have told the customer the backup target is fine while its drive was gone. **R-114's correctness was resting on R-116's bug** — a coupling invisible until the payload existed. Taken instead: the Observe row gets the **guest path only** (`mount_path` stays `""`, which is true) from a new `ConfigPath` (`json:"-"`, so the cross-repo golden + key-set contract is untouched), and the union row is deduped **on guest path** — the join being CONFIGURATION (`storage.cfg`'s `path` vs the `.mount` unit's `Where`), the only identity that survives the device. Tests 845→849; 4 red-proofs each asserted to land, and red-proof 1 replays v0.115.0's code and fails, which is the empirical proof it was inert. Its green test had supplied a `MountPath` production never supplies AND left `DriveTargets` nil so the union loop never ran — both corrected. v0.115.0 left in place (inert, harmless). Teardown all 3 layers; hub layer gate-blocked on ONLINE with the command recorded. **Caveat: the drill's controller was 0.185.1 from the golden, which PREDATES R-114**, so its absent-state banner showed the old false copy — the golden being a release behind, not a regression → **R-120** | — | | **R-120** | ~~**The golden baked a controller that predated R-114 + R-112, so a FRESH box showed the customer the WRONG absent-target message**~~ | **CLOSED — golden rebaked + PROVEN-LIVE, and the class now has an ENFORCED gate** (golden 0.186.0 + hub v0.82.0, 2026-07-30) | — | **`audits/R120-golden-rebake-2026-07-30.md`.** **Half 1 — the artifact.** Golden **0.186.0** baked from `main`'s controller in the DooPlex bake fixture (overlay2 OK, **3 mounts**, FATAL 0, exclusions 0, 618 MB, upload **201**, `GOLDEN_SHA256=b760ac6a33e70700…`, token-leak grep 0, GL-1 teardown, `drill.qcow2` back to `virgin`). Three observables: **published** — anonymous GET (what the installer does) 200 / 648930639 bytes / sha identical to the bake; **vouched** — manifest read BACK; **resolved** — `Artifact manifest served for customer sess-f (agent=0.116.0 golden=0.186.0)`. Floor **untouched** per publish-train rule 2 (`min_controller_version` still 0.156.0; it is a separate form); MinAgent left 0.113.0 as 0.186.0 declares. **Proven on a REAL day-0, not the fixture** (per the Part-1 rule now in `runbooks/target-selection.md`): VM 9402 on demo-hp from the v1.25.0 ISO → `Controller elindult (0.186.0)`. With the target detached the endpoint returned the **`TargetAbsent`** copy — *„A rendszermentés meghajtója nem érhető el — amíg vissza nem csatlakoztatod…"* — **and `offer_path` absent entirely**; the day-old read on the 0.185.1 golden had returned the false system-disk message **plus** an offer of the other drive. **Half 2 — the mechanism, operator ruling REFUSE.** hub **v0.82.0**: the gate sits in `hub/internal/web/configs.go` `handleSetArtifacts` immediately before the only write — the sole UI path to `SetArtifactManifest` — so it runs on every vouch without anyone choosing to, and it **refuses** rather than warning. Signal: `store.NewestReportedControllerVersion()` over `reports.controller_version`, **semver-compared in Go** (`MAX()` in SQL ranks 0.99.0 above 0.186.0 — a pair this fleet has shipped). Fail-open in exactly two deliberate cases: empty golden field, unknown fleet version. **NEAR-MISS RECORDED: the first draft read `guests.controller_version`, a column that exists and that NOTHING writes** — it would always have seen `""` and failed open, i.e. inert, this gate's own failure shape, one grep from shipping. 4 tests through the **production handler** over httptest (never a seam), the refusal asserting **both** the flash **and** that the manifest was not written; red-proof: deleting the block makes the stale golden vouchable again. **PROVEN LIVE on the deployed hub by re-attempting the original mistake:** vouching 0.185.1 → `HTTP 303 …flash=golden_behind_fleet` + `[WARN] artifact vouch REFUSED: golden 0.185.1 is older than the newest controller the fleet reports (0.186.0)`, and the manifest read back **unchanged at 0.186.0**. Recorded on **R-29's audit list** (`ROADMAP.md`) as the **first enforced gate** beside its three orphans, so the contrast is kept — the orphans are unchanged. Teardown all 3 layers; hub layer gate-blocked on ONLINE with the command recorded, exactly as `sess-e` was (and `sess-e` was deleted this run) | — | | **R-117** | ~~**A drive's guest bind becomes a DEAD MOUNT while every signal reads healthy — and it happens in TWO ways, only one of which the original framing covered.** (a) *after a detach/return*: the host raw mount heals onto the NEW device via its fs-UUID-keyed unit while the bind still names the OLD one, so the gate takes its `Return` branch and restarts the customer's apps onto a namespace that `EIO`s on every call; (b) *in STEADY STATE, no cycle at all* — a device that errors without disappearing leaves the raw mount `active`, `BoundUnderParent` `true` and the drive never `Disconnected`, so **the gate produces no action and NOTHING is emitted on any channel**~~ | **SHIPPED + PROVEN-LIVE** (agent **v0.117.0**, 2026-07-30) | — | **CLOSED. `audits/R117-v0117-2026-07-30.md`.** `BoundUnderParent` gains a THIRD term at both /disks sites: `bindLiveness` reads `/proc` only and requires (a) **the bind names the same device as the raw mount** and (b) **the filesystem has not aborted** (`shutdown` **or** `emergency_ro`, both measured). **BOTH CHECKS ARE LOAD-BEARING and this is the substantive part:** R-117 was filed as a detach/return defect, but a device that fails WITHOUT disappearing gives the identical all-signals-healthy state with the **devnos EQUAL** and the drive never `Disconnected`, so the gate emits nothing at all, indefinitely (R-117a) — the device comparison alone cannot see it, and a P1-only fix passes every payload test (red-proof RP3 exists for exactly that). **THREE states, never a bool:** `{Unknown, Live, StaleDevice, Aborted}`, `Unknown` is the zero value, and every caller reads `Usable()` where unknown counts **PRESENT** (absent stops a customer's apps — the `newestArchiveOn` trap). **NO NEW RECOVERY PATH:** `AttachDrive`'s normalize leg already did the repair and three call sites already invoked it (20 s ticker, agent startup, and **the controller's `Return` branch BEFORE `restartStacks`**); all three were defeated by `if n == 1 && GuestSeesMount(...)` logging *"fully live, no-op"* about an EIO namespace. **RULING (asked for, given, flagged for overrule):** `StaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, guest never restarts — init PID identical); `Aborted` ⇒ **quiet no-op and SURFACE**, because a re-bind lands on the SAME dead superblock and this runs every 20 s = an infinite silent retry that masks the state. No operator decision required: it routes an already-broken state into the **existing** gate, event types and Hungarian copy — no new customer-facing concept — and the alternative is apps writing documents into a filesystem that rejects every write. **ORDERING TRAP caught by a test:** abort-first classifies the real return state as aborted (its stale bind carries `shutdown` too) and refuses the repair **while still reporting correctly**, so the abort flag is read off the RAW mount in the stale case. **LIVE on demo-hp** (brought 0.113.0 → 0.117.0 first — see R-121): RETURN `raw 8:32 / bind 8:16 shutdown` ⇒ `stale-device`, usable **false**; IN-PLACE `both 252:11 emergency_ro`, raw unit still `active` ⇒ `filesystem-aborted`, usable **false**; healthy ⇒ `live`; **340–497 µs**. **No block I/O proven by strace** (only `/proc/self/mountinfo`, **0** statfs) — the Part 1 `CLAUDE.md` fence applied to its own first consumer. **No regression through the REAL pipeline:** `GET /disks` with the controller's own credential shows the live backup-target drive `bound_under_parent=True`, with 32 gate lines in 3 min as the positive observable and zero spurious transitions. Tests **849→863**, 29/29 green, **6 red-proofs each verified to land** — and **RP1 failing to fail exposed a HOLLOW test**: the aborted fixture used a `/dev/mapper` device, for which `RoleForStorage` derives `role=system`, and a system row never runs the conjunction, so it reported false by DEFAULT and no mutation could fail it. Fixtures now assert the production row shape first. Teardown all 3 layers; hub layer = the vouched manifest, **retained** (it is the product, not scratch). **NOT covered:** the stale-bind repair on hardware — `StablePathForRaw` hardcodes the live parent, so it would write into guest 9201's namespace (R-117h); and sustained-load behaviour, still unmeasured. Follow-ups **R-117g** (no guided recovery for an aborted fs), **R-117h** (parent dir not test-seamable), **R-121** | CC | @@ -92,7 +92,10 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha | **R-174** | ~~**The app-stop guard's crash recovery started apps onto MISSING drives — a regression in v0.189.0 code.**~~ | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.191.0, 2026-08-02) | — | **Found by REVIEW on 2026-08-02, in code shipped 2026-08-01, and closed the same session — R-171 one path over.** `appStopGuard.SetStarter(stackMgr)` handed `Recover` the RAW stack manager, whose `StartStack` has no drive gate, and `Recover` runs **at startup** — exactly when an external drive may not have come back. So: a backup stops an app, the box loses power, the drive does not remount, and the app is started on a missing drive. The rule was not new — the API's own `startGatedByMissingDrive` already refused this to the customer; the guard bypassed it. **`bootDriveGate` could NOT be reused whole**, and the reason is recorded in the code: its holder #2 reads `bootAppStopGuard.HeldStacks()`, which during `Recover` is **the guard's own marker** — it would refuse every recovery it was meant to perform — and holders #1/#2 read package-level vars assigned AFTER `Recover()` runs, so a whole-gate reuse would be correct only by accident of nil-safety. Holder #3 is extracted into a shared `driveStartGate` with **two callers, one implementation**, and `TestBootDriveGateAndAppStopShareTheDrivePredicate` pins the delegation. **A REFUSAL IS NOT A FAILURE:** new `ErrStartRefused` + a `Refused` bucket — both keep the marker, only `Failed` alarms, because routing a deliberate hold into `NotifyBackupFailed` (customer-enabled by default) is the very R-171 false alarm this fixes. `main.go` guards on `Alarming()`, not `!= nil`, and the pre-existing seam test was TIGHTENED to require it. **Live on 9201, both directions:** drive held unmounted → `refusing to restart "calibre-web" … drive /mnt/felhom-drives/hdd_1 is not a live mountpoint`, marker retained byte-identical, zero containers started, `not alarming`; drive returned → `restarted calibre-web`, marker CLEARED. **ID established free:** `grep -ro "R-174\b" documentation/ *.md` → 0 hits | — | | **R-175** | ~~**`07-backup-architecture.md` §7.5 states ONE box's size bound as if it were the fleet's.**~~ | **CLOSED — FIXED 2026-08-03** (same pass as R-165) | — | **Measured, not inferred** (`audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1: `pct config 9201` on both hosts). Independent of the merge — the sentence is wrong today and will be wrong differently after R-165. **The fix is to state the bound as a FUNCTION of `mp1`, not a constant**, and to say which box any quoted figure came from. Same class as the comment-asserting-an-invariant rule: a doc stating a fleet-wide number that only one machine satisfies reads as settled and is not. **ID established free:** `grep -ro "R-175\b" documentation/ *.md` → 0 hits **FIXED.** §7.5 gained a **7.5.1** which (a) states plainly that the bound is a FUNCTION of `mp1` and applies only to a box still on the split layout, naming all three real shapes, and (b) records that the ceiling itself has been removed by R-165 for boxes built from golden ≥ 0.192.0. Fixed in the same pass as the merge rather than filed and forgotten, because the section would otherwise have been wrong in two ways at once | CC | | **R-176** | **Two prerequisites for the R-165 merge are UNMEASURED, and both are cheap.** (a) Whether a **pre-merge archive** (carrying `mp1`) restore-tests cleanly into a **merged-layout** guest — reading `mountParity` (`felhom-agent/internal/reconcile/restoretest.go:347`) says it should, because the restore recreates `mp1` from the archive so archive and restored guest agree; **that was reasoned from source and never executed.** (b) The in-place per-box migration (move `/felhom-data` onto `mp0`, drop the slot, verify) has **never been rehearsed even once**, so "is the box restorable at every point of it?" is currently unknown | **(a) ANSWERED 2026-08-03 (P1: PASS). (b) NOT REQUIRED — operator ruling: every node is reinstalled, none migrated** | blocks R-165 landing safely | **Filed because this project's own record is that FOUR production designs specced against unvalidated mechanisms were all wrong** — which is exactly why R-165's own spike refused to design. Both are one command on a **Tier-0** box (D-d: both demo boxes are disposable). (b) is only required work if Peti's box turns out to need migrating rather than reinstalling — the hub cannot answer that (M5: `peti-felhom` exists as a customer with **no host in the register**), so it is the operator's input. **ID established free:** `grep -ro "R-176\b" documentation/ *.md` → 0 hits **UPDATE 2026-08-03.** **(a) is measured and passed** — `audits/SPIKE-r165-phase0-2026-08-03.md` P1: a real pre-merge archive (`mp0+mp1`, confirmed from its own vzdump log) restore-tested on demo-hp, `pass: true`, `mount_parity: ok`, 84 s, with `mountParity` untouched. One limit stated rather than glossed: it ran with the pre-merge agent because the merged one did not exist yet, and the comparison is archive-vs-its-own-restore which never consults the host layout — **re-run it once against agent v0.120.0**, which is one command. **(b) is withdrawn, not deferred:** the operator ruled that every node is REINSTALLED rather than migrated in place (both demo boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed in a few weeks), so the in-place migration rehearsal has no consumer. Recorded explicitly rather than silently skipped | CC | -| **R-178** | **The merged golden (0.192.0) is built and published but NO BOX HAS BEEN REINSTALLED FROM IT, and it is deliberately UNVOUCHED.** `build-golden.sh` v3.0.0 baked it with variant V-c and every retargeted assertion passed on the real bake (`including mount point mp0 ('/var/lib/felhom')`, no `mp1` line, `both paths are ONE filesystem`); it is in the registry (HTTP 200, sha `54e2a4c431daf580…`). What has NOT happened is Part 4: reinstall each demo box from it and prove claim → deploy an app → back up → restore | **READY (M) — NEW 2026-08-03** | blocks R-165 reaching PROVEN-LIVE; blocks the capability-map row | **The golden is UNVOUCHED ON PURPOSE and that is the safe state**, not an oversight: vouching is what makes a fresh install pick it up, so vouching a golden no box has been proven from would put an unproven disk layout in front of the next install anywhere. **Prove first, then vouch** — the bake script's own output treats the hub record as a separate deliberate step for this reason. **Everything else for the merge is shipped and green:** controller v0.192.0 (the B2 floor) is live on 9201, agent v0.120.0 is live on BOTH hosts, and `felhom-host-install.sh` computes the single grow from the thin pool. So a reinstall is now a self-contained piece of work with no code left to write. **Order matters: ONE box at a time**, demo-hp first, proven end to end, and only then demo-felhom — two in parallel leaves no working reference to compare against. Note demo-felhom carries the PBS-DR/offsite tier, so it is the one whose backup chain a reinstall actually disturbs. **ID established free:** `grep -ro "R-178\b" documentation/ *.md` → 0 hits | CC | +| **R-181** | **The capture floor guards the cheap leg and not the leg that fills the volume — and its refusal message asserts an invariant the code does not provide.** B2 (controller v0.192.0) is recorded on R-165 as the deliberate replacement for the bulkhead the `mp1` partition used to give. It is consulted in exactly one place — `m.unitFloorBlocked(stack.Name)` at `recovery_unit.go:328`, inside `captureAllRecoveryUnits`, which writes a manifest and a compose copy: **a few KB.** The leg that writes the bulk, `runVolumeDumps` (`backup.go:535`), has **no floor check at all** — its gates are protected-stack, volume-less, disconnected, decommissioned — and it runs FIRST, by design (*"MUST run before captureAllRecoveryUnits so the manifests enumerate the fresh tars"*, `backup.go:483`). So the write that fills the filesystem is unguarded, and the floor then refuses the write that would have cost almost nothing. **Second limb: the refusal message is false.** `recovery_unit.go:331` prints *"the previous unit is untouched and NOTHING was deleted"*. Nothing was deleted — true. Untouched — **measured false**: privatebin's `volume-dumps/privatebin_privatebin_data.tar` went `26c546c2…` → `b538ab89…` and opengist's went **182,272 B → 2,147,666,432 B**, both rewritten by the earlier leg, while each unit's `manifest.json` kept `created_at: 2026-08-03T06:34:26Z` and its `checksums` block covers only the three compose files — so a unit's payload can be swapped under a stale descriptor and **nothing in the unit can detect it** | **READY (M) — NEW 2026-08-03** | blocks **R-165** reaching PROVEN-LIVE | **FIRST LIVE FIRING OF B2, and it is why the runbook asked for one.** Proven on demo-hp 2026-08-03 06:40:03 on a box reinstalled from the merged golden (R-178). Method: a real 2 GiB file in opengist's data volume, then `fallocate` to bring the filesystem to 96 % used / 3.0 GiB free — both floor terms deliberately still clear, so the run started. **The `fallocate` instrument was proven before use** (5 GiB moved guest `df` 977M→6.0G while thin-pool `data_percent` stayed 29.03 → 29.03: zero blocks allocated), because demo-hp's thin pool is 53.93 GiB and a real fill to 97 % of a 70 G volume would have exhausted it and corrupted every guest on the box including the `drill-r50` fixture. Sequence observed: opengist's volume dump wrote **2.0 GB unguarded** → free fell to 1.0 GB → **both** apps' recovery-unit captures were then REFUSED on the `1.0 GiB free` term, each pushing `recovery_unit_capture_failed` (severity `error`) to the hub, accepted HTTP 200. **What DOES hold: it refuses per app rather than aborting the run, it never deletes, and the alert reaches the operator.** **Fix shape, not written this session by design (§7 of the runbook):** the floor belongs before the write in `runVolumeDumps` too, the message must stop claiming what the earlier leg has already falsified, and per `CLAUDE.md` *"a comment asserting an invariant needs a test pinning it"* the pinning test must assert the **consequence** (after a refusal, is the previous unit's payload byte-identical?) and not the mechanism. **Class:** the sixth entry in `CLAUDE.md`'s own table of shipped guarantees the code did not provide — found, as four of those were, only on live hardware | CC | +| **R-180** | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** `felhom-host-install.sh` validates the archive storage EXISTS (`pvesm status --storage`, `:1583`) and that the golden volid RESOLVES on it (`:1661`), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default `local local-lvm felhom-pbs` (`--acl-storages`, which `runbooks/day0-install.md` tells the operator **not** to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step | **READY (S) — NEW 2026-08-03** | — | **Hit live on demo-hp 2026-08-03** during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on `felhom-backup` (the enrolled NVMe, where the box's vzdumps live) and `--archive-storage felhom-backup` passed. Pre-flight passed; steps 1–7 ran; step 8 returned `reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)`. **The cost is the ORDER, not the error** — by the time it fires, step 2 has minted the PVE token, step 4b has **rotated root@pam and vaulted it** (so the old console password is already dead), and step 5 has installed the agent. Recovery was `--resume` after moving the golden to `local`, which worked cleanly. **This is statically checkable in pre-flight**: `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a one-line assertion over two variables both known at `:1583`. Same class as R-29 — the checkable thing that nothing checks | CC | +| **R-179** | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** The teardown's residue-diff provenance (`day0-install.md` Part E: *"a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers"*) is from **v1.9.1**, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps `/etc/systemd/system/mnt-felhom\x2ddrives-.mount` and `.automount` after a full uninstall | **READY (S) — NEW 2026-08-03** | — | **Observed on demo-hp 2026-08-03** after `--uninstall --vmid 9201`: `mnt-felhom\x2ddrives-Felhom\x2dShare.automount` **loaded failed failed**, its `.mount` `loaded inactive dead`, and `mnt-felhom\x2ddrives.mount` still `active mounted` — the uninstall's own output had warned `/mnt/felhom-drives/Felhom-Share is busy — NOT forcing` and `/mnt/felhom-drives root bind left mounted`, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, `daemon-reload`, unmount the autofs then the parent. **NEGATIVE CONTROL, same day:** demo-felhom's uninstall left **nothing** (`ls /etc/systemd/system | grep -i felhom` → only the unrelated `felhom-bootstrap.service`; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** `felhom-bootstrap.service` is NOT residue — it is the ISO first-boot unit, `disabled`+`inactive`, exactly-once and already fired | CC | +| **R-178** | **The merged golden (0.192.0) is built and published but NO BOX HAS BEEN REINSTALLED FROM IT, and it is deliberately UNVOUCHED.** `build-golden.sh` v3.0.0 baked it with variant V-c and every retargeted assertion passed on the real bake (`including mount point mp0 ('/var/lib/felhom')`, no `mp1` line, `both paths are ONE filesystem`); it is in the registry (HTTP 200, sha `54e2a4c431daf580…`). What has NOT happened is Part 4: reinstall each demo box from it and prove claim → deploy an app → back up → restore | **CLOSED — BOTH BOXES REINSTALLED AND PROVEN (2026-08-03)** | blocks R-165 reaching PROVEN-LIVE; blocks the capability-map row | **The golden is UNVOUCHED ON PURPOSE and that is the safe state**, not an oversight: vouching is what makes a fresh install pick it up, so vouching a golden no box has been proven from would put an unproven disk layout in front of the next install anywhere. **Prove first, then vouch** — the bake script's own output treats the hub record as a separate deliberate step for this reason. **Everything else for the merge is shipped and green:** controller v0.192.0 (the B2 floor) is live on 9201, agent v0.120.0 is live on BOTH hosts, and `felhom-host-install.sh` computes the single grow from the thin pool. So a reinstall is now a self-contained piece of work with no code left to write. **Order matters: ONE box at a time**, demo-hp first, proven end to end, and only then demo-felhom — two in parallel leaves no working reference to compare against. Note demo-felhom carries the PBS-DR/offsite tier, so it is the one whose backup chain a reinstall actually disturbs. **ID established free:** `grep -ro "R-178\b" documentation/ *.md` → 0 hits. **CLOSED 2026-08-03 — both boxes reinstalled from the merged golden, by two DIFFERENT supply paths, and proven end to end** (`REPORT.md`). **demo-hp — the layout proof**, installed with `--golden local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst` (installer v1.22.0, sha `ed02acb2…`, byte-identical to the repo copy): `mp0 …mp=/var/lib/felhom,backup=1,size=70G`, **no `mp1`**; `/var/lib/docker` → `…disk--1[/docker]` and `/mnt/sys_drive` → `…disk--1[/sys_drive]`, both real mounts, both writable, both in `/etc/fstab`; ONE `df` figure (69G/65G) and `stat -c %d` = `64519` on all three paths; **reboots 3/3** (08:13:33 / 08:13:59 / 08:14:19, controller healthy in 12s/7s/7s, all three still mountpoints after each). **demo-felhom — the pipeline proof**, installed with `--force-gitea-golden` and NO local golden used (preflight logged *"golden: none local — will fetch + verify from Gitea in step 7/8"*, bypassing the 06:58 bake artifact sitting on the same box): **`verified sha256 54e2a4c431daf580… matches the hub manifest`** for the golden and **`verified sha256 a7763d31b55b5ce7…`** for the agent — the observable this second box exists to produce; `mp0 …size=250G`, `grep -c '^mp1:'` → 0, one `df` figure (246G/233G), reboots 3/3 (09:18:55 / 09:19:12 / 09:19:30). **Journey proven on BOTH**, endpoint-level (no browser on DooPlex — the exact endpoints the dashboard's own JS calls): claim (`POST /claim` with the pre-auth HMAC CSRF + `felhom_claim_csrf` cookie; gate discriminator flipped `dashboard not yet claimed` → `authentication required`) → deploy (`POST /api/stacks//deploy`) → capture (`POST /api/debug/backup/dbdump`, which runs the production `RunDBDumps`) → **restore** (`POST /backup/restore`): a planted marker deleted from the live volume came back with an **identical sha256** on each box (`ac1faae6…ae861` privatebin/demo-hp in 9.2s; `bc550798…b59e939` opengist/demo-felhom in 9.4s), recovery units on the single volume in both cases. **Ceiling gone, measured:** 65 GiB (demo-hp) and 233 GiB (demo-felhom) available to a recovery unit, against the 19 GiB and 45 GiB their pre-wipe `mp1` slices offered. **Two deviations, both the operator's call and both recorded:** the golden was ALREADY vouched when the session opened (hub log `2026/08/03 07:23:26 Artifact manifest set: agent=0.119.0 golden=0.192.0`, ~10 min before the first read of this session — so §7's prove-then-vouch order was already spent and the operator elected to accept it); and agent **0.120.0 had never been published**, so the vouched agent was 0.119.0 — published + vouched before the reinstalls (→ **R-115** third instance). **Three new findings: R-179, R-180, R-181** | — | | **R-177** | **There is no operator-triggerable "run the fill check now" path.** `fill-watch` is reachable only on its daily 03:30 schedule plus the once-at-startup run added in controller v0.191.1 — so the only way to exercise it on demand is to restart the controller | **READY (S) — NEW 2026-08-02** | — | **Noticed while live-validating R-167 on 9201, not by a failure.** It cost a controller restart per observation during validation, and it costs the same on a support call: after a customer frees space, nobody can confirm the warning has cleared without restarting their controller or waiting until 03:30. **Partially mitigated already** — v0.191.2 makes every run log a positive observable (`checked N filesystem(s), M unreadable/skipped, K notification(s)`), so at least a run that DID happen is visible; the gap is triggering one. The scheduler has `GetJobs` but no run-now, so this is a general affordance, not a fill-watch one — **scope it as "run a named scheduler job now", operator-gated.** **ID established free:** `grep -ro "R-177\b" documentation/ *.md` → 0 hits | CC | | **R-173** | **The hub's SQLite PVC is excluded from every Longhorn backup job.** `pvc/hub-data` carries `recurring-job-group.longhorn.io/default: disabled`, and `backup-daily` + `backup-weekly` (04:00 / Sun 05:00) are the ONLY recurring jobs and both target the `default` group — so the 128 MB `/data/hub.db` has **no volume-level backup**. That database holds `host_recovery` (every managed box's break-glass root password), `host_escrow` + `host_escrow_superseded` (escrow custody), `host_pbs_secrets`, `customer_configs`, `dr_recipe` and the wg endpoints/peers — i.e. the material several documented recovery routes depend on | **READY (M) — NEW 2026-08-02** | — | **Noticed while checking the blast radius of the R-172 WAL change, not by a failure** — the WAL work needed to know who copies this file, and the answer turned out to be nobody on a schedule. **Establish before designing:** (a) whether the exclusion is deliberate (a 1 Gi RWO Longhorn volume snapshotting a 128 MB SQLite file is cheap, so the label looks like a leftover rather than a decision) and by whom; (b) whether anything else backs it up out-of-band that this census missed — the `_recovery-inventory-2026-07-28.md` records a MANUAL hot copy, which is not a backup. **When it is designed, it must be WAL-aware** (R-172): a volume snapshot of a live WAL database is crash-consistent and replays on open, which is fine, but any file-level copy must take `hub.db-wal` too or it silently loses the newest writes. **Grep establishing the ID was free:** `grep -ro "R-173\b" documentation/ *.md` → 0 hits | CC | | **R-158** | ~~**A local Tier-1 app-data backup failure reaches no hub channel.**~~ | **CLOSED BY R-167 — SHIPPED + PROVEN-LIVE** (controller v0.191.0 + hub v0.89.0, 2026-08-02) | — | **Closed by the wire it named; no second row was filed for it** (R-167 subsumes and widens it). New `unitNotify` seam + `SetUnitNotify` beside the manager's existing three, called from `captureAllRecoveryUnits` **per app with the loop continuing**, carrying the target filesystem's used/free bytes at the moment of failure — the cause is usually a full filesystem and those numbers answer *why* without an operator logging in. **ROUTED TO THE OPERATOR, NOT `backup_failed`, AND THAT OVERRIDES THIS ROW'S OWN PROPOSAL.** The proposal above said *"emitting the existing `backup_failed`"*; that type carries a `customerMessages` entry AND sits in `settings.DefaultEnabledEvents`, so it would email the customer in Hungarian about a failure they cannot act on — precisely the mistake R-97a avoided by minting `whole_guest_backup_failed`. **Decision D-c routes it to the operator and D-c wins.** New `recovery_unit_capture_failed` in `allowedEventTypes` **and** `notify.operatorOnlyEvents`; `notify.IsOperatorOnly` added so ONE test pins both registers (allowlisted-but-not-operator-only is invisible when they are checked separately — the v0.78.0 defect). **Red-proof:** removing the register entry shows the customer being emailed. **Live on 9201:** two events accepted and stored, `operator | sent`, and the positive observable `customer | recovery_unit_capture_failed | skipped | operator_only` read from the hub's `notification_log` | — | @@ -102,7 +105,7 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha | **R-162** | **`docker diff` is the gate's only witness, and its failure mode is quiet.** The gate's power comes from `docker diff` excluding mounted paths, which makes "in the writable layer" mechanically decidable — an implementation detail of the overlay driver. On a driver where `docker diff` is unsupported or lies, the gate degrades to the mount-occupancy and writability legs **and would not say so**. | **WATCHING** — a limitation, not a defect | — | It **fails closed**: the canary self-test would stop reporting BROKEN and the gate would then refuse to report at all. What is wrong is the message — it would blame the prober rather than the driver. Revisit only if a non-overlay storage driver ever ships | CC | | **R-163** | ~~**`mp1` is RETENTION, not staging — and it is sized as if it were neither.**~~ | **CLOSED by R-165 — the ceiling it describes no longer exists** (golden v3.0.0, 2026-08-03) | — (the sizing question is answered; the work is **R-165**) | **Closed, not merely re-framed.** This row was the record of a constraint that was to stay open *"until the merge lands"*. It has landed: the golden ships ONE data volume, so there is no separate 20 G area for a driveless app's recovery unit to outgrow, and the free space an app can use is the box's actual free space. **What replaced the constraint is recorded on R-165**: the bulkhead the partition also provided is now B2's explicit capture floor (controller v0.192.0), and the measured 2× DB-app unit size this row documented is what justifies the floor's reserve being a reserve rather than a working budget. **Caveat carried forward, deliberately:** no box has been reinstalled from the merged golden yet (R-178), so every box in the field still has the split layout and this row's consequences remain live ON THOSE BOXES until they are reinstalled. Original finding unchanged below | CC | | **R-164** | **C2's chain: the DB volume tar cannot be dropped until a SOUND dump predicate exists.** The unit carries both a volume tar and a SQL dump; the restore uses **both** — the dump is authoritative and replayed *after* the tar so it WINS (F17), with only the DB service up (R-47) — `internal/backup/restore_unit.go:262-266`. Dropping the DB container's tar would halve DB-app units **and** close the R-127(b) initdb-skip password trap (restored PGDATA ⇒ `POSTGRES_PASSWORD` ignored). | **BLOCKED** — on the predicate | a dump-validity predicate that is not `accounts has rows` | **The obvious gate is DEAD, measured:** `ValidateDump` warns when the `accounts` table is empty, and that warning was **correct** — the live DB genuinely had 0 accounts, and seeding one stopped the warning and put the row in the dump. But **a fresh appliance legitimately has zero accounts**, so promoting that predicate to a gate would **block every new customer's first backup**. Order: (1) a sound predicate — dump vs **live** per-table counts, not an absolute expectation; (2) warn→gate; (3) tar-drop. **Until (1), the tar is load-bearing** — not because dumps are bad, but because nothing can yet prove one is good. Pairs with **R-127** | CC | -| **R-165** | ~~**Merge `mp1` into `mp0` — the dedicated 20 G backup partition stops existing.**~~ | **SHIPPED — golden `build-golden.sh` v3.0.0 + agent v0.120.0 + controller v0.192.0 (B2), 2026-08-03. NOT YET PROVEN-LIVE: no box has been reinstalled from the merged golden** | — | **Variant V-c chosen by the operator on MEASURED evidence, not by reading** (`audits/SPIKE-r165-phase0-2026-08-03.md`): one volume at the NEUTRAL path `/var/lib/felhom`, with `/var/lib/docker` and `/mnt/sys_drive` both binds of subdirectories. Three shapes were built and rebooted; **all three boot, reboot 3/3, give ONE `df` figure and keep a container's `statfs("/")` on the merged volume — the ordering worry that motivated the probe did not materialise.** They differ only in which documented guarantee they break: volume-at-`/var/lib/docker` puts customer backups INSIDE Docker's data-root (so the ordinary "clear /var/lib/docker" reflex destroys every local unit); volume-at-`/mnt/sys_drive` puts Docker's ENTIRE data-root under `/mnt`, which the controller container mounts wholesale — **measured: it then sees `/mnt/sys_drive/docker`**, falsifying the bootstrap's own scoping claim. V-c breaks neither. **P1 answered R-176(a):** a pre-merge archive (`mp0+mp1`) restore-tests clean with `mount_parity: ok` in 84 s; `mountParity` was not weakened. **P3: the four golden assertions were RETARGETED, never deleted, and each was RUN against a deliberately wrong shape — 8 checks, 8 passed**, including a NEW 2b asserting both paths are ONE filesystem (which catches the S2 shape the spike ranked worse than the split) and a new guard for a leftover `mp1` (the old "was mp1 excluded?" pattern could no longer match — a guard that cannot match has silently stopped guarding). **B2 shipped first, in controller v0.192.0**: a two-term capture floor (97% / 1 GiB) in `fillwatch`'s shape, deliberately beyond its critical band so the customer is always warned before a refusal; it refuses per app and **never deletes**, because nothing here is generational. **Golden 0.192.0 is published (registry HTTP 200, sha `54e2a4c4…`) but DELIBERATELY NOT VOUCHED** — vouching is what makes fresh installs pick it up, and the right order is prove-then-vouch. **Remaining: reinstall both demo boxes from it, prove end to end, then vouch → the work is R-178** | CC | +| **R-165** | ~~**Merge `mp1` into `mp0` — the dedicated 20 G backup partition stops existing.**~~ | **SHIPPED — golden `build-golden.sh` v3.0.0 + agent v0.120.0 + controller v0.192.0 (B2), 2026-08-03. IMPLEMENTED — the LAYOUT is proven live on both boxes (R-178, 2026-08-03); the BULKHEAD'S REPLACEMENT IS NOT (→ R-181)** | — | **Variant V-c chosen by the operator on MEASURED evidence, not by reading** (`audits/SPIKE-r165-phase0-2026-08-03.md`): one volume at the NEUTRAL path `/var/lib/felhom`, with `/var/lib/docker` and `/mnt/sys_drive` both binds of subdirectories. Three shapes were built and rebooted; **all three boot, reboot 3/3, give ONE `df` figure and keep a container's `statfs("/")` on the merged volume — the ordering worry that motivated the probe did not materialise.** They differ only in which documented guarantee they break: volume-at-`/var/lib/docker` puts customer backups INSIDE Docker's data-root (so the ordinary "clear /var/lib/docker" reflex destroys every local unit); volume-at-`/mnt/sys_drive` puts Docker's ENTIRE data-root under `/mnt`, which the controller container mounts wholesale — **measured: it then sees `/mnt/sys_drive/docker`**, falsifying the bootstrap's own scoping claim. V-c breaks neither. **P1 answered R-176(a):** a pre-merge archive (`mp0+mp1`) restore-tests clean with `mount_parity: ok` in 84 s; `mountParity` was not weakened. **P3: the four golden assertions were RETARGETED, never deleted, and each was RUN against a deliberately wrong shape — 8 checks, 8 passed**, including a NEW 2b asserting both paths are ONE filesystem (which catches the S2 shape the spike ranked worse than the split) and a new guard for a leftover `mp1` (the old "was mp1 excluded?" pattern could no longer match — a guard that cannot match has silently stopped guarding). **B2 shipped first, in controller v0.192.0**: a two-term capture floor (97% / 1 GiB) in `fillwatch`'s shape, deliberately beyond its critical band so the customer is always warned before a refusal; it refuses per app and **never deletes**, because nothing here is generational. **Golden 0.192.0 is published (registry HTTP 200, sha `54e2a4c4…`) but DELIBERATELY NOT VOUCHED** — vouching is what makes fresh installs pick it up, and the right order is prove-then-vouch. **Remaining: reinstall both demo boxes from it, prove end to end, then vouch → the work is R-178**. **STATUS SETTLED 2026-08-03, operator ruling: IMPLEMENTED, not PROVEN-LIVE, and the reason is the substantive part.** R-178 proved the *layout* on both boxes past any doubt — one volume, no `mp1`, both binds real mounts, one `df` figure, 3/3 reboots each, claim→deploy→backup→restore, and the ceiling's removal measured at 65 GiB / 233 GiB against the old 19 GiB / 45 GiB slices. **But B2, which this row records as the bulkhead's deliberate replacement, does not guard the leg that fills the volume** — proven live on demo-hp at 06:40:03 and filed as **R-181**: the floor is consulted ONLY in `captureAllRecoveryUnits` (`recovery_unit.go:328`), while `runVolumeDumps` (`backup.go:535`) writes the bulk with no floor check at all, and its refusal message's claim *"the previous unit is untouched"* was measured FALSE. **This row's own framing is what makes that gate the status:** it says the partition's bulkhead "is now B2's explicit capture floor". Until R-181 closes, the merge has removed a bulkhead and its stated replacement covers the cheap leg only — and post-merge the unguarded leg can reach Docker's data-root, which pre-merge it could not (it could only fill the dedicated 20 G `mp1`). PROVEN-LIVE when R-181 closes and a fill is re-run | CC | | **R-166** | ~~**App state gets a desired/observed model with its own store.**~~ Operator decision **D-b**, 2026-08-02 (`CONTEXT.md` S-5) | **SHIPPED + PROVEN-LIVE** (controller v0.189.0, 2026-08-02) | — | **Both blocking facts were established at source before any code was written, and the answers changed the shape.** **(a) Does a crash-safe journal already exist for the in-flight case?** YES, twice — `internal/quiesce/quiesce.go` (marker + `Recover`, proven on live hardware by Campaign 8 fault 10) and `internal/stacks/migrate.go` (`migration.json` + `RecoverMigration`) — but **neither covers the app-data path**: `DumpAppVolumesSafe` stopped and restarted an app with **no marker, no journal and not even a `defer`**. So the pattern existed and the coverage did not; `backup.AppStopGuard` copies the proven shape into its **own** file (one file, one writer). **(b) Is the SQLite store reachable?** Irrelevant, and deliberately unused: `metrics.db` is optional by design (the controller runs with it absent), and operational state must not live in a store designed to be droppable. **Shipped:** tri-state `desired_state` in `app.yaml` written ONLY by the customer's action (API action switch, `DeployStack`, `UpdateOptionalConfig`'s redeploy branch, `.fab` import — a 14-caller census established that `StartStack`/`StopStack` must NOT be writers); `isBootOrphan` reads intent instead of `len(Containers) > 0`; **absent means UNKNOWN, never running**, so a legacy `app.yaml` keeps byte-identical pre-v0.189.0 behaviour; running-only backfill. D-b's every-container requirement was already met by `aggregateState` and was NOT re-implemented. **Live on 9201:** all three flows (stop survives a restart; a zero-container `running` app is recovered by name; a legacy app.yaml is skipped and never inferred as stopped). **Also fixed en route:** `SaveAppConfig` rebuilt `AppConfig` field-by-field (the R-100 shape) and would have dropped the new field on every save across nine call sites | — | | **R-167** | ~~**Storage monitoring and backup alerts.**~~ | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.191.0/.1/.2 + hub v0.89.0, 2026-08-02) | — | **Operator decision D-c. It shipped BEFORE the R-165 merge, not with it** — D-a's condition (2) says the monitoring lands in the same step and never after, and landing it first is strictly better and costs nothing. **Customer half:** new `internal/fillwatch`, per FILESYSTEM (never per app — one full disk holding ten apps would fire ten times). **It emits the PRE-EXISTING `disk_warning`/`disk_critical` pair, which was allowlisted, copy'd, in `DefaultEnabledEvents` and checkbox'd with NO PRODUCER IN ANY REPO** — a complete customer pipeline with no producer, the **sixth** *built-but-never-wired* instance here; minting a new near-duplicate type would have left it inert forever. **Two threshold terms, whichever trips first** (85% / 5 GiB; critical 95% / 2 GiB) because a percentage alone lies at both ends of this fleet's size range — **proven live: the critical crossing fired on the FREE-BYTE term (1.7 GB) at only 91% used.** Edge-triggered on escalation, state persisted, hysteresis dead zone at 75% / 7 GiB pinned by a test; a nil usage read never warns and never clears one (§8.4). The hub's two generic `customerMessages` entries were **removed** — `FormatCustomerEmail` prefers the entry over the message, so keeping them would discard the drive label and the byte figures. **Operator half: see R-158.** **Live on 9201, all three flows:** `disk_warning` then `disk_critical` both `customer | sent` with the Hungarian rendered, exactly two events across three boots (the edge trigger held on the one between), then a silent clear that re-armed. **v0.191.1** added the once-at-startup run (Daily/Every both wait for their first tick, so a box BOOTING over the line would have stayed silent up to 24 h — the R-100 shape); **v0.191.2** added a per-run positive observable, earned when a quiet run during this session's own validation proved unreadable as evidence. Follow-ups: **R-177** (no run-now path) | — | | **R-168** | ~~CI: no runner exists, and with trunk-based pushes CI can DETECT but not BLOCK~~ | **SHIPPED — and the alarm is DEMONSTRATED** (2026-08-02) | — | **Runner live**: `homelab-manifests/gitea-system/act-runner.yaml`, an unprivileged host-mode `act_runner` in `gitea-system`, one owner-scoped registration serving all four repos (measured: tasks 7-10 all claimed by `felhom-gates-runner`). `.gitea/workflows/gates.yml` in each repo runs that repo's entry point with `--fast` and nothing else; no `uses:` step anywhere. **Six probes, all answered, none STOPped** — `audits/SPIKE-ci-runner-2026-08-02.md`. The two that changed the design: **P2** (stock image has git but NO python3 → custom image `felhom-act-runner:0.1.0`, base pinned, python3 and nothing else) and **P6** (a runner that loses `/data/.runner` re-registers and leaves a dead record behind → the PVC is load-bearing, measured both ways). **P5 is the one that mattered**: a failed run produced NO mail, NO notification row and NO log line from Gitea, so the run now sends its own alarm via Resend and prints the provider's accepted id. **Proven end to end, not asserted**: a deliberately broken commit pushed with `--no-verify` → run #6 `failure` → `RESEND-ACCEPTED id=5ff34766-c5f8-4588-8104-08296aeb45ab`. Posture shown from the live pod spec: `privileged: false`, all caps dropped, no docker socket, no hostPath, `automountServiceAccountToken: false`, sized at half Gitea's limits so it cannot crowd out the service holding every repository on the same node. **The standing limit stays true and is written into the manifest and every workflow: it DETECTS, it does not BLOCK** — making it block is → R-169 | — | diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index b6163872..059e92fa 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,3 +1,32 @@ +## docs — v1.22.0 exercised end to end on two real reinstalls (2026-08-03, R-178) — **no script change** + +**Nothing shipped.** `felhom-host-install.sh` stayed at **v1.22.0**; the published copy at +`https://felhom.eu/scripts/felhom-host-install.sh` was confirmed byte-identical to the repo copy +(`sha256 ed02acb2da46c8d2b5c486ce99d5b9a2747e8786c6eb03652cf755ed1abdd9f4`) before use. Both demo +boxes were uninstalled and reinstalled with it, by **two deliberately different supply paths**: +demo-hp with `--golden ` (the `:2584` alternative), demo-felhom with +`--force-gitea-golden` (the canonical C.3 customer command). The merge-aware `step_grows` produced +`data +46G (->70G, ONE volume)` and `+226G (->250G)` respectively, and `fetch_verify` was observed +succeeding against the vouched manifest for **both** artifacts on demo-felhom +(`verified sha256 a7763d31b55b5ce7…` agent, `verified sha256 54e2a4c431daf580…` golden). + +**Two script-side findings, filed not fixed** (the session was a runbook; §7 forbade code): + +- **R-180** — `--archive-storage` is validated for existence (`:1583`) and for golden resolution + (`:1661`), but never against the ACL storage set it is about to grant (the fixed default + `local local-lvm felhom-pbs`). Staging the golden on `felhom-backup` therefore passed every + pre-flight gate and died at **step 8/8**: `HTTP 403: permission denied at /storage/felhom-backup + (missing privilege Datastore.AllocateSpace)` — *after* step 2 minted the token, step 4b **rotated + and vaulted root@pam**, and step 5 installed the agent. `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a + one-line assertion over two variables both known at `:1583`. +- **R-179** — `--uninstall` leaves the NAS network-storage systemd units behind + (`mnt-felhom\x2ddrives-.{mount,automount}`; automount left `failed`, parent bind left + mounted). The Part E residue-diff provenance is from **v1.9.1**, which predates the feature — and + demo-felhom, which never had a share configured, left nothing, which is exactly why a diff on such + a box reported clean. + +Full evidence: root `REPORT.md`. + ## host-install: one data volume, derived from the disk (2026-08-03, R-165) **Forced by a census, not planned.** `felhom-agent` v0.120.0 merges the appliance's two data volumes From 6b5d64c1fa7329b2aebbd90bbf8a0e32c196a1b1 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 09:35:16 +0200 Subject: [PATCH 06/13] =?UTF-8?q?REPORT:=20CI=20run=20ids=20and=20conclusi?= =?UTF-8?q?ons=20for=20both=20commits=20(46,=2047=20=E2=80=94=20both=20suc?= =?UTF-8?q?cess)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- REPORT.md | 20 ++++++++++++++++++-- 1 file changed, 18 insertions(+), 2 deletions(-) diff --git a/REPORT.md b/REPORT.md index b1cb4cf8..44c9e395 100644 --- a/REPORT.md +++ b/REPORT.md @@ -309,9 +309,25 @@ unused. --- -## 9. CI +## 9. CI — checked by run id, not assumed -Appended after the push — run id and conclusion, per `CLAUDE.md`'s pull-check rule. +Two commits, both docs-only. + +| repo | commit | CI run | result | +|---|---|---|---| +| `felhom-agent` | `9dfd89c` — *docs: agent 0.120.0 published + vouched…* | **run 46**, `head_sha 9dfd89cb` | **success** (`gates`) | +| `felhom.eu` | `aa62449` — *R-178 CLOSED: both demo boxes reinstalled…* | **run 47**, `head_sha aa624496` | **success** (`gates`) | + +Queried with +`curl -s "https://gitea.dooplex.hu/api/v1/repos/admin//actions/tasks?limit=3"` and matched on +`head_sha`, per `CLAUDE.md`'s pull-check rule — CI mails on failure, which is a push signal; this is +the pull check that catches a lost or unread mail. + +**`--no-verify` was NOT used**, anywhere. This clone is armed (`core.hooksPath = .githooks`). Both +gate entry points were also run by hand before committing: +`felhom.eu/scripts/repo_gates.py` → **all five OK** (site, hostinstall, hub-confirm, +manifest-bearer, reuse-refs), and `felhom-agent/scripts/agent_gates.py` → **OK** (reuse-refs). +No test suite was run because no code changed in either repo. --- From 06cbf8df29073f710375d039b55d0125c84ffee7 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 10:54:33 +0200 Subject: [PATCH 07/13] =?UTF-8?q?skill(build-deploy):=20the=20installer=20?= =?UTF-8?q?ISO=20=E2=80=94=20the=20one=20artifact=20the=20skill=20promised?= =?UTF-8?q?=20and=20omitted?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The skill's own description claimed 'ANY Felhom artifact' and 'publish', and had no ISO section — a description asserting coverage that did not exist. Description corrected and a section added. POINTERS, NOT COPIES. The 13-criterion release gate stays in documentation/runbooks/iso-release-gate.md and the measurements stay in the four spike audits; duplicating them into a skill guarantees drift (the R-94/R-128 class). What the skill adds is the ROUTING that was missing: nothing told anyone the gate exists, which is R-29's exact shape. Records the two modes (--release public vs --pairing appliance) because picking the wrong one ships the wrong product, the build and publish commands (rclone env-only, so no credential file is ever written), and the round-trip verification. The traps it carries existed only in commit messages until now, and each cost a wrong diagnosis: - 'qm set --scsi0 ... --boot order=' in ONE call silently yields boot: order=net0;ide2 - after install the CD must be detached, or a COMPLETED install looks exactly like a stuck one - verify focus by screendump before every Enter (GTK Enter lands in fields, not Next) - proof installs register appliances; the verb is POST /appliances//discard, not /delete - scratch storage at the /mnt/nvme-1tb ROOT (a subdirectory reads disconnected forever) Also flags that the ISO gate is NOT wired into repo_gates.py, so nothing reminds you to run it. Docs only. python3 scripts/repo_gates.py --fast: all gates OK (rc=0). --- skills/felhom-build-deploy/SKILL.md | 68 ++++++++++++++++++++++++++++- 1 file changed, 67 insertions(+), 1 deletion(-) diff --git a/skills/felhom-build-deploy/SKILL.md b/skills/felhom-build-deploy/SKILL.md index 135fedb2..6cdd2ab4 100644 --- a/skills/felhom-build-deploy/SKILL.md +++ b/skills/felhom-build-deploy/SKILL.md @@ -1,6 +1,6 @@ --- name: felhom-build-deploy -description: Build, deploy, publish, or verify ANY Felhom artifact — felhom-controller image (guest 9201 bootstrap deploy), felhom-agent binary (felhom-pve), felhom-hub (GitOps/ArgoCD), the felhom.eu website (git-sync), or the app catalog. Use whenever the task says build, deploy, ship, release, publish, bump version, restart the controller/agent/hub, or verify what version is live. Contains the exact verified commands and the gotchas that silently break deploys. +description: Build, deploy, publish, or verify ANY Felhom artifact — felhom-controller image (guest 9201 bootstrap deploy), felhom-agent binary (felhom-pve), felhom-hub (GitOps/ArgoCD), the felhom.eu website (git-sync), the app catalog, or the PUBLIC installer ISO (iso.felhom.eu). Use whenever the task says build, deploy, ship, release, publish, bump version, restart the controller/agent/hub, or verify what version is live. Contains the exact verified commands and the gotchas that silently break deploys. --- # Felhom build & deploy runbooks @@ -75,6 +75,72 @@ Publish to Gitea (so Day-0 self-install can fetch it): `scripts/publish-agent.sh `REGISTRY_*` creds. The hub's Day-0 artifact manifest must then vouch the new version — that UI is operator-password-gated (CC cannot); flag it as an operator follow-up. +## Installer ISO (felhom.eu/scripts/iso → iso.felhom.eu) — PUBLIC, irreversible + +**Two modes, and picking the wrong one ships the wrong product.** + +| Mode | What it is | Menu | +|---|---|---| +| `--release` | the **public** image. No `answer.toml`, no root password, no SSH key, no disk profile. Day-0 rides a `.deb`. | TWO interactive entries, graphical default, timeout 15s | +| `--pairing` / `--bootstrap-env` | operator-built **appliance** image: baked answer file, baked root hash, a disk profile pinned to one machine | ONE automated entry | + +```bash +# build the public image (clean-tree gate first — an unpushed change does not exist) +export FELHOM_ISO_OUT=$FELHOM_ROOT/felhom-iso/out +bash scripts/iso/build-felhom-iso.sh \ + --pve-iso $FELHOM_ROOT/drill/proxmox-ve_9.2-1.iso \ + --iso-sha256 4e88fe416df9b527624a175f24c9aa07c714d3332afb1ee3dbf3879573ef2c6c --release +# -> $FELHOM_ISO_OUT/felhom-installer--pve.iso + .sha256 + .manifest.txt (NO .rootpw.txt) +``` + +**Before publishing, two hard gates — neither is optional and neither is a script yet:** + +1. **`documentation/runbooks/iso-release-gate.md`** — 13 criteria, run against **the exact file you + will upload**, not the build inputs. It is a manual checklist; `repo_gates.py` does **not** cover + it, so nothing will remind you (R-29's shape — say so if you skip it). +2. **A proof install from the built image on BOTH menu entries** (graphical *and* Terminal UI), each + showing: package installed, unit `enabled`, unit fired on first boot, and a pairing code in + `/etc/felhom/appliance-pairing-code`. Spike 4 *reasoned* the graphical path follows from shared + `Install.pm`; the 1.26.0 run proved that reasoning insufficient in a different place — do both. + +```bash +# publish — rclone in a container, env-only config, so NO credential file is ever written +source ~/.config/credentials # ISO_S3_CLIENT_AK / _SK / ISO_S3_URL — never echo, never log +docker run --rm -v $FELHOM_ISO_OUT:/data:ro \ + -e RCLONE_CONFIG_R2_TYPE=s3 -e RCLONE_CONFIG_R2_PROVIDER=Cloudflare \ + -e RCLONE_CONFIG_R2_ACCESS_KEY_ID="$ISO_S3_CLIENT_AK" \ + -e RCLONE_CONFIG_R2_SECRET_ACCESS_KEY="$ISO_S3_CLIENT_SK" \ + -e RCLONE_CONFIG_R2_ENDPOINT="$ISO_S3_URL" \ + -e RCLONE_CONFIG_R2_REGION=auto -e RCLONE_CONFIG_R2_NO_CHECK_BUCKET=true \ + rclone/rclone:latest copy /data R2:felhom-iso --include "felhom-installer-*" --s3-chunk-size 64M +# verify by ROUND TRIP — the downloaded bytes, not the local file +curl -fsSL -o /tmp/rt.iso https://iso.felhom.eu/felhom-installer--pve.iso && sha256sum /tmp/rt.iso +``` + +`ListBuckets` 403s — the token is object-scoped; list with `lsf R2:felhom-iso`, not `lsd R2:`. + +### Proof-VM traps — every one of these cost a wrong diagnosis + +- **Set `--boot` in a SEPARATE `qm set`, after the disk exists.** `qm set --scsi0 … --boot order="scsi0;ide2"` + in one call silently yields `boot: order=net0;ide2`; the VM netboots, fails, falls through to the CD. +- **After the install, detach the CD** (`qm set --delete ide2; qm set --boot order="scsi0"`) + or the machine re-enters the installer on reboot — **a completed install looks exactly like a stuck one.** + Judge completion from `qm config` + disk usage, never from the screen. +- **Verify focus by screendump before every `Enter`.** TUI: red-highlighted button, tab order. GTK: + dashed focus ring, and `Enter` lands in text *fields*, not `Next`. Not checking once aborted an install. +- **Proof installs register unclaimed appliances at the hub — discard them** or R-131 grows: + `curl -u ":$HUB_PW" -X POST http://:8080/appliances//discard` → 303. + The verb is **`/discard`**, POST only (`hub/internal/web/server.go:345`); `/delete` 404s. +- Venue: `demo-hp`, scratch `dir` storage at **`/mnt/nvme-1tb` root** (a subdirectory reads + `disconnected` forever — the agent's `exactMount` check). Never `local-lvm`. Remove the storage at teardown. + +**Why the shape is what it is** (do not re-derive; four spikes measured it): +`documentation/audits/SPIKE-universal-iso-{1,2,3,4}-2026-07-31.md`. In short — no udev property +distinguishes an internal disk from a customer's backup drive and a two-disk filter match silently +wipes one, so **there is no safe automated disk selection for unseen hardware**; and `[first-boot]` +is never placed on the system by an interactive install, so day-0 rides a `.deb` in +`/proxmox/packages/` instead (`Install.pm:1343-1372`). + ## Hub (felhom.eu/hub → k3s, GitOps via ArgoCD app `felhom`) **The manifest is the truth.** A code push + image build deploys NOTHING until `manifests/hub.yaml`'s From fb652024eaf87819a4d1f2e5ccb98dc0aaf2702c Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 11:36:16 +0200 Subject: [PATCH 08/13] docs: R-181 closed, R-156 closed, R-110 + R-115 rulings recorded, R-182 filed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit R-181 CLOSED (controller v0.193.0 + v0.193.1) and proven live on demo-hp for BOTH reserve terms. The reserve is now a per-app, per-run ADMISSION decision taken before the app's first write and covering all three write legs, and it gained a size term. The refusal's wording was not weakened; the behaviour moved so it became true, verified by sha256 tree fingerprint. R-156 CLOSED — papra's template mounts the app's own data root. Precondition re-measured rather than inherited (both boxes were wiped today). Part 4, documentation only, nothing built: - R-110 WAITING-ON-OPERATOR -> READY. Ruling: option (b), the installer's publish channel moves to a TAG. Recorded with the condition that decides whether it works at all — it must cover BOTH the /scripts/ git-sync and the nine files the installer fetches from raw/branch/main. - R-115 WAITING-ON-OPERATOR -> READY. Ruling: mechanism (b), a build-side gate refusing to deploy or vouch an unpublished version. The third instance (agent v0.120.0) would have silently downgraded both demo boxes while succeeding. R-182 NEW: the periodic status refresh has no admission scope, so a refused app re-alerts on every poll (measured: a second alert pair 13s after the run's). Pre-existing in v0.192.0; deliberately not fixed in the R-181 task. capability map: the local-backup row moves to PROVEN-LIVE in BOTH halves. ROADMAP: R-165 collapses to CLOSED; R-181 collapsed into it. 07-backup-architecture.md: the reserve's contract stated as what the code provides (S-1 — an architectural contract changed in the same session). STATUS.md trimmed 150 -> 111 lines, "What's broken" no longer holds shipped work, and the stale "After:" line (pointing at work that shipped on 2 August) is fixed. --- CONTEXT.md | 40 +- REPORT.md | 532 +++++++----------- STATUS.md | 145 ++--- .../architecture/00-capability-map.md | 2 +- .../architecture/07-backup-architecture.md | 37 +- documentation/backlog/OPEN-ITEMS.md | 9 +- documentation/backlog/ROADMAP.md | 8 +- 7 files changed, 324 insertions(+), 449 deletions(-) diff --git a/CONTEXT.md b/CONTEXT.md index da6085e2..6fd56751 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -60,10 +60,42 @@ demo boxes were wiped and reinstalled from golden 0.192.0 and taken through clai proved the disk shape *and* the delivery route rather than one of them twice. Live shape on both: `mp0` at `/var/lib/felhom`, `backup=1`, **no `mp1`**; `/var/lib/docker` and `/mnt/sys_drive` both real mounts of its subdirectories via `/etc/fstab`; ONE `df` figure and one device id on all three paths; -3/3 reboots each with the binds surviving every time. **What is NOT proven is B2** → **R-181**: the -floor guards `captureAllRecoveryUnits` and not `runVolumeDumps`, which is the leg that fills the -volume, and its refusal message's "the previous unit is untouched" was measured false. R-165 is -therefore **IMPLEMENTED**, not PROVEN-LIVE. +3/3 reboots each with the binds surviving every time. B2 was **not** proven on that pass → **R-181**: +the floor guarded `captureAllRecoveryUnits` and not `runVolumeDumps`, the leg that fills the volume, +and its refusal's "the previous unit is untouched" was measured false. **R-181 CLOSED the same day +(controller v0.193.0 + v0.193.1), so R-165 is now PROVEN-LIVE in both halves** — see S-14. + +**S-14 — the reserve is a per-app, per-run ADMISSION decision, not a capture check (2026-08-03, R-181; +controller v0.193.0 + v0.193.1).** B2 as first shipped was consulted in exactly one place — +`captureAllRecoveryUnits`, a few KB — while `RunDBDumps`' database leg and `runVolumeDumps` wrote the +bulk into the same `backups/primary/` tree, first and unguarded. The reserve was therefore +consumed by the very write it exists to bound, and the refusal then claimed *"the previous unit is +untouched"* about a tree the earlier leg had already rewritten (182,272 B → 2,147,666,432 B under a +manifest that had not moved). **Sixth entry in `CLAUDE.md`'s table of shipped guarantees the code did +not provide, and the fourth of those found on live hardware rather than by review.** + +- **`internal/backup/admission.go` — `admitApp` is now THE gate**, and every per-app write leg calls + it. One verdict per app per run covers all three; they share one per-app root, which is what makes + that honest. +- **Decided lazily at the app's first write, never once at run start** (app A's dump can put app B + under the reserve), **never re-decided between an app's own legs** (that is the split it closes), + and **reset per run**. +- **Ahead of `DumpAppVolumesSafe`**, which stops the stack as its first act — a refusal decided + inside it has already bounced the app. **After** the volume-less check, which has no write to gate. +- **Size term added:** *would THIS app's write cross the reserve?*, estimated from the app's previous + `.sql` + `.tar`. **No history → headroom-only**, deliberately — otherwise the first backup is the + one that can never happen. +- **A container-based `du` was MEASURED and rejected**, not waved away: median **~355 ms/volume** over + 66 runs on demo-hp, on volumes holding tens of KB (container start-up, not the walk). Decisive on + top: `docker run` needs the writable layer, so the instrument can fail under exactly the pressure + the reserve handles. +- **The wording was NOT weakened; the behaviour moved so it became true**, and it is checked by + sha256 tree fingerprint, never by reading the log line — the log line is what lied. +- **v0.193.1**, found by the proof run itself: a 178 KB estimate printed as `0.00 GiB`, which reads as + *no estimate available*. Rendering moved to `humanizeBytes`; arithmetic still in GiB. +- **New finding, deliberately not fixed here → R-182**: `GetFullStatus`'s periodic capture sweep has + no run scope, so a refused app re-alerts on every status refresh (measured: a second identical alert + pair 13 s after the run's). Pre-existing in v0.192.0; R-181 changed neither caller. **S-11 — D-c's routing, and why R-158's own proposal was overruled (2026-08-02, R-167 SHIPPED).** Decision D-c splits two signals by AUDIENCE, and the split is the ruling: **a fill warning is the diff --git a/REPORT.md b/REPORT.md index 44c9e395..196ed85d 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,358 +1,216 @@ -# REPORT — R-178: both demo boxes reinstalled from the merged golden and proven (2026-08-03) +# REPORT — R-181 (the reserve guards the write that fills the disk) + R-156 (papra) + two operator rulings -**Overwritten** per the standing rule. The prior contents (R-165, the bake, 2026-08-03) have their -durable record in `documentation/backlog/OPEN-ITEMS.md` R-165 and the per-repo CHANGELOGs. +**Date:** 2026-08-03 · **Repos:** `felhom-controller` (v0.192.0 → **v0.193.1**), `app-catalog-felhom.eu`, `felhom.eu` (docs only — **no hub change, no hub version bump**) -**Runbook, not a task.** No repo got a version bump and nothing was built. One artifact was -**published** (agent 0.120.0) on an explicit operator ruling — see §2. All code findings are filed as -register rows, not commits, per the runbook's §7. +## 1. Baselines — re-read on arrival, all matched §1 ---- - -## 1. Preconditions P1–P6, each as measured - -| # | Precondition | Measurement | -|---|---|---| -| **P1** | Golden sha256 matches R-178's record | **PASS** — computed from the artifact itself on demo-hp: `sha256sum` → `54e2a4c431daf580d2807b82d810be36a7c6a9f697b27094dc2912d26a43b3e0`, 649,547,835 B. Matches R-178's recorded prefix and the hub manifest's stored value exactly. Re-verified after copying to `local`: identical | -| **P2** | A rollback golden still exists | **PASS, three depths.** Split-layout golden **0.188.0** fetchable from Gitea (`HTTP 200`, 649,310,288 B). A **local** split-layout golden sat on each box — `local:backup/vzdump-lxc-9100-2026_07_21-18_24_49.tar.zst` (demo-hp), `…2026_07_20-17_50_57…` (demo-felhom). Best: **full guest vzdumps**, three per box, newest `vzdump-lxc-9201-2026_08_03-04_54_28.tar.zst` (2,351,557,870 B, demo-hp) and `…2026_08_03-04_44_50…` (6,323,506,165 B, demo-felhom), both ~3 h old — so each box could be put back exactly as it was | -| **P3** | Neither demo box holds anything wanted | **PASS, stated explicitly.** demo-hp: paperless-ngx (webserver/postgres/redis) + filebrowser + samba shares. demo-felhom: immich (×4), docmost (×3), calibre-web, bookstack (×2) + filebrowser + samba. Both are demo customers (`Demo HP`, `Demo Ügyfél`); no real customer data. **All of it was destroyed by the wipes and none of it was restored** — that was the point, and the operator confirmed each wipe | -| **P4** | The colleague's box untouched | **PASS** — `peti-felhom` is Tier 2 and out of scope. No command was sent to it; it appears in the hub customer list as `DOWN`, exactly as before | -| **P5** | Both hosts' agent is v0.120.0 | **PASS** — `felhom-agent 0.120.0` on both, before and after | -| **P6** | The hub's host register, BEFORE | **Captured**: `demo-felhom-8363b5 / Demo Ügyfél / 0.120.0 / ONLINE / 1-of-2 guests`; `demo-hp-bb76ea / Demo HP / 0.120.0 / ONLINE / 1-of-2`; `drill-r50-0a4f9a / drill-r50 / 0.113.0 / DOWN` (the fixture, untouched throughout) | - ---- - -## 2. Two findings that contradicted the runbook's premise, both surfaced before any wipe - -**(a) The golden was ALREADY vouched.** Hub log, a positive observable: -`2026/08/03 07:23:26 [INFO] Artifact manifest set: agent=0.119.0 golden=0.192.0 min_agent="0.113.0" -wrapper_sha=true` — roughly ten minutes before this session's first hub read, and this session had -POSTed nothing. The configuration page confirmed `Artifacts.GoldenVersion=0.192.0`, -`GoldenSHA256=54e2a4c4…3b3e0`. So §7's prove-then-vouch order was already spent, and Phase A's stated -safety ("nothing is official yet, so a failure reaches nobody") was void. **Operator ruling: accept it -and proceed**, keeping the two different supply paths. Recorded on `CONTEXT.md` S-14, now marked SPENT -with the reason — the rule lived only in prose and nothing enforced it. - -**(b) Agent v0.120.0 had never been published — the vouched agent was 0.119.0.** -`GET …/generic/felhom-agent/0.120.0/felhom-agent` → **HTTP 404** (0.119.0 → 200). Installer step 5's -idempotent skip requires `installed == vouched` exactly, so a documented-path reinstall would have -**downgraded both boxes** to the pre-merge 0.119.0 — and would have *succeeded* while doing it, since -the current `step_grows` sets `SYSDATA_GROW=0` and 0.119.0's `mp1` resize (`bringup.go` 4c, fatal on -error) therefore never fires. The session would have proven a stack nobody ships. **Operator ruling: -publish and vouch first.** `scripts/publish-agent.sh 0.120.0` from a clean tree at `4bb84fc3` -(`git status --porcelain` empty, `HEAD == origin/main`): upload HTTP 201, **round-trip GET verified**, -`AGENT_SHA256=a7763d31b55b5ce75457b4dba7b06aa300325811834b0be78af4587b47110b9d`. Vouched via -`POST /configuration/artifacts` → `303 …flash=artifacts_set`, hub log -`Artifact manifest set: agent=0.120.0 golden=0.192.0`; the hub resolved the sha authoritatively from -Gitea rather than trusting the submitted value. Filed as the **third instance of R-115**, not a new -ID — R-115 is that row, and it has been WAITING-ON-OPERATOR since 2026-07-29. - ---- - -## 3. demo-hp — the layout proof - -**Install path.** Uninstall: `./felhom-host-install.sh --uninstall --vmid 9201` (typed-vmid -confirmation supplied over a pty). Install: - -``` -./felhom-host-install.sh --customer-id demo-hp --mode appliance --vmid 9201 \ - --cores 7 --memory 26906 \ - --golden local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst \ - --passphrase-file /root/.pp-demo-hp -``` - -Script fetched from `https://felhom.eu/scripts/felhom-host-install.sh`, **v1.22.0**, sha256 -`ed02acb2da46c8d2b5c486ce99d5b9a2747e8786c6eb03652cf755ed1abdd9f4` — byte-identical to the repo copy -that was read. The passphrase went file→file into a 0600 file and never onto a command line. A -`--dry-run` preceded the real run and resolved `manifest: agent v0.120.0 (sha a7763d31…), golden -v0.192.0` with `grows: rootfs +0G (->32G), data +46G (->70G, ONE volume)`. - -**One deviation, mine, and it cost a restart.** The first attempt staged the golden on -`felhom-backup` with `--archive-storage felhom-backup`. Pre-flight passed; **step 8/8** failed: -`HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)` — -the default pool-scoped ACL grants `local local-lvm felhom-pbs` only. Fixed by copying the golden to -`local` (sha re-verified after the copy) and `--resume`. Filed as **R-180**: the condition is -statically checkable in pre-flight, and the failure lands *after* step 4b has rotated and vaulted -root@pam. - -**Layout evidence.** - -``` -mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=70G # and NO mp1 line -rootfs: local-lvm:vm-9201-disk-0,size=32G - -/var/lib/felhom /dev/mapper/pve-vm--9201--disk--1 ext4 rw,relatime,stripe=16 -/var/lib/docker /dev/mapper/pve-vm--9201--disk--1[/docker] ext4 -/mnt/sys_drive /dev/mapper/pve-vm--9201--disk--1[/sys_drive] ext4 - -Filesystem Size Used Avail Use% Mounted on -/dev/mapper/pve-vm--9201--disk--1 69G 977M 65G 2% /var/lib/felhom -/dev/mapper/pve-vm--9201--disk--1 69G 977M 65G 2% /var/lib/docker -/dev/mapper/pve-vm--9201--disk--1 69G 977M 65G 2% /mnt/sys_drive - -stat -c %d → 64519 for all three paths -/etc/fstab: /var/lib/felhom/docker → /var/lib/docker ; /var/lib/felhom/sys_drive → /mnt/sys_drive -both binds writable (touch succeeded on each) -``` - -`/mnt/sys_drive` shows a second `findmnt` row — it is the controller container's -`-v /mnt:/mnt:rslave` propagation (peer group 383 vs the fstab bind's 333), the same shape the split -layout had, not a stacked bind. - -**Reboots — each individually, minimum three:** - -| # | started | controller healthy | outcome | +| Repo | `main` @ arrival | Version | Shipped | |---|---|---|---| -| 1 | 08:13:33 | 08:13:55, `Up 12 seconds (healthy)` | one filesystem, 69G/65G on both paths | -| 2 | 08:13:59 | 08:14:15, `Up 7 seconds (healthy)` | same | -| 3 | 08:14:19 | 08:14:35, `Up 7 seconds (healthy)` | same | +| `felhom-controller` | `4be6467b501b` | v0.192.0 | **v0.193.0 `fef07c3`** → **v0.193.1 `6c43bf6`** | +| `app-catalog-felhom.eu` | `7cb58ecdf8e7` | n/a | `122bbee` | +| `felhom.eu` | `6b5d64c1fa73` | hub v0.89.0 | docs only, **no bump** | -After all three, `mountpoint -q` returns true for **all three paths** and both binds still resolve to -the single volume's subdirectories — which is what the reboots exist to test. `uptime -s` = -`2026-08-03 06:14:24 UTC`, matching reboot 3, so these were real reboots. +All three clean (`git status --porcelain` empty, `HEAD == origin/main`) before every build. -**The journey — method stated: endpoint-level, not a browser.** `claude-in-chrome` does not exist on -DooPlex; every step below invoked the exact endpoint the dashboard's own JavaScript calls. +## 2. The fix -| Leg | Endpoint | Observable | -|---|---|---| -| Claim | `POST /claim` with the pre-auth HMAC CSRF token (64 chars) **and** its `felhom_claim_csrf` cookie | `302 → /`; gate discriminator flipped `{"error":"dashboard not yet claimed"}` → `{"error":"authentication required"}`. **The code is emailed-only (R-119) — the operator supplied it**, after an operator resend rotated the generation (the previous code had been consumed 12 d earlier) | -| Deploy | `POST /api/stacks//deploy` with `{"values":{…}}`, field values scraped from the deploy form exactly as the form's own `fetch` does | `{"ok":true}`; `privatebin Up (healthy)`, then `opengist Up (healthy)` | -| Back up | `POST /api/debug/backup/dbdump` — this runs the **production** `RunDBDumps` path (DB leg → `runVolumeDumps` → `captureAllRecoveryUnits`); the debug route only starts it instead of waiting for 03:30 | `Volume dump: opengist/… → 178.0 KB`, `privatebin/… → 2.5 KB`, `App-data backup completed … 2 volume dump(s) (3.962s)`, `Recovery unit captured for …` ×2 | -| **Restore** | `POST /backup/restore` `stack_name=privatebin snapshot_id=primary` | A marker planted in the live volume was **deleted**, then restored: `{"ok":true,"message":"privatebin visszaállítva (primary)."}` in **9.2 s**, and the marker returned with an **identical sha256 `ac1faae6134871d640d5cb1bbc6b5092d1e96b7e46c6c3bcb6e8236cc33ae861`**. App `Up (healthy)` afterwards | +**One admission verdict per app per run** (`controller/internal/backup/admission.go`), taken before +that app's **first** write and consulted by all three legs — DB dump, volume dump, unit capture. The +three write under one per-app root (`appbackup.RecoveryUnitPath`), which is what makes one verdict +able to cover them honestly. -**Recovery unit path, and it lands on the single volume:** -`/mnt/sys_drive/felhom-data/backups/primary//{manifest.json,compose/,volume-dumps/}`, whose `df` -is `/dev/mapper/pve-vm--9201--disk--1` — the merged volume. +- **Lazy, not run-wide.** App A's dump can put app B under the reserve; a run-start verdict reads a + disk that no longer exists. **Never re-decided between an app's own legs** — that is the split being + closed. **Reset per run.** +- **Ahead of `DumpAppVolumesSafe`**, which stops the stack as its first act, so a refused app is never + bounced. **After** the volume-less check, which has no write to gate. +- **Exactly one operator alert per refused app per run.** Leg order unchanged. +- **Size term added:** *would this app's write cross the reserve?* — estimated from its previous + `.sql` + `.tar`. **No history → headroom-only**, or the first backup becomes the one that can never + happen; the alert says so when that applies. -**Note on the first app chosen.** `privatebin`'s catalog entry declares no `backup:` section, so its -first per-stack capture produced `"volume_dumps": null` — correct for that declaration, not a defect, -but it means a per-stack `POST /stacks//backup` writes compose+config only; the volume leg lives in -the full pass. A second app (`opengist`) was deployed so the run had both a capture and, later, a -refusal. +## 3. Files -**The ceiling is gone, measured:** a recovery unit can use **65 GiB** — the whole volume — against the -**19 GiB** the pre-wipe `mp1` slice offered (`/dev/mapper/pve-vm--9201--disk--2 20G 95M 19G 1%`). - ---- - -## 4. The floor's first live firing — and it does not do what it says - -**Instrument, proven before use.** demo-hp's thin pool is **53.93 GiB** and the auto-sized volume is -70 G, so a real fill to 97 % would have exhausted the pool and corrupted every guest on the box, -including the protected `drill-r50` fixture. A 5 GiB `fallocate` probe moved guest `df` from -`977M used / 65G avail` to `6.0G / 60G` while thin-pool `data_percent` stayed **29.03 → 29.03** — -zero blocks allocated — and cleanup returned both to baseline. The floor reads `statfs`, which is -exactly what `fallocate` moves, so the condition it guards is genuinely present. - -**Setup.** A **real** 2 GiB file in opengist's data volume (so its capture writes 2 GiB), then -`fallocate` to leave `69G / 63G used / 3.0G avail / 96%` — **both floor terms deliberately still -clear**, so the run would start. Capture order was established empirically from the previous run's log -(opengist first, privatebin second), not assumed. - -**What happened, 06:40:03:** - -``` -Volume dump: opengist/opengist_opengist_data → 2.0 GB # unguarded -Volume dump: privatebin/privatebin_privatebin_data → 2.5 KB -App-data backup completed … 2 volume dump(s) (20.737s) -[WARN] Recovery unit capture REFUSED for opengist — … 1.0 GB free; the previous unit is untouched and NOTHING was deleted -[WARN] Recovery unit capture REFUSED for privatebin — … 1.0 GB free; the previous unit is untouched and NOTHING was deleted -[INFO] Event pushed: recovery_unit_capture_failed (error) — … ×2 (HTTP 200) -``` - -**What holds:** it refuses **per app** rather than aborting the run; **nothing was deleted** (both -units present afterwards); and the operator alert reached the hub — `recovery_unit_capture_failed`, -severity `error`, accepted `HTTP 200`, twice. - -**What does not hold — measured, not inferred.** The refusal message claims *"the previous unit is -untouched"*: - -| unit file | before | after | -|---|---|---| -| `privatebin/volume-dumps/privatebin_privatebin_data.tar` | `26c546c2…` | **`b538ab89…`** | -| `opengist/volume-dumps/opengist_opengist_data.tar` | 182,272 B | **2,147,666,432 B** | - -Both were rewritten by the earlier leg, while each `manifest.json` kept -`created_at: 2026-08-03T06:34:26Z` and its `checksums` block covers only the three compose files — so -a unit's payload can be swapped under a stale descriptor and nothing inside the unit can detect it. - -**Cause, in the code and not the log.** The floor is consulted in exactly one place — -`m.unitFloorBlocked(stack.Name)` at `recovery_unit.go:328`, inside `captureAllRecoveryUnits`, which -writes a manifest and a compose copy: a few KB. `runVolumeDumps` (`backup.go:535`) — the leg that -writes the bulk, and the leg that consumed the reserve — has **no floor check at all**; its gates are -protected-stack, volume-less, disconnected, decommissioned. And it runs first *by design* -(`backup.go:483`). **The floor guards the cheap leg and not the leg that fills the volume.** - -This is why R-165 is IMPLEMENTED and not PROVEN-LIVE: that row records B2 as the deliberate -replacement for the bulkhead the `mp1` partition provided, and pre-merge the unguarded leg could only -fill a dedicated 20 G partition — post-merge it can reach Docker's data-root. Filed as **R-181**. -**No code was written**, per the runbook's §7. - -**Cleanup:** fill file and payload removed; `df` back to `1.2G used / 65G avail`; a clean re-run left -both units valid (`Volume dump … 178.0 KB` / `2.5 KB`, `App-data backup completed … (2.399s)`). - ---- - -## 5. demo-felhom — the pipeline proof - -**Install path — deliberately different, and this is the reason the second box exists.** - -``` -./felhom-host-install.sh --customer-id demo-felhom --mode appliance --vmid 9201 \ - --cores 3 --memory 12288 \ - --force-gitea-golden \ - --passphrase-file /root/.pp-demo-felhom -``` - -A copy of golden 0.192.0 already sat on this box's `local` storage — it is where the golden was -**baked** at 06:58 (a `.log` beside it), sha `54e2a4c4…`. `--force-gitea-golden` is the documented -C.3 customer path and overrides local discovery in **both** pre-flight and step 7, which pre-flight -confirmed: `golden: none local — will fetch + verify from Gitea in step 7/8`. The bake artifact was -left untouched. - -**`fetch_verify` succeeding against the vouched sha — the observable this box exists to produce:** - -``` -5/8 fetching agent binary v0.120.0 from Gitea … - verified sha256 a7763d31b55b5ce7… matches the hub manifest -7/8 fetching golden v0.192.0 from Gitea → /var/lib/vz/dump/vzdump-lxc-9100-2026_08_03-09_15_34.tar.zst - verified sha256 54e2a4c431daf580… matches the hub manifest - golden imported + verified: local:backup/vzdump-lxc-9100-2026_08_03-09_15_34.tar.zst -Day-0 provision SUCCESS — vmid=9201 host_id=demo-felhom-8363b5 customer=demo-felhom -``` - -Both artifacts — agent and golden — were fetched anonymously (the normal customer shape) and -sha-verified against the manifest. Controller `0.192.0` healthy. - -**Layout:** `mp0: local-lvm:vm-9201-disk-1,mp=/var/lib/felhom,backup=1,size=250G`; -`pct config 9201 | grep -c '^mp1:'` → **0**. Both binds real mounts -(`…disk--1[/docker]`, `…disk--1[/sys_drive]`), `df` one figure — `246G 977M 233G 1%` on all three -paths — `stat -c %d` = `64519` on all three, fstab carries both binds. - -**Reboots:** 1 — 09:18:55 → healthy 09:19:09; 2 — 09:19:12 → 09:19:26; 3 — 09:19:30 → 09:19:45. All -three paths still mountpoints after each; `uptime -s` = `2026-08-03 07:19:34 UTC`, matching reboot 3. - -**Journey:** claim (`302 → /`, discriminator flipped to `authentication required`; code supplied by -the operator after a resend) → deploy `opengist` (`{"ok":true}`, `Up (healthy)`) → capture -(`Volume dump: opengist/… → 178.0 KB`, `Recovery unit captured`, unit on -`/dev/mapper/pve-vm--9201--disk--1`, marker present inside the tar) → **restore** -(`{"ok":true,"message":"opengist visszaállítva (primary)."}` in **9.4 s**, marker back with identical -sha256 `bc5507987f3f56dc19a9c24785826c304f1a986998d4aeb0b06d1a810b59e939`, app healthy). - -**Ceiling:** 233 GiB available to a recovery unit, against the 45 GiB the pre-wipe `mp1` offered. - -**Step 8 was NOT repeated here, deliberately** — stating it rather than leaving it ambiguous. The -floor fired on demo-hp and R-181 characterises it fully; re-firing would add no information and would -mean filling a 246 G volume. - ---- - -## 6. Vouching - -Already done before the session (§2a), at **07:23:26 CEST on 2026-08-03**, by an operator action in -the hub UI. The session's own manifest write was the **agent** half, at **07:44:31**: -`Artifact manifest set: agent=0.120.0 golden=0.192.0 min_agent="0.113.0" wrapper_sha=true`. The -manifest afterwards, read back: `agent_sha256=a7763d31b55b5ce7…10b9d`, -`golden_sha256=54e2a4c431daf580…43b3e0`, `min_agent=0.113.0`, `wrapper_sha256=104db0a4…` (preserved -verbatim). The R-120 gate did not block: golden 0.192.0 equals the newest controller the fleet -reports. - ---- - -## 7. Teardown — all three layers - -1. **The machine.** No throwaway guest was created this session, so there is none to delete. VM 300 - `drill-r50` on demo-hp — the protected drift fixture — was never touched and is still `stopped`; - it is not in the `felhom` pool (`pvesh get /pools/felhom` listed only `lxc/9201`), so the - uninstall's shared-box logic never reached it. -2. **The host.** `local-lvm`, before → after: **demo-felhom 29.31 % → 1.37 %** (the old 200 G + 50 G - volumes returned; the new 250 G volume is thin and barely allocated). **demo-hp 39.13 % → - 36.75 %** — *higher than a clean reinstall would leave it*, because the floor test's 2 GB tar - blocks cannot be reclaimed: `fstrim` inside an unprivileged LXC returns - `FITRIM ioctl failed: Operation not permitted`. No operational impact (the guest shows 65 G free of - 69 G, the host 35.7 GiB free of 53.93), but it is real residue and is stated rather than rounded - away. Each box's data volume is now the only data volume; no old guest volumes remain. - **Residue found and cleared by hand on demo-hp:** `--uninstall` left the NAS network-storage units - `mnt-felhom\x2ddrives-Felhom\x2dShare.{mount,automount}` behind (automount `failed`, parent bind - still mounted) → **R-179**. demo-felhom left none, because it had no network share configured. -3. **The hub.** **No old records exist to dispose of, and this is the honest finding, not an - omission:** both enrollments were **idempotent** — `host REUSED (idempotent — existing - credential)`, `host_id: demo-hp-bb76ea` and `demo-felhom-8363b5`, the same ids as before. The - reinstalls therefore produced **no new host records**, so nothing was orphaned and nothing needed - deleting. Final register: `demo-felhom-8363b5 ONLINE 0.120.0`, `demo-hp-bb76ea ONLINE 0.120.0`, - `drill-r50-0a4f9a DOWN 0.113.0` — the same three rows as at P6. No scratch customers were created. - `/appliances` returns 404 on hub 0.89.0 — there is no appliance-record surface to clean. - -**Secrets:** both retrieval passphrases were moved file→file into 0600 files, used via -`--passphrase-file`, and `shred -u`'d afterwards on both hosts along with the session cookie files and -helper scripts; the local scratch copies were deleted. Nothing was written to a committed file and -`curl -w '%{redirect_url}'` was never used (R-132). - ---- - -## 8. Registers changed - -| Row | Change | +| File | | |---|---| -| **R-178** | **CLOSED** — both boxes reinstalled and proven, by two different supply paths | -| **R-165** | **IMPLEMENTED**, not PROVEN-LIVE — operator ruling; the layout half is proven, the B2 half is not (→ R-181) | -| **R-115** | **Third instance recorded** — agent 0.120.0 built, deployed to both hosts, never published | -| **R-181** *(new)* | The capture floor guards the recovery-unit leg and not `runVolumeDumps`; its "previous unit is untouched" claim measured false | -| **R-180** *(new)* | `--archive-storage` is not cross-checked against the ACL grant; the 403 lands at step 8/8, after root@pam has been rotated | -| **R-179** *(new)* | `--uninstall` leaves NAS network-storage systemd units behind when a share was configured | +| `controller/internal/backup/admission.go` | **new** — the gate, the memo, the estimator | +| `controller/internal/backup/admission_test.go` | **new** — 11 tests | +| `controller/internal/backup/backup.go` | run scope + gates in the DB and volume legs | +| `controller/internal/backup/recovery_unit.go` | `floorVerdict` size-aware; capture leg via `admitApp` | +| `controller/internal/backup/capture_floor_test.go` | 3 call sites updated for the new signature | +| `controller/README.md`, `REUSE.md`, `CHANGELOG.md` | | +| `app-catalog-felhom.eu/templates/papra/docker-compose.yml` | mount moved to `/app/app-data` | -**IDs established free before minting:** `grep -ro "R-179\b\|R-180\b\|R-181\b\|R-182\b"` over -`documentation/` and `*.md` → **0 hits**, and over all four repo roots (`felhom-agent`, -`felhom-controller`, `felhom.eu`, `app-catalog-felhom.eu`) → **0 hits**. `R-182` was checked and left -unused. +## 4. Tests — 28 packages `ok`, `rc=0` (read separately from any commit) ---- +All 11 new tests pass, plus the pre-existing floor suite. Refusal assertions are **sha256 tree +fingerprints before and after**, never log lines — the defect being fixed *is* a log line the tree +contradicted. -## 9. CI — checked by run id, not assumed +The DB leg cannot run without Docker (`DiscoverDatabases` shells out), so its gate is pinned by an +**AST walk** of `backup.go` asserting `admitApp` precedes `DumpOne`. `strings.Contains` is +insufficient: a commented-out call still contains the string. -Two commits, both docs-only. +### Red-proofs — each demonstrated failing, then restored -| repo | commit | CI run | result | -|---|---|---|---| -| `felhom-agent` | `9dfd89c` — *docs: agent 0.120.0 published + vouched…* | **run 46**, `head_sha 9dfd89cb` | **success** (`gates`) | -| `felhom.eu` | `aa62449` — *R-178 CLOSED: both demo boxes reinstalled…* | **run 47**, `head_sha aa624496` | **success** (`gates`) | +| # | Mutation | Result | +|---|---|---| +| 1 | **Both** dump-leg `admitApp` gates removed (= exactly v0.192.0) | Scenario A **RED** — *"the VOLUME leg ran for a refused app"*; with the leg assertions temporarily made non-fatal, the **tree fingerprint changed** too. Also red: Scenario C, Scenario D, and the AST wiring test (which named the DB leg specifically) | +| 2 | The entire size term removed from `floorVerdict` (both its thresholds) | Scenario D **RED** — 0 alerts where 1 was required | +| 3a | The reserve removed entirely | Scenario F **PASSED — recorded honestly.** The specified mutation does not exercise the assertion: removing the reserve makes every app write, which overwrites and adds but **deletes nothing**, so a deletion-watching test correctly stays green | +| 3b | A prune injected into the refusal path | Scenario F **RED** — this is the mutation that proves the test watches deletion | +| 4 | Floor moved above the warning band (90% / 6 GiB) | `TestFloorSitsBelowTheCriticalWarningBand` **RED** | -Queried with -`curl -s "https://gitea.dooplex.hu/api/v1/repos/admin//actions/tasks?limit=3"` and matched on -`head_sha`, per `CLAUDE.md`'s pull-check rule — CI mails on failure, which is a push signal; this is -the pull check that catches a lost or unread mail. +Every mutation removed **every** guard its test covers (#1 removed both dump-leg gates, not one). -**`--no-verify` was NOT used**, anywhere. This clone is armed (`core.hooksPath = .githooks`). Both -gate entry points were also run by hand before committing: -`felhom.eu/scripts/repo_gates.py` → **all five OK** (site, hostinstall, hub-confirm, -manifest-bearer, reuse-refs), and `felhom-agent/scripts/agent_gates.py` → **OK** (reuse-refs). -No test suite was run because no code changed in either repo. +## 5. Live validation — demo-hp guest 9201 (Tier 0), the method that found the defect ---- +**Method:** endpoint-level — `POST /api/debug/backup/dbdump`, the exact endpoint the debug UI button +calls, which runs the production `RunDBDumps`. No browser on DooPlex. -## 10. Observations — noticed and NOT acted on +**The instrument was re-proven before use.** demo-hp's thin pool is 53.93 GiB, so a real fill of a +70 G volume would exhaust it and corrupt every guest. A 5 GiB `fallocate` step moved guest `df` +1.2G → 6.2G while thin-pool `data_percent` held **36.83 → 36.83** — zero blocks allocated. Re-checked +at every step of the fill. -- **The runbook's central claim was wrong in a way that mattered.** R-178 said *"a reinstall is now a - self-contained piece of work with no code left to write."* True about code; false about the artifact - channel — the agent half of the merge was unpublished, and following the documented path without - checking would have downgraded both boxes and produced a green, meaningless result. **No code change - was needed to make any step pass** — §10 asks this loudly, and the answer is no. What was needed was - a publish. -- **Prove-then-vouch was already spent when the session opened.** Not a defect in anything, but the - rule protected nothing because nothing enforced it. R-120's gate is the shape that would. -- **A per-stack `POST /stacks//backup` does not produce volume dumps** — the volume leg lives in the - full app-data pass. Not wrong, but the endpoint's name suggests otherwise and it cost time here. -- **privatebin's recovery unit carries no user data** (its catalog entry declares no `backup:` - section), so restoring it loses every paste. That may be intentional for an expiring, E2E-encrypted - paste bin — but nothing in the app's description tells the customer so. Catalog question, not filed. -- **V-c doubles systemd's mount-unit count** — every docker overlay appears twice, - `var-lib-docker-…-merged.mount` and `var-lib-felhom-docker-…-merged.mount`, because `/var/lib/docker` - is a bind of `/var/lib/felhom/docker`. Cosmetic, inherent to the chosen variant, no action. -- **`fstrim` cannot run inside the guest** (`EPERM`, unprivileged LXC), so space freed inside the guest - is not returned to the thin pool. It did not matter here; on a box that fills and empties repeatedly - it would. -- **The floor's own message mixes units** — it reports `64.2/68.7 GB used (93%)` against a threshold - stated as `97% used or 1.0 GiB free`, while `df` showed 96 %. GB-vs-GiB, so the percentage term - fires later than an operator reading `df` would expect. Minor; noted on R-181's fix shape rather - than filed separately. +### Headroom term — 08:59:46, 906 MB free / 99% used + +| Observable | Result | +|---|---| +| Tree fingerprint before | `TREE_SHA=111d1760c18d3440f700634ab325f8b8` (10 files; opengist's tar **182,272 B** — R-181's own "before" figure) | +| Tree fingerprint after | **`111d1760c18d3440f700634ab325f8b8` — identical** | +| Volume dumps written | **0** (baseline run at 08:58 wrote 2) | +| `Stopping for safe volume dump` | **absent** — and this is evidence, not an absence, because that line **is** present in the 08:58 baseline | +| Operator alerts | one `recovery_unit_capture_failed` per app, severity `error`, HTTP 200 | + +Free space restored → re-run at **09:01:33**: both apps captured normally. + +### Size term — 09:03:00, proven separately + +Reproducing the original sequence: a real 2 GiB file planted in opengist's volume, backed up so its +**previous** tar became **2,147,666,432 B** (the exact live figure), then the filesystem set to +**91% used / 2.9 GB free — both headroom terms deliberately clear**. + +- **opengist refused `(size)`** — *"this app's last backup was 2.0 GB and writing it again would cross the reserve"* +- **privatebin ADMITTED and dumped normally** — the term is per-app, not a global halt +- Tree unchanged; 1 volume dump instead of 2 + +### One honest correction to the "app not stopped" claim + +`StartedAt` on both apps *did* move, 26 s **after** the refusal. It was the **quiesce loop** for the +whole-guest PBS backup, which my fill had broken — not the app-data path. Its own backoff logic then +behaved correctly (*"deferring its next quiesce by 15m so the apps are not stopped again for a backup +that cannot succeed"*). The app-data claim rests on the **absence of the `Stopping … for safe volume +dump` line**, which is the line that appears when that leg bounces an app. + +## 6. The `du` measurement (§Part 1.3) — measured, then rejected + +**66 timed runs** on demo-hp guest 9201, `docker run --rm -v :/v alpine du -sb /v`: +**median ~355 ms per volume, range 341–404 ms** — on volumes holding **tens of KB**. The cost is +container start-up, not the walk, so it does not shrink for small apps and only grows for real ones. + +**Rejected**, on two grounds beyond the number: `docker run` needs the writable layer, so the +measurement mechanism can fail under exactly the disk pressure the reserve exists to handle; and the +previous-dump estimate measures the **artifact that will be written** rather than the live volume, +which is the truer predictor. The previous-dump estimate stands. + +## 7. The refusal message as shipped, and what it guarantees + +``` +[WARN] [backup] App backup REFUSED for opengist (headroom) — refused: backing up this app would +leave the filesystem below the reserve (reserve: 97% used or 1.0 GiB free; the filesystem is already +below it, before this app's estimated 178.0 KB write) — /mnt/sys_drive: 64.3/68.7 GB used (94%), +0.9 GB free; NO database dump, NO volume dump and NO recovery-unit capture was written for it, the +previous unit is untouched and NOTHING was deleted +``` + +**It guarantees, for that app in that run:** no DB dump, no volume dump and no capture were written; +every file under `backups/primary/` is byte-identical; the app was not stopped; nothing anywhere +was deleted; exactly one operator alert was sent. All five verified by fingerprint above. + +**The wording was not weakened to fit the behaviour** — the behaviour moved so the wording became +true. What was *added* is the bound term (`headroom` / `size`) and the estimate. + +**v0.193.1 — found by this very proof run.** The estimate was rendered fixed to two-decimal GiB, so +opengist's real **178 KB** printed as `estimated 0.00 GiB write`, which reads as *no estimate was +available* — the opposite of what happened. Shipped the same session because it is the same defect +class the whole task is about. Re-verified live after redeploy: `estimated 178.0 KB write`. + +## 8. papra (R-156, last leg) + +**Precondition checked, not inherited** — both boxes were wiped and rebuilt today, so the 2 August +evidence was re-measured: `docker ps -a` (**including stopped**) on **both** demo guests → no papra; +hub `/hosts` → exactly two enrolled hosts (`demo-felhom-8363b5`, `demo-hp-bb76ea`), **zero** papra. + +**Decided from the image, not the README:** `WORKDIR=/app`, `DATABASE_URL=file:./app-data/db/db.sqlite`, +`DOCUMENT_STORAGE_FILESYSTEM_ROOT=./app-data/documents`, `PAPRA_CONFIG_DIR=./app-data` — and +**`/app/data` does not exist in the image at all**. + +**Departure from the task's stated preference order, stated because it was deliberate.** Option (1) +(reconfigure the app to write to `/app/data`) *was* available — all three paths are env-settable. Not +taken: it enumerates data paths, so a fourth added upstream would silently escape to the writable +layer again — this defect re-armed and invisible. Mounting the app's own data **root** captures every +current and future path by construction. + +**Gate output — the arbiter, run in both directions:** + +- fixed → `papra CLEAN`, with the self-test passing on that run: *"prober flags the R-156 signature and clears a correct template — trustworthy"* +- reverted to `/app/data` (red-proof on the **real template**, not just the canary) → `BROKEN`: *"mount /app/data is NOT writable by the app's own uid=999"*, *"DATA in the writable layer at /app/app-data/db (db_signature=True, e.g. ['db.sqlite'])"*, *"declared volume /app/data is EMPTY"* +- `catalog_gates.py papra` (full, not `--fast`) → **rc=0**, all three gates OK + +**Two operational findings about the gate:** it needs **root** (it reads `/var/lib/docker/volumes`, +mode `drwx--x---`; as a normal user its own canary fails UNDETERMINED and it correctly refuses a +verdict — fail-closed working as designed), and it hardcodes scratch path `/srv/felhom-gate`, created +on DooPlex. Unscoped it deploys all 53 templates; that run was aborted after 10 minutes and its +`volgate-*` scratch projects were cleaned up. + +## 9. §3's correction — confirmed in passing, not chased + +`restore_points.go:57-59` takes the manifest's mtime and then `newestArtifact` over the `.sql` and +`.tar` files, so **the newest of the three wins**. The restore point does **not** show a stale +timestamp. Confirmed and dropped, as instructed. + +## 10. Register + +| ID | Change | +|---|---| +| **R-181** | **CLOSED — SHIPPED** (v0.193.0 + v0.193.1), with the live evidence above | +| **R-156** | **CLOSED** — all three apps fixed | +| **R-110** | WAITING-ON-OPERATOR → **READY**, ruling attached: **option (b), tag-tracked**, and it must cover **both** channels (the `/scripts/` git-sync *and* the nine files fetched from `raw/branch/main`) or it only half-works | +| **R-115** | WAITING-ON-OPERATOR → **READY**, ruling attached: **mechanism (b)**, a build-side gate refusing to deploy or vouch an unpublished version; the third instance (agent v0.120.0) would have silently downgraded both demo boxes while reporting success | +| **R-182** | **NEW.** ID established free: `grep -ro "R-182\b"` over `documentation/` and `*.md` → 2 hits, both prose in `REPORT.md` recording it as *"checked and left unused"*; `R-183` → 0 hits and remains free | + +**R-165** is collapsed to CLOSED/PROVEN-LIVE in `ROADMAP.md`; the capability map's local-backup row +moves to **PROVEN-LIVE, both halves**, because the live fill proved the fixed behaviour for **both** +reserve terms. + +## 11. Observations — noticed, documented, NOT acted on + +1. **R-182 (filed).** The periodic status refresh (`GetFullStatus` → `captureAllRecoveryUnits`) runs + with no admission scope, so a refused app re-alerts on every poll — measured live: a second + identical alert pair 13 s after the run's. **Pre-existing in v0.192.0**; R-181 changed neither + caller. Its mitigation is a *comment* claiming the hub owns cooldown — which is exactly the + "invariant asserted in a comment with no test pinning it" shape, so verify at the hub before + scoping. +2. **A reserve refusal does not make the run fail.** The DB and volume legs record `SKIP`, not `FAIL`, + so `lastDBDump.Success` stays true and the customer-facing status does not turn red. Deliberate and + consistent with v0.192.0 (the capture refusal never set it either), and the operator alert is the + signal — but it means "backup succeeded" and "every app was backed up" are not the same statement. +3. **`UnitSpace.UsedPercent` and `df` disagree** — `df` reported 99% where the alert said 94%, because + `df`'s figure accounts for ext4 reserved blocks and the floor's does not. Harmless here (the + free-byte term bound), but a percent-term threshold is being compared against a number the operator + cannot reproduce with `df`. +4. **The whole-guest PBS backup fails when the volume is near-full**, pushing + `whole_guest_backup_failed` (severity `error`). Expected under a deliberate fill, and its backoff + behaved correctly; noted because it is collateral any future fill test will also produce. + +## 12. Teardown + +Fill file removed; the planted 2 GiB file removed; a final backup regenerated a correct 178 KB tar; +`pct fstrim 9201` returned 67.5 GiB and the thin pool settled at **29.43%**, *below* its 36.83% +baseline. The backups tree is byte-identical to the pre-test fingerprint. Guest helper scripts and the +credential file `shred`-ed. `volgate-*` scratch compose projects removed; the unrelated 9-day-old +`jarr-*` containers on DooPlex were left untouched. papra is **not** left deployed. + +No `--no-verify` was used on any push; the `felhom-controller` pre-push hook ran and reported +`gates OK` on both pushes. diff --git a/STATUS.md b/STATUS.md index bb515504..87883e1d 100644 --- a/STATUS.md +++ b/STATUS.md @@ -27,107 +27,74 @@ time, and an app switched off deliberately stayed off every time. also delete it. A daily snapshot is armed as a stopgap, and we have never restored from that copy. *(R-95, R-87)* -**Three apps out of fifty-three kept their data where backups never looked.** They reported healthy; -the data would vanish on the next update. Two are fixed, the third is now clear to fix because it is -installed nowhere. *(R-156)* +**A full disk emails you repeatedly instead of once.** When the reserve refuses an app's backup you +are told once by the backup run — correctly — but the page showing backup status re-checks on a timer +and sends the same message again each time. Not new: as old as the reserve itself, and seen only +because we watched the alerts closely while proving the fix below. Harmless if the hub already +collapses repeats — which a comment claims and nobody has checked. *(R-182)* -**The reserve does not guard the step that actually fills the disk.** The rule that refuses a backup -when space runs low is checked at the wrong moment: the big write happens first, unchecked, and only -the small write after it is refused. So the thing meant to stop a runaway backup is the thing it -runs past. Worse, when it does refuse it says *"your last good copy is untouched"* — and we measured -that copy being overwritten by the earlier step anyway. Found by deliberately filling a rebuilt demo -machine. Nothing was deleted and you were emailed, both correctly. *(R-181)* +## What shipped recently -**When that happens, only one page says so** — no email, no alert. The page that answers "is this app -backed up?" is the one that stays silent. *(R-158)* +**The backup partition is gone, and both demo machines run on the new shape.** Wiped and rebuilt on +3 August and taken through the whole customer journey — set up, install an app, back it up, restore +it. One storage area instead of two; the space a backup can use went from 19 GB to 65 GB on the small +machine and 45 GB to 233 GB on the big one. Three reboots each, correct every time. The two were +rebuilt deliberately differently — one from a local copy of the image, one by the ordinary customer +route with the published fingerprint checked — so the disk shape and the delivery route are both +proven, rather than one proven twice. Their previous demo apps and data are gone; that was the point +of a wipe, and you approved it. *(R-165, R-178)* -**The checks now have two nets, and the second one emails you.** Every repository has one command -that runs all of its checks; it runs by itself before every push and refuses a push that fails. That -one lives on the workstation and can be skipped. So the build server now runs the same checks again, -on a machine that does not care who pushed or what they typed — and **when they fail it sends you an -email**, because a red mark on a page nobody watches is not a warning. Proven with a real broken -change, not assumed. The one thing it still cannot do is *stop* the change: every change here goes -straight to the main copy with no review step, so there is no point in the road for it to stand at. -It notices, quickly, and tells you. *(R-29, R-161, R-168, R-169)* +**What replaced the wall — and it now watches the right moment.** The wall was quietly doing a second +job: keeping a runaway backup from eating the space the machine needs to keep running. That job is now +explicit, and as first built it was checked too late — the big write happened first, unchecked, and +only the small write after it was refused, while the message still promised your last good copy was +untouched. **Fixed and proven on 3 August.** The machine now decides once, per app, **before it writes +anything at all**, and that one answer covers all three steps: a refused app writes nothing, is not +restarted, and the promise is now literally true — checked by fingerprinting every file before and +after. It also stopped being blind to size, so an app is no longer waved through at 96% full and then +allowed to write two gigabytes. Proven by deliberately filling a demo machine, once for each way it +can refuse. Nothing is ever deleted to make room: every app has only one local copy, so "delete the +oldest" would always mean destroying some other app's only copy. *(R-181)* -**A filling disk now warns the customer before anything breaks, and a failed backup now reaches you.** -Until today the first sign that a disk was filling up was a backup that did not happen — nothing said -anything beforehand. Two things changed. The customer is now warned while there is still room to act, -naming the drive and how much space is left, in plain Hungarian that says what to do about it. And -when one app's backup fails for any reason, **you** are told which app and why, with the disk figures -attached — the page that answers "is this app backed up?" was, until now, the one page that never -said. The customer is deliberately *not* told about that second one: they can free up space, but they -can do nothing about a backup that failed, so telling them would only alarm them. +**The last of the three apps that never saved their data is fixed.** Installed nowhere, so nothing was +stranded — checked on both demo machines and in the fleet list rather than assumed. Proven by the check +that caught it, run in both directions: it clears the fixed version and still convicts the old one. +*(R-156)* -Both were proven on the demo machine by actually filling a disk. One detail is worth knowing because -it is why there are two rules and not one: the serious warning fired when free space dropped below a -fixed amount while the disk was only 91% full — a percentage on its own would have missed it. +**A filling disk warns the customer before anything breaks, and a failed backup reaches you** — the +customer while there is still room to act, naming the drive and the space left; you when one app's +backup fails, with the disk figures. The customer is deliberately not told about the second: they can +free space, but they can do nothing about a failed backup. Both proven by filling a real disk. There +are two rules and not one because the serious warning fired on free space while the disk was only 91% +full — a percentage alone would have missed it. *(R-167, R-158)* -**These went in *before* the partition change deliberately.** The partition being removed is also a -barrier against a runaway backup filling the space the machine needs to run; putting the warnings in -first means that when it comes down, the thing watching is already working and already tested. -*(R-167, R-158)* - -**The backup partition is gone — and both demo machines now run on the new shape.** Each was wiped -and rebuilt from the new base image on 3 August and taken through the whole journey a customer takes: -set the machine up, install an app, back it up, restore it. One storage area instead of two, and the -space a backup can use went from 19 GB to 65 GB on the small machine and from 45 GB to 233 GB on the -big one. The wall does not move; it stops existing. Each machine was rebooted three times over and -came back correctly every time. **The two machines were rebuilt deliberately differently** — the -first from a copy of the image held locally, the second by the ordinary route a real customer takes, -fetching the image and checking it against the fingerprint we publish, so both the disk shape and the -delivery route are now proven rather than one twice. - -**Their previous demo apps and data are gone.** That was the point of a wipe and you approved it; the -machines now carry a couple of small test apps instead. - -**What replaced the wall.** It was quietly doing a second job — keeping a runaway backup from eating -the space the machine needs to keep running. That job is now explicit: if a backup would push the disk -below a safe reserve, **that one app's backup is refused, its last good copy is left exactly as it -was, and you are told.** Nothing is ever deleted to make room; every app has only one local copy, so -"delete the oldest" would always mean destroying some other app's only copy. - -**How the shape was chosen — worth one line, because it was not the obvious one.** Three ways of doing -it were built and rebooted rather than argued about. All three worked. They differed in what they -quietly broke: one put your backups inside Docker's own storage, where the normal way of fixing a sick -Docker would wipe them; another exposed all of Docker's internals to the part of the system that -manages your drives. The third does neither, and costs one extra line of configuration. - -**The new base image is now switched on**, so any machine installed from here on gets the new shape. -The tester's box is untouched and keeps working exactly as before; it gets the new shape whenever it -is reinstalled. *(R-165, R-178)* +**The checks have two nets and the second emails you.** Every repository has one command that runs all +its checks, before every push. That one can be skipped, so the build server runs them again and emails +you on failure. It cannot *stop* a change — everything goes straight to the main copy with no review +step — but it notices quickly and tells you. *(R-29, R-161, R-168, R-169)* ## What we're working on -- **Now:** the last app whose data was never saved; today's decisions written down. -- **Next:** fixing the reserve so it guards the step that fills the disk, and stopping it from - claiming your last good copy is untouched when it is not *(R-181)*. The partition merge itself is - done and proven on both machines. -- **After:** rebuilding how the machine records whether an app is meant to be running. +- **Now:** both of today's items are done — the reserve and the last unsaved app. Your two decisions + are written down and are ours to build. +- **Next:** building those two — moving the installer onto a labelled version so publishing is one + step you can undo, and a check that refuses to install a version nobody can download *(R-110, R-115)*. +- **After:** the off-site copy that the machine making it can still erase *(R-95, R-87)*. ## Waiting on you -- **How a new version reaches a machine — and it has now been forgotten a third time.** Pushing the - installer publishes it: half a minute later every new machine downloads it, with no staging and no - way back but another push. And publishing is a step we remember rather than one the release - performs. On 3 August the new agent — the half of the partition merge that runs on the machine — - turned out to have been built and installed on both demo machines but **never published**, so a - rebuild would have quietly put the *old* one back and proved a version nobody ships. Caught before - the wipe and fixed in ten minutes, but only because someone happened to look. Nothing is installing - today, so this is the cheapest moment to settle it. *(R-110, R-115)* - **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a session log; nothing suggests anyone else saw it. *(R-132)* -- **Nothing on the partition merge itself** — the shape is chosen, built, and now proven on both demo - machines. The tester's box needs no conversion: it will simply be reinstalled. What is left is the - reserve defect above, which is ours to fix, not yours to decide. *(R-165, R-181)* +- **Nothing else.** You settled both open questions on 3 August — the installer moves onto a labelled + version, and a check will refuse to install a version nobody can download. Both are written down and + are ours to build. *(R-110, R-115)* ## Changed since last update -- **2026-08-03** — Both demo machines wiped and rebuilt from the new base image, and taken through - set-up → install an app → back it up → restore it. The backup space ceiling is gone and measured. - The hard stop that replaced the old partition was fired for the first time on real hardware — it - refused, it deleted nothing, it emailed you, **and it turned out to be watching the wrong step**; - that is now the top thing to fix. +- **2026-08-03** — The reserve now guards the step that fills the disk, and its promise is true; the + last app whose data was never saved is fixed. Both proven on a demo machine, not just in tests. + Earlier the same day: both demo machines wiped and rebuilt from the new base image and taken through + set-up → install an app → back it up → restore it, with the backup space ceiling gone and measured. - **2026-08-02** — The false "host offline" warning is fixed. The hub's database was supposed to be in a mode where reading a page cannot block a machine's status update; a one-word difference meant that @@ -140,11 +107,5 @@ is reinstalled. *(R-165, R-178)* and fixed the same day. - **2026-08-02** — Thirteen mechanical checks had built up and nothing ran most of them; two were - failing quietly. Fixed, and the arrangement that replaced them is described above. -- **2026-08-02** — Decided: the 20 GB backup partition goes away and shares space with app data. That - changes the disk layout, so it happens before any machine is installed outside the house. **Measured - since:** no machine outside the house is registered yet, so this is as cheap now as it will ever be; - and the two demo machines can simply be reinstalled rather than converted. -- **2026-08-02** — Decided: only this machine and the tester's box are protected; every other box, - demo boxes included, may be broken or reinstalled freely. Two of the three apps that never saved - their data are fixed; this page created. + failing quietly. Fixed. Decided the same day: the 20 GB backup partition goes away; and only this + machine and the tester's box are protected, every other box may be broken or reinstalled freely. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 00fec6ee..89fc4a9f 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -86,7 +86,7 @@ | Box survives a **site/network change** (relocation, different subnet, DHCP re-lease) with the control plane intact | agent v0.96.0 (island NIC), host-install v1.19.0, controller (unchanged), bootstrap | **PROVEN-LIVE (2026-07-25)** | **R-50 SHIPPED and deployed to the whole fleet.** The control plane now rides a host-internal, portless island bridge (`vmbr9`, `169.254.253.1/30`↔`.2/30`) with a fixed private address that no LAN/DHCP/site move can invalidate. Proven end-to-end: the spike's F1 replay (renumber the LAN → agent stays bound on the island, control plane HTTP 200; the LAN-literal contrast reproduces the original `bind: cannot assign requested address` daemon-death) + cold-reboot survival (`SPIKE-island-bridge-2026-07-25.md`), the migration runbook run verbatim (`RUNBOOK-island-migration.md`), a fresh provision auto-attaching the island `net1` (A4), and the live migration of **both demo boxes** (demo-hp + demo-felhom, 2026-07-25) — island `/storage` HTTP 200, LAN DNS pinned to the LAN IP (Finding-1), **apps served throughout (0 container restarts)**, hub reporting 0.96.0. **Origin:** `audits/AUDIT-vacation-remote-ops-2026-07-20.md` — the real relocation where the agent's LAN-literal bind took storage/PBS/quiesce/restore-test/DR down silently; that is now structurally impossible on a migrated box | **Fleet: DONE.** Remaining: **R-74** — bring the island to Peti's 2-node cluster (SDN vnet / bridge parity), its own supervised runbook. Related historical: R-51 (dead-primary alerting), R-52 (boot desired-state reconciliation), both shipped | | **The customer is warned BEFORE a filesystem fills** — per filesystem, in Hungarian, naming the drive and the free space, edge-triggered | controller **v0.191.0/.1/.2**, hub **v0.89.0** (R-167, decision D-c) | **PROVEN-LIVE (2026-08-02)** | `audits/SPIKE-r165-mp1-merge-2026-08-02.md` (context) + `felhom-controller/REPORT.md`. Exercised on guest 9201 against a REAL filesystem (`/mnt/sys_drive` filled with `fallocate`): **`disk_warning` at 90% used / 4.7 GB free** → hub `notification_log` `customer | disk_warning | sent` with the dynamic Hungarian rendered; grown to 1.7 GB free → **`disk_critical`** → `customer | sent`; file removed → `critical → ok … cleared silently, re-armed` and the persisted state emptied. **Exactly two events across three boots** — the boot in between produced none, which is the edge trigger holding | **Nothing warned before this.** The only prior signal was the healthcheck's generic `health_degraded` at 90%, for REGISTERED STORAGE PATHS ONLY — it never looked at the docker area or the system-data area, never gave a free-byte figure and never named a drive. **The two event types already existed with NO PRODUCER** (`disk_warning`/`disk_critical`: allowlisted, copy'd, in `DefaultEnabledEvents`, checkbox'd) — the **sixth** *built-but-never-wired* instance here; this ships their producer rather than a seventh near-duplicate type. **Two threshold terms, whichever trips first, and the live proof vindicated the design:** the critical crossing fired on the FREE-BYTE term (1.7 GB) at only **91%** used — a percentage-only rule would have missed it. The hub's generic `customerMessages` entries were REMOVED, because `FormatCustomerEmail` prefers the entry over the message and would discard the label and figures. **Known gap → R-177:** there is no operator-triggerable run-now path; the check is daily 03:30 + once at startup, so confirming a cleared warning on a support call needs a controller restart or a wait | | **A failed per-app Tier-1 backup reaches the OPERATOR** (app, error, and the target filesystem's used/free bytes at the moment of failure) | controller **v0.191.0**, hub **v0.89.0** (R-158, closed by R-167) | **PROVEN-LIVE (2026-08-02)** | `felhom-controller/REPORT.md`. Two real capture failures on guest 9201 (`mkdir …/backups: permission denied`) → both accepted and stored by the hub, `operator | recovery_unit_capture_failed | sent`, and the positive observable **`customer | recovery_unit_capture_failed | skipped | operator_only`** read from the hub's `notification_log`. One event per app, loop continuing | **Before this the failure was a `[WARN]` line and nothing else** — the manager carried three notify seams and none for the unit capture, so `/backups/apps`, the page you open to ask whether ONE app is backed up, was the one page that never said. **Deliberately NOT `backup_failed`:** that type is customer-enabled by default and carries Hungarian copy, so reusing it — which R-158's own proposal said — would email the customer about a failure they cannot act on. **D-c routes it to the operator and overrides the proposal.** Operator-only is enforced by `notify.operatorOnlyEvents`, NOT by the absence of a `customerMessages` entry (the v0.78.0 defect); a red-proof removing the register entry shows the customer receiving it | -| **A local backup is bounded by the box's FREE SPACE, not by a partition set at build time** — the appliance ships ONE data volume, and a capture that would exhaust it is refused per app rather than allowed to stop the container runtime | golden `build-golden.sh` **v3.0.0**, agent **v0.120.0**, controller **v0.192.0** (R-165 / D-a / B2) | **IMPLEMENTED — the LAYOUT half is PROVEN-LIVE (2026-08-03); the REFUSAL half is not (R-181)** | `REPORT.md` (R-178 reinstalls) + `audits/SPIKE-r165-phase0-2026-08-03.md` (P1/P2/P3) + the bake transcript. **The golden bake is real evidence and is cited as such:** `build-golden.sh v3.0.0` produced `including mount point mp0 ('/var/lib/felhom')` with **no `mp1` line at all**, and its own guards printed `/var/lib/docker is a real mount`, `/mnt/sys_drive is a real mount` and `both paths are ONE filesystem`. Archive published (registry HTTP 200, sha `54e2a4c4…`). The B2 floor is unit-proven with 3 red-proofs and live on 9201 | **The row's FIRST clause is now PROVEN-LIVE; its SECOND is not, and they are separated deliberately.** **Proven (R-178, 2026-08-03):** *"a local backup is bounded by the box's FREE SPACE, not by a partition set at build time"* — both demo boxes reinstalled from this golden, by two different supply paths (demo-hp `--golden `; demo-felhom the normal manifest route with **`verified sha256 54e2a4c431daf580… matches the hub manifest`**), each showing `mp0` at `/var/lib/felhom` with **no `mp1`**, both consumer paths real mounts on ONE filesystem (`stat -c %d` = `64519` on all three), 3/3 reboots each, and claim → deploy → backup → **restore** with a planted marker returning byte-identical. Space available to a recovery unit measured at **65 GiB / 233 GiB**, against the **19 GiB / 45 GiB** those boxes' `mp1` slices offered. **NOT proven — and measured FALSE in part:** *"a capture that would exhaust it is refused per app rather than allowed to stop the container runtime"*. The floor fired live for the first time (demo-hp 06:40:03) and does refuse per app, delete nothing, and alert — **but it is checked only in `captureAllRecoveryUnits`, while `runVolumeDumps` writes the bulk with no floor check at all**, so the leg that exhausts the volume is the unguarded one; and the refusal's claim that the previous unit is untouched was measured false (a 182,272 B dump replaced by 2,147,666,432 B under a manifest still dated 06:34:26). → **R-181**. The golden **is now VOUCHED** (2026-08-03, hub `Artifact manifest set: … golden=0.192.0`), so fresh installs pick up the merged layout. Every box in the field that has not been reinstalled is still on the SPLIT layout and is unaffected: nothing assumes the merged shape at runtime, the controller's system_data_path is a path rather than a volume, and agent v0.120.0 FOLDS the retired `-sysdata-grow` into the single grow so an older `felhom-host-install.sh` still provisions the same total capacity | +| **A local backup is bounded by the box's FREE SPACE, not by a partition set at build time** — the appliance ships ONE data volume, and a capture that would exhaust it is refused per app rather than allowed to stop the container runtime | golden `build-golden.sh` **v3.0.0**, agent **v0.120.0**, controller **v0.193.1** (R-165 / D-a / B2, completed by R-181) | **PROVEN-LIVE (2026-08-03) — BOTH halves** | `REPORT.md` (R-178 reinstalls) + `audits/SPIKE-r165-phase0-2026-08-03.md` (P1/P2/P3) + the bake transcript. **The golden bake is real evidence and is cited as such:** `build-golden.sh v3.0.0` produced `including mount point mp0 ('/var/lib/felhom')` with **no `mp1` line at all**, and its own guards printed `/var/lib/docker is a real mount`, `/mnt/sys_drive is a real mount` and `both paths are ONE filesystem`. Archive published (registry HTTP 200, sha `54e2a4c4…`). The B2 floor is unit-proven with 3 red-proofs and live on 9201 | **The row's FIRST clause is now PROVEN-LIVE; its SECOND is not, and they are separated deliberately.** **Proven (R-178, 2026-08-03):** *"a local backup is bounded by the box's FREE SPACE, not by a partition set at build time"* — both demo boxes reinstalled from this golden, by two different supply paths (demo-hp `--golden `; demo-felhom the normal manifest route with **`verified sha256 54e2a4c431daf580… matches the hub manifest`**), each showing `mp0` at `/var/lib/felhom` with **no `mp1`**, both consumer paths real mounts on ONE filesystem (`stat -c %d` = `64519` on all three), 3/3 reboots each, and claim → deploy → backup → **restore** with a planted marker returning byte-identical. Space available to a recovery unit measured at **65 GiB / 233 GiB**, against the **19 GiB / 45 GiB** those boxes' `mp1` slices offered. **NOT proven — and measured FALSE in part:** *"a capture that would exhaust it is refused per app rather than allowed to stop the container runtime"*. The floor fired live for the first time (demo-hp 06:40:03) and does refuse per app, delete nothing, and alert — **but it is checked only in `captureAllRecoveryUnits`, while `runVolumeDumps` writes the bulk with no floor check at all**, so the leg that exhausts the volume is the unguarded one; and the refusal's claim that the previous unit is untouched was measured false (a 182,272 B dump replaced by 2,147,666,432 B under a manifest still dated 06:34:26). → **R-181, CLOSED THE SAME DAY (controller v0.193.0 + v0.193.1) and the second half is now PROVEN-LIVE TOO.** The reserve became a **per-app, per-run ADMISSION decision** taken before the app's FIRST write and covering all three legs (DB dump, volume dump, capture) — they write under one per-app root, which is what lets one verdict cover them honestly — and it gained a **size term**, so an app is no longer admitted at 96% and then allowed to write 2 GB. **Re-proven by filling demo-hp deliberately, once for EACH term, using the method that found the defect.** *Headroom @ 08:59:46* (906 MB free / 99%): both apps refused, **the whole `backups/primary` tree byte-identical — `TREE_SHA` 111d1760c18d3440f700634ab325f8b8 before and after**, opengist's tar still at its original 182,272 B; **no `Stopping for safe volume dump` line at all**, which is the positive-by-absence observable that matters because that line IS present in the 08:58 baseline run; 0 volume dumps; one alert per app, HTTP 200. Space freed, re-run @ 09:01:33 → both captured normally. *Size @ 09:03:00*, reproducing the original sequence with a real 2 GiB file in opengist's volume (previous tar **2,147,666,432 B**, the exact figure the defect was measured at) and the filesystem at **91% used / 2.9 GB free — both headroom terms deliberately clear**: opengist refused `(size)` while **privatebin was ADMITTED and dumped normally**, proving the term is per-app rather than a global halt. **The refusal's wording was NOT weakened to fit** — the behaviour moved so the wording became true, and it is verified by tree fingerprint rather than by reading the log line, which is what lied. The `fallocate` instrument was re-proven on the rebuilt box before use (5 GiB step moved guest `df` while thin-pool `data_percent` held **36.83 → 36.83**), and teardown returned the pool to **29.43%**, below its own baseline. The golden **is now VOUCHED** (2026-08-03, hub `Artifact manifest set: … golden=0.192.0`), so fresh installs pick up the merged layout. Every box in the field that has not been reinstalled is still on the SPLIT layout and is unaffected: nothing assumes the merged shape at runtime, the controller's system_data_path is a path rather than a volume, and agent v0.120.0 FOLDS the retired `-sysdata-grow` into the single grow so an older `felhom-host-install.sh` still provisions the same total capacity | | Soft-quota: usage bar, pre-push enlargement block, customer notification | controller v0.109/134, hub v0.41/55 | **PROVEN-LIVE** | 6D/6E; hub OffsiteChecker | | | **A customer (not the operator) performs a restore via UI alone** | all | **MISSING** (as evidence) | — | Alpha will produce this; script it into R-3. **2026-07-19:** the C6 evidence attempt ran and found a **product gap instead of evidence** — `audits/DIAG-immich-restore-2026-07-19.md`. A customer-driven UI restore of a DB-indexed app cannot currently succeed (R-43 file-only restore, R-44 stale dump), so this row cannot flip until those close. Row stays MISSING **by finding, not by absence of attempt** — the rehearsal system working, not failing. **2026-07-19: the blocking product gaps are CLOSED in controller v0.148.0** (R-43 + R-44 shipped), so this row is now blocked only on the evidence run itself, not on missing capability. It flips the moment the §9 acceptance produces screenshots + the outcome flash + a snapshot ID. **2026-07-19 round 2 — PARTIAL EVIDENCE ONLY, row NOT flipped** (`audits/DIAG-immich-restore-round2-2026-07-19.md`): a deliberate run from snapshot `49e7cb46` did recover all 11 assets (`status=active`, files resolve), but the operation **reported failure** and left immich reporting schema drift, because the replay aborted against the running app (H4). Photos back ≠ clean acceptance. **2026-07-20: H4 closed in controller v0.153.0 (R-47) on BOTH paths, AND THE EVIDENCE RUN HAPPENED.** *(The "closing in v0.149" wording above was wrong — v0.149.0 was the F3 dashboard fix; R-47 shipped in v0.153.0.)* The C6 drill ran end-to-end **through the UI**: photos deleted, **trash emptied**, the full files+database restore pressed on `/backups/restore`, 40 files placed + 1 DB dump replayed rc-0, 11 assets back, no drift, timeline visually confirmed. The method note below is now DEMONSTRATED, not merely written down. Evidence: `felhom-controller/REPORT.md` 4e. **Residual: the run was performed by the OPERATOR, not by a customer** — for this row literal wording the alpha still owes one genuinely customer-driven pass, but no product gap blocks it. Method note for R-3's script: deleting in an app's own UI usually means *trash*, not deletion, so a drill written that way merges 0 files, flashes success and proves nothing — a real drill must empty the trash **and** verify the app's *content*, not the file count **Lane split → `07-backup-architecture.md` §3**: this row is Lane 1 (customer, unassisted). §8 rows 1–5 are the routes it would exercise | diff --git a/documentation/architecture/07-backup-architecture.md b/documentation/architecture/07-backup-architecture.md index 1d56e22e..dc78a02f 100644 --- a/documentation/architecture/07-backup-architecture.md +++ b/documentation/architecture/07-backup-architecture.md @@ -566,17 +566,40 @@ mismatch table above (`mp0` 50 G vs `mp1` 20 G) describes what a merged box no l before provision grew them). **Read it as a function of `mp1`, and only for a box still on the split layout.** Measured: `audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1. -**What replaced the partition's second job.** `mp1` was also a BULKHEAD: an overflow was refused per -app with the last good unit byte-identical, and it **could not reach `/var/lib/docker`**, because that -was a different filesystem. On a merged box it can. Decision **B2**, shipped in controller -**v0.192.0**, is that bulkhead made deliberate — a two-term capture floor (97% used or 1 GiB free) in +**What replaced the partition's second job — the reserve.** `mp1` was also a BULKHEAD: an overflow was +refused per app with the last good unit byte-identical, and it **could not reach `/var/lib/docker`**, +because that was a different filesystem. On a merged box it can. Decision **B2**, shipped in controller +**v0.192.0**, is that bulkhead made deliberate — a two-term reserve (97% used or 1 GiB free) in `internal/fillwatch`'s shape, sitting beyond its critical band so the customer is always warned first. It **refuses per app and never deletes**: nothing on this filesystem is generational, so pruning could only destroy a different app's only local copy. -**Status caveat, deliberately explicit:** as of 2026-08-03 **no box has been reinstalled from the -merged golden** (R-178), so every box in the field is still on the split layout and everything above -still describes them exactly. This subsection describes what a box built from golden ≥ 0.192.0 gets. +**THE CONTRACT, stated as what the code provides (controller v0.193.0, R-181).** The reserve is a +**per-app, per-run ADMISSION decision, not a capture check.** It is taken once for an app, immediately +before that app's FIRST write of the run, and it covers **all three write legs — the database dump, the +volume dump and the recovery-unit capture**. Those three write under one per-app root +(`backups/primary/`), which is what makes one verdict able to cover them honestly. + +- **What it guarantees.** A refused app has **nothing written for it in that run**, its previous unit + is **byte-identical**, it is **not stopped**, nothing anywhere is deleted, and the operator gets + **exactly one** alert naming the app, the term that bound and the disk figures. +- **Two terms, two questions.** *Headroom*: is the filesystem already below the reserve? *Size*: would + THIS app's write take it below? The size estimate is the app's previous `.sql` + `.tar` on disk; + with no history the decision degrades to headroom alone, deliberately — otherwise the first backup + is the one that can never happen. +- **Why it is decided lazily and not once per run.** Space changes during a run: app A's dump can put + app B under the reserve, so a verdict taken at run start reads a disk that no longer exists. +- **Why it is never re-decided between an app's own legs.** That is precisely the shape v0.192.0 had — + the two dump legs unguarded and only the capture refused — under which the reserve was consumed by + the very write it exists to bound, and the refusal's *"the previous unit is untouched"* was measured + false. Proven live on demo-hp 2026-08-03 (R-181), fixed the same day, and re-proven by filling the + box for each of the two terms. +- **It sits ahead of `DumpAppVolumesSafe`**, which stops the stack as its first act — a refusal + decided inside it would already have bounced the app it is refusing to back up. + +**Status caveat, deliberately explicit:** every box in the field that has not been reinstalled is still +on the split layout and everything above still describes them exactly. This subsection describes what a +box built from golden ≥ 0.192.0 gets. Both demo boxes were reinstalled from it on 2026-08-03 (R-178). Two things are deliberately **not** recorded here. **The sizing ratio is the operator's ruling** (**R-163**) — this section states the constraint, not a number. And **the same-device placement is diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 5767345f..e045ace8 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -13,9 +13,9 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha | **R-88b** | ~~`/backup/due` cannot say *unknown*~~ | **SHIPPED + PROVEN-LIVE** (agent v0.105.0 + controller v0.178.0, 2026-07-27) | — | `age_state=unknown` captured on real hardware during a deliberate ep0 outage; controller deferred, **zero app stacks stopped** | — | | **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC | | **R-94** | ~~A hand-synced version constant drifts, and the gate that would catch it is never run~~ | **CLOSED — SHIPPED** (hub v0.87.0, 2026-08-02) | — | **All three legs closed.** **(a) closed by DELETION, not derivation** — deriving is not achievable honestly: the Setup command fetches `felhom-host-install.sh` at RUN TIME from a website that git-syncs `main` every 30 s (R-110), so no build-time value in the hub can be true, and a number that is wrong carries a version number's authority while being a guess. The const, the `pageData.ScriptVersion` field, its assignment and the rendered label are gone; a NOTE stands where the const was so it is not helpfully re-added. **(b)** `hostinstall_gates.py` gate 1 INVERTED — it now asserts the hub carries **no** host-install version literal, in six code shapes across every `.go`/`.html` under `hub/`; and the gate is now invoked, by `scripts/repo_gates.py` and the pre-push hook (→ R-29). **(c)** the tautological `render_test.go:219` assertion is deleted, not replaced — there is no version to assert. It was demonstrated PASSING with the const at `9.9.9` while the script was 1.22.0. The label had been wrong for 19 days (since 2026-07-14) | — | -| **R-110** | **`main` is the installer's publish channel — there is no staging.** `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a 30 s period and nginx serves that working tree directly (`location /scripts/`, `root …/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`) receives. There is no tag, no pinned-version path, no staging copy and no rollback other than another push — for the artifact that runs as **root on a virgin box**, the single most privileged thing Felhom ships | **WAITING-ON-OPERATOR (S)** | operator ruling | **Two consequences worth stating:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29 — and the precaution recorded on the old R-94 row as "do not point every new box at an installer that has never run" **was never available to take**. **Open question for the operator, not a defect to fix blind:** whether `/scripts/` should serve a pinned release (tag-tracked path, or a versioned directory with the customer command naming a version) or whether `main`-tracking is the accepted shape for a one-operator product. Exposure today is zero — there are no boxes installing — which is exactly why it is cheap to decide now. **SECOND INSTANCE, found 2026-07-29 by the E-2d run and filed here rather than as a new ID:** `felhom-host-install.sh` fetches **nine** files from `raw/branch/main` (`:2072`–`:2206`) and the hub manifest vouches a sha for exactly **one** (`wrapper_sha256` → `felhom-pbs-apply`; re-checked this run, no drift). E-2a's `felhom-backup-target-apply` (`:2116`) is installed **0755 to `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n` — a root-executed artifact taken from `main` with no pinned integrity, which is this row's class exactly | CC | +| **R-110** | **`main` is the installer's publish channel — there is no staging.** `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a 30 s period and nginx serves that working tree directly (`location /scripts/`, `root …/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`) receives. There is no tag, no pinned-version path, no staging copy and no rollback other than another push — for the artifact that runs as **root on a virgin box**, the single most privileged thing Felhom ships | **READY (S)** — operator ruling taken 2026-08-03: **option (b), tag-tracked** | — | **Two consequences worth stating:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29 — and the precaution recorded on the old R-94 row as "do not point every new box at an installer that has never run" **was never available to take**. **Open question for the operator, not a defect to fix blind:** whether `/scripts/` should serve a pinned release (tag-tracked path, or a versioned directory with the customer command naming a version) or whether `main`-tracking is the accepted shape for a one-operator product. Exposure today is zero — there are no boxes installing — which is exactly why it is cheap to decide now. **SECOND INSTANCE, found 2026-07-29 by the E-2d run and filed here rather than as a new ID:** `felhom-host-install.sh` fetches **nine** files from `raw/branch/main` (`:2072`–`:2206`) and the hub manifest vouches a sha for exactly **one** (`wrapper_sha256` → `felhom-pbs-apply`; re-checked this run, no drift). E-2a's `felhom-backup-target-apply` (`:2116`) is installed **0755 to `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n` — a root-executed artifact taken from `main` with no pinned integrity, which is this row's class exactly. **OPERATOR RULING 2026-08-03 — option (b) chosen: the publish channel moves from `main`-tracking to a TAG.** Publishing becomes *moving the tag*, and rollback becomes *moving it back* — the property `main`-tracking cannot have at any price. Recorded here, **not built this session, by instruction**. **The ruling carries a condition that decides whether the fix works at all: it must cover BOTH channels.** (i) the nginx-served `/scripts/` git-sync (`manifests/webpage.yaml`, `--branch=main`, 30 s period) that every `felhom-bootstrap.sh` fetch and every operator day-0 command reads, AND (ii) **the nine files `felhom-host-install.sh` fetches from `raw/branch/main`** (`:2072`–`:2206`), of which the hub vouches a sha for exactly ONE. Fixing only (i) leaves a tagged installer pulling nine untagged files from `main` at run time — a staging story that is false in the place it matters most, since one of those nine (`felhom-backup-target-apply`) is installed **0755 into `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n`. **Exposure is still zero** (no boxes installing), which is exactly why it stays cheap. Now CC's to build | CC | | **R-111** | ~~**The Day-0 artifact channel is 17 agent releases stale — a box installed today gets agent `0.96.0`, not `0.113.0`.**~~ `felhom-host-install.sh` does not use `main`: it reads the hub-vouched manifest (`:423-436`) and fetches Gitea generic packages (agent `:1945`, golden `:2573`). Gitea's newest are **agent 0.96.0** and **golden 0.161.0**, and the hub's manifest selects exactly those — so a fresh box lands on **agent 0.96.0 + controller 0.161.0** (global floor `v0.156.0` < the golden's 0.161.0, so no self-update) against `main`'s 0.113.0 / 0.185.1. Agent 0.113.0 reached both demo boxes by **direct deploy and was never published** | **SHIPPED 2026-07-29 — the channel now serves agent 0.113.0 + golden 0.185.1** | — | **FIXED the same day it was found.** Agent **0.113.0** built from the clean tree @ `58b598b` and published (`scripts/publish-agent.sh`), sha `5f3247f756cb658e…`, round-trip GET verified. Golden **0.185.1** baked on the nested drill VM embedding controller `0.185.1`, published, sha `dba00f3e845c415e…` — bake clean: `Result=success`, overlay2, **all 3 mounts included** (rootfs+mp0+mp1), 0 FATAL/exclusions, upload HTTP 201, token-leak grep 0; log `drill/bake-0.185.1.log`; GL-1 teardown done (guest 9100 purged, secrets shredded, disk restored to `virgin`). Hub Day-0 manifest moved **both together in one POST** so it never vouched a new agent against an old golden; `min_agent` **0.93.0 → 0.113.0**, which is what controller v0.185.0 declares (`felhom-controller/CHANGELOG.md:15`) — **zero fleet impact, verified: all three enrolled hosts already run agent 0.113.0, so no box is held.** `wrapper_sha256` preserved verbatim (re-checked against `configs/felhom-pbs-apply` — no drift). **The global controller floor was deliberately NOT raised**: the golden now bakes 0.185.1, so a fresh box needs no self-update, and raising it would have been an unnecessary fleet-wide write. Original finding follows. **Found 2026-07-29 by the E-2d Phase 0 gate, which stopped the run before a VM was created.** 17 unpublished releases (v0.97.0–v0.113.0) strand the **entire R-82 tiered-backup arc** plus **F-CRIT-2** (a failed backup looking fresh — 7 days silent) and **F-REBOOT** (a guest rebooted mid-backup never returns): a new customer's box would install without them. **Blocks E-2d's C3/C4/C5** — those test endpoints and events that do not exist in 0.96.0/0.161.0. The **controller is fine** (registry has 0.185.1, floor-driven self-update), so the gap is specific to the two Gitea-generic artifacts. **Mirror of R-110, not a duplicate:** R-110 = the installer publishes instantly with no staging; R-111 = the agent/golden publish gate exists and was never walked. Fix should decide whether publishing joins the release train rather than staying a remembered step (R-29's shape, one layer up). Evidence: `audits/E2D-fresh-vm-2026-07-29.md` **DEFERRED LEG, AND IT RECURRED → R-115.** This row's shipped half stands and is not reopened: the bump happened, was verified, and was proven end-to-end by the E-2d install. But its own closing line — *decide whether publishing joins the release train rather than staying a remembered step* — was never acted on, and agent 0.114.0 reproduced the exact condition the same afternoon. The recurrence is filed as **R-115**, not as a reopen, because the stale-channel finding is closed while the process defect that caused it is a distinct problem with a distinct fix. | CC | -| **R-115** | **Publishing is a remembered step, and it was forgotten within eight hours of being documented as forgettable.** A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published — so "deployed" and "installable" are independent states that drift silently. **Two instances, both real:** **R-111** (2026-07-29 morning) — 17 agent releases v0.97.0–v0.113.0 stranded, so a new customer would have installed without the entire R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT; found only because the E-2d Phase 0 gate happened to look. **Agent 0.114.0** (same afternoon) — the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, and **unpublished until this task**, which blocked Session C: a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix | **WAITING-ON-OPERATOR (M)** | operator ruling on the release process | **The finding is the RECURRENCE, not either instance** — both instances are fixed. R-111's own text already named this leg (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; the leg then recurred the same day, which is the evidence that a note is not a mechanism. **Class: → R-29, one layer up** — a control that exists and is never walked; deliberately NOT given its own ID. **The decision is the operator's; the options, mechanisms first:** (a) **publish as a step in the build/release path**, so deployed and installable cannot diverge; (b) **a gate that refuses to deploy a version that is not published+vouched** — the strongest, and it fails closed; (c) a session-end checklist entry; (d) accept it as manual and add a pre-Session-C verification. **(a) and (b) are mechanisms; (c) and (d) are reminders — and R-29's whole finding is that reminders do not hold.** No code this session by design. **THIRD INSTANCE, 2026-08-03 — and it was found by a runbook that had been told there was nothing left to do.** Agent **v0.120.0** — the agent half of the R-165 merge — was built, committed at `cd6e267`, and deployed to BOTH demo hosts, and was **never published**: `GET …/generic/felhom-agent/0.120.0/felhom-agent` → **HTTP 404** (0.119.0 → 200), and the hub manifest accordingly vouched **0.119.0**. The consequence is the sharpest yet, because installer step 5's idempotent skip requires `installed == vouched` EXACTLY: a documented-path reinstall would have **downgraded both boxes** from the merge-aware 0.120.0 to the pre-merge 0.119.0 — silently, since the current `step_grows` sets `SYSDATA_GROW=0` so 0.119.0's `mp1` resize (`bringup.go` 4c, fatal on error) never fires and the install would have *succeeded* while proving a stack nobody ships. **R-178's own row asserted `agent v0.120.0 is live on BOTH hosts` and `no code left to write`; both were true and both were beside the point** — the gap was publication, which no one checks. Fixed in-session on the operator's ruling: `scripts/publish-agent.sh 0.120.0` (sha `a7763d31b55b5ce75457b4dba7b06aa300325811834b0be78af4587b47110b9d`, round-trip GET verified) then vouched, and both reinstalls then fetched and sha-verified it from Gitea. **This is the third instance of a row that has been WAITING-ON-OPERATOR since 2026-07-29; option (b) — a gate that refuses to deploy or vouch an unpublished version — would have caught all three** | CC | +| **R-115** | **Publishing is a remembered step, and it was forgotten within eight hours of being documented as forgettable.** A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published — so "deployed" and "installable" are independent states that drift silently. **Two instances, both real:** **R-111** (2026-07-29 morning) — 17 agent releases v0.97.0–v0.113.0 stranded, so a new customer would have installed without the entire R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT; found only because the E-2d Phase 0 gate happened to look. **Agent 0.114.0** (same afternoon) — the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, and **unpublished until this task**, which blocked Session C: a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix | **READY (M)** — operator ruling taken 2026-08-03: **mechanism (b), build-side gate** | — | **The finding is the RECURRENCE, not either instance** — both instances are fixed. R-111's own text already named this leg (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; the leg then recurred the same day, which is the evidence that a note is not a mechanism. **Class: → R-29, one layer up** — a control that exists and is never walked; deliberately NOT given its own ID. **The decision is the operator's; the options, mechanisms first:** (a) **publish as a step in the build/release path**, so deployed and installable cannot diverge; (b) **a gate that refuses to deploy a version that is not published+vouched** — the strongest, and it fails closed; (c) a session-end checklist entry; (d) accept it as manual and add a pre-Session-C verification. **(a) and (b) are mechanisms; (c) and (d) are reminders — and R-29's whole finding is that reminders do not hold.** No code this session by design. **THIRD INSTANCE, 2026-08-03 — and it was found by a runbook that had been told there was nothing left to do.** Agent **v0.120.0** — the agent half of the R-165 merge — was built, committed at `cd6e267`, and deployed to BOTH demo hosts, and was **never published**: `GET …/generic/felhom-agent/0.120.0/felhom-agent` → **HTTP 404** (0.119.0 → 200), and the hub manifest accordingly vouched **0.119.0**. The consequence is the sharpest yet, because installer step 5's idempotent skip requires `installed == vouched` EXACTLY: a documented-path reinstall would have **downgraded both boxes** from the merge-aware 0.120.0 to the pre-merge 0.119.0 — silently, since the current `step_grows` sets `SYSDATA_GROW=0` so 0.119.0's `mp1` resize (`bringup.go` 4c, fatal on error) never fires and the install would have *succeeded* while proving a stack nobody ships. **R-178's own row asserted `agent v0.120.0 is live on BOTH hosts` and `no code left to write`; both were true and both were beside the point** — the gap was publication, which no one checks. Fixed in-session on the operator's ruling: `scripts/publish-agent.sh 0.120.0` (sha `a7763d31b55b5ce75457b4dba7b06aa300325811834b0be78af4587b47110b9d`, round-trip GET verified) then vouched, and both reinstalls then fetched and sha-verified it from Gitea. **This is the third instance of a row that has been WAITING-ON-OPERATOR since 2026-07-29; option (b) — a gate that refuses to deploy or vouch an unpublished version — would have caught all three.** **OPERATOR RULING 2026-08-03 — mechanism (b), build-side: a gate that REFUSES to deploy or vouch a version that is not published.** It is the strongest of the four options and the only one that fails closed; (c) and (d) were reminders, and R-29's whole finding is that reminders do not hold. Recorded here, **not built this session, by instruction**; it is now CC's to build. **The third instance is the argument for the ruling and belongs inside it:** agent **v0.120.0** (the agent half of the R-165 merge) was built, committed and deployed to BOTH demo hosts while `GET …/generic/felhom-agent/0.120.0/felhom-agent` returned **HTTP 404**, so the hub vouched 0.119.0. Because installer step 5's idempotent skip requires `installed == vouched` EXACTLY, a documented-path reinstall would have **silently downgraded both boxes** to the pre-merge agent — and would have *succeeded* while doing it, since the current `step_grows` sets `SYSDATA_GROW=0` so 0.119.0's fatal `mp1` resize never fires. A gate at deploy/vouch time is the only one of the four standing between that and the operator | CC | | **R-116** | ~~**The drive-absent alarm and its recovery were a MISMATCHED PAIR — absent fired the GENERIC `storage_disconnected`, return the SPECIFIC `backup_target_restored`; `backup_target_absent` never fired at all**~~ | **SHIPPED + PROVEN-LIVE** (agent v0.116.0, 2026-07-30) | — | **CLOSED. The full four-event sequence, on the wire, on a fresh box** (`audits/R116-v0116-2026-07-30.md`): `backup_target_absent (error)` on detach → `backup_target_restored (info)` on return for the TARGET, and `storage_disconnected (error)` → `storage_reconnected (info)` for a NON-target drive on the same box four minutes apart. **Two matched pairs, correctly discriminated — and discrimination is proven NON-trivially for the first time**, since both prior runs had the target itself emit the generic event. Gate fired in **3 s**; all four events reached the hub, so the specific alarm, its severity, its Hungarian copy and the hub routing are now exercised end-to-end. **Over-correction PASSES with a positive observable** (0 ABSENT lines / 0 drive events over 2m14s with both drives present, target `degraded:false`, while 2 `RETURNED` lines prove the gate was ticking). **NARROWED by the R-117 spike (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §12), and it stands as written:** the 2 `RETURNED` lines are a genuine positive observable, so rule 3 is satisfied — but `degraded:false` over that window was read off a drive whose bind was **dead** (R-117), so the window evidences **"the gate did not over-fire"** and **NOT** **"the drive was healthy."** No other part of this row changes: every input to the pairing fix is configuration-derived (`storage.cfg`'s `path` vs the `.mount` unit's `Where`), which R-117 does not touch. Ran on a nested PVE on **demo-hp** per `runbooks/target-selection.md` — through the **real day-0** from the v1.25.0 ISO, with the agent **installed unaided from the vouched Day-0 manifest** (published sha `b47c5c4dab641ee5…`, independent registry GET verified, manifest read back), drives enrolled through the real endpoints, device loss a real hot-detach. **THE FIX, and the ruling is the substantive part:** the mechanism was first isolated from the captured payload (`DIAG-r116-disks-payload-2026-07-30.md`) after two fixes aimed at shapes that do not occur. **Both smaller-looking options were REJECTED because they regress R-114** — `backup_target_offer.go:79` reads `BackupTarget && MountPath != ""` as *"a real drive with its own mountpoint — healthy"* and returns before its `TargetAbsent` branch, so back-filling `MountPath` on the Observe row **or** flagging the registry row (whose `MountPath` is the stale unit-file value) would have told the customer the backup target is fine while its drive was gone. **R-114's correctness was resting on R-116's bug** — a coupling invisible until the payload existed. Taken instead: the Observe row gets the **guest path only** (`mount_path` stays `""`, which is true) from a new `ConfigPath` (`json:"-"`, so the cross-repo golden + key-set contract is untouched), and the union row is deduped **on guest path** — the join being CONFIGURATION (`storage.cfg`'s `path` vs the `.mount` unit's `Where`), the only identity that survives the device. Tests 845→849; 4 red-proofs each asserted to land, and red-proof 1 replays v0.115.0's code and fails, which is the empirical proof it was inert. Its green test had supplied a `MountPath` production never supplies AND left `DriveTargets` nil so the union loop never ran — both corrected. v0.115.0 left in place (inert, harmless). Teardown all 3 layers; hub layer gate-blocked on ONLINE with the command recorded. **Caveat: the drill's controller was 0.185.1 from the golden, which PREDATES R-114**, so its absent-state banner showed the old false copy — the golden being a release behind, not a regression → **R-120** | — | | **R-120** | ~~**The golden baked a controller that predated R-114 + R-112, so a FRESH box showed the customer the WRONG absent-target message**~~ | **CLOSED — golden rebaked + PROVEN-LIVE, and the class now has an ENFORCED gate** (golden 0.186.0 + hub v0.82.0, 2026-07-30) | — | **`audits/R120-golden-rebake-2026-07-30.md`.** **Half 1 — the artifact.** Golden **0.186.0** baked from `main`'s controller in the DooPlex bake fixture (overlay2 OK, **3 mounts**, FATAL 0, exclusions 0, 618 MB, upload **201**, `GOLDEN_SHA256=b760ac6a33e70700…`, token-leak grep 0, GL-1 teardown, `drill.qcow2` back to `virgin`). Three observables: **published** — anonymous GET (what the installer does) 200 / 648930639 bytes / sha identical to the bake; **vouched** — manifest read BACK; **resolved** — `Artifact manifest served for customer sess-f (agent=0.116.0 golden=0.186.0)`. Floor **untouched** per publish-train rule 2 (`min_controller_version` still 0.156.0; it is a separate form); MinAgent left 0.113.0 as 0.186.0 declares. **Proven on a REAL day-0, not the fixture** (per the Part-1 rule now in `runbooks/target-selection.md`): VM 9402 on demo-hp from the v1.25.0 ISO → `Controller elindult (0.186.0)`. With the target detached the endpoint returned the **`TargetAbsent`** copy — *„A rendszermentés meghajtója nem érhető el — amíg vissza nem csatlakoztatod…"* — **and `offer_path` absent entirely**; the day-old read on the 0.185.1 golden had returned the false system-disk message **plus** an offer of the other drive. **Half 2 — the mechanism, operator ruling REFUSE.** hub **v0.82.0**: the gate sits in `hub/internal/web/configs.go` `handleSetArtifacts` immediately before the only write — the sole UI path to `SetArtifactManifest` — so it runs on every vouch without anyone choosing to, and it **refuses** rather than warning. Signal: `store.NewestReportedControllerVersion()` over `reports.controller_version`, **semver-compared in Go** (`MAX()` in SQL ranks 0.99.0 above 0.186.0 — a pair this fleet has shipped). Fail-open in exactly two deliberate cases: empty golden field, unknown fleet version. **NEAR-MISS RECORDED: the first draft read `guests.controller_version`, a column that exists and that NOTHING writes** — it would always have seen `""` and failed open, i.e. inert, this gate's own failure shape, one grep from shipping. 4 tests through the **production handler** over httptest (never a seam), the refusal asserting **both** the flash **and** that the manifest was not written; red-proof: deleting the block makes the stale golden vouchable again. **PROVEN LIVE on the deployed hub by re-attempting the original mistake:** vouching 0.185.1 → `HTTP 303 …flash=golden_behind_fleet` + `[WARN] artifact vouch REFUSED: golden 0.185.1 is older than the newest controller the fleet reports (0.186.0)`, and the manifest read back **unchanged at 0.186.0**. Recorded on **R-29's audit list** (`ROADMAP.md`) as the **first enforced gate** beside its three orphans, so the contrast is kept — the orphans are unchanged. Teardown all 3 layers; hub layer gate-blocked on ONLINE with the command recorded, exactly as `sess-e` was (and `sess-e` was deleted this run) | — | | **R-117** | ~~**A drive's guest bind becomes a DEAD MOUNT while every signal reads healthy — and it happens in TWO ways, only one of which the original framing covered.** (a) *after a detach/return*: the host raw mount heals onto the NEW device via its fs-UUID-keyed unit while the bind still names the OLD one, so the gate takes its `Return` branch and restarts the customer's apps onto a namespace that `EIO`s on every call; (b) *in STEADY STATE, no cycle at all* — a device that errors without disappearing leaves the raw mount `active`, `BoundUnderParent` `true` and the drive never `Disconnected`, so **the gate produces no action and NOTHING is emitted on any channel**~~ | **SHIPPED + PROVEN-LIVE** (agent **v0.117.0**, 2026-07-30) | — | **CLOSED. `audits/R117-v0117-2026-07-30.md`.** `BoundUnderParent` gains a THIRD term at both /disks sites: `bindLiveness` reads `/proc` only and requires (a) **the bind names the same device as the raw mount** and (b) **the filesystem has not aborted** (`shutdown` **or** `emergency_ro`, both measured). **BOTH CHECKS ARE LOAD-BEARING and this is the substantive part:** R-117 was filed as a detach/return defect, but a device that fails WITHOUT disappearing gives the identical all-signals-healthy state with the **devnos EQUAL** and the drive never `Disconnected`, so the gate emits nothing at all, indefinitely (R-117a) — the device comparison alone cannot see it, and a P1-only fix passes every payload test (red-proof RP3 exists for exactly that). **THREE states, never a bool:** `{Unknown, Live, StaleDevice, Aborted}`, `Unknown` is the zero value, and every caller reads `Usable()` where unknown counts **PRESENT** (absent stops a customer's apps — the `newestArchiveOn` trap). **NO NEW RECOVERY PATH:** `AttachDrive`'s normalize leg already did the repair and three call sites already invoked it (20 s ticker, agent startup, and **the controller's `Return` branch BEFORE `restartStacks`**); all three were defeated by `if n == 1 && GuestSeesMount(...)` logging *"fully live, no-op"* about an EIO namespace. **RULING (asked for, given, flagged for overrule):** `StaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, guest never restarts — init PID identical); `Aborted` ⇒ **quiet no-op and SURFACE**, because a re-bind lands on the SAME dead superblock and this runs every 20 s = an infinite silent retry that masks the state. No operator decision required: it routes an already-broken state into the **existing** gate, event types and Hungarian copy — no new customer-facing concept — and the alternative is apps writing documents into a filesystem that rejects every write. **ORDERING TRAP caught by a test:** abort-first classifies the real return state as aborted (its stale bind carries `shutdown` too) and refuses the repair **while still reporting correctly**, so the abort flag is read off the RAW mount in the stale case. **LIVE on demo-hp** (brought 0.113.0 → 0.117.0 first — see R-121): RETURN `raw 8:32 / bind 8:16 shutdown` ⇒ `stale-device`, usable **false**; IN-PLACE `both 252:11 emergency_ro`, raw unit still `active` ⇒ `filesystem-aborted`, usable **false**; healthy ⇒ `live`; **340–497 µs**. **No block I/O proven by strace** (only `/proc/self/mountinfo`, **0** statfs) — the Part 1 `CLAUDE.md` fence applied to its own first consumer. **No regression through the REAL pipeline:** `GET /disks` with the controller's own credential shows the live backup-target drive `bound_under_parent=True`, with 32 gate lines in 3 min as the positive observable and zero spurious transitions. Tests **849→863**, 29/29 green, **6 red-proofs each verified to land** — and **RP1 failing to fail exposed a HOLLOW test**: the aborted fixture used a `/dev/mapper` device, for which `RoleForStorage` derives `role=system`, and a system row never runs the conjunction, so it reported false by DEFAULT and no mutation could fail it. Fixtures now assert the production row shape first. Teardown all 3 layers; hub layer = the vouched manifest, **retained** (it is the product, not scratch). **NOT covered:** the stale-bind repair on hardware — `StablePathForRaw` hardcodes the live parent, so it would write into guest 9201's namespace (R-117h); and sustained-load behaviour, still unmeasured. Follow-ups **R-117g** (no guided recovery for an aborted fs), **R-117h** (parent dir not test-seamable), **R-121** | CC | @@ -84,7 +84,7 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha | **R-137** | **Cloudflare geo-WAF rules are zone-scoped and non-namespaced — four cross-tenant faults.** `globalRuleDesc = "[felhom-geo] Global"` (`waf.go:18`) is one literal description per ZONE; `appRuleDescPrefix` keys by app name with no customer (`waf.go:21`); `BuildGlobalExpression` has no positive hostname scoping (`waf.go:241`); `applyDiff` deletes every `[felhom-geo]` rule not in THIS box's desired set (`geosync.go:320`) | READY (M) — **blocks shared-zone onboarding** | — | With two customers in one zone: they overwrite each other's Global rule forever; one customer's country policy applies zone-wide; per-app rules collide by name; and disabling the feature for one (or the hub's `RemoveGeoRules`) wipes them all. Interim mitigation, no code: keep geo-restriction OFF for every shared-zone customer. Fix = namespace descriptions by `customer_id` + add `http.host ends_with ""` to both expressions — a TWO-REPO change (controller + hub `RemoveGeoRules`). Same audit §5.1 | CC | | **R-138** | **A shared-zone `cf_api_token` is a zone-wide DNS-write capability on a customer's box** — written 0600 to `/opt/docker/stacks/traefik/.env` (`controller/internal/infra/infra.go:123`) | READY (S) | — | Today each box holds a token for a zone nobody else uses, so the blast radius is one customer. Under a shared customer zone, one compromised tester box could repoint every other tester's DNS. The ACME path is already switchable — an empty token selects HTTP-01 (`traefik.yml.tmpl`) — so the fix is policy plus a guard that refuses to hand a shared-zone customer a zone-scoped token. Same audit §5.2 | CC | | **R-133** | **The vaulted break-glass console credential is PLAINTEXT AT REST — every hub DB backup is a fleet-wide console-credential dump.** `host_recovery.secret` holds each managed box's `root@pam` password verbatim, so any copy of the SQLite DB (Longhorn snapshot, PBS backup of the hub PVC, a hand-taken copy during a diagnosis) carries root console access to every Felhom host in one file | **READY (M) — NEW 2026-07-31** | — | **The deferred leg of hub v0.84.0** (Console access card), filed separately because v0.84.0 changed only WHO can retrieve the secret, never how it is stored. v0.84.0 makes it more worth doing, not more broken: retrieval now rides the hub SESSION, so the DB and the login password are jointly the whole protection (ruling **S-4**, `CONTEXT.md`). Fix shape: **envelope-encrypt the `host_recovery.secret` column under a KEK held outside the DB** — the hub already proves it can hold something it cannot itself read (escrow blobs), and that contrast is the argument. Two constraints the design must respect: the credential must stay retrievable **when the box is unreachable** (that is the whole point of break-glass), so the KEK cannot live on the box or depend on the agent; and the global-key API path must keep working with the hub UI down. Would flip the capability-map row **"Break-glass management-plane recovery"**, which today reads IMPLEMENTED with this as its caveat | CC | -| **R-156** | **An app's data is neither persisted nor backed up, and it reports healthy.** A template mounts a volume at a path the application never writes, so the data sits in the container's **writable layer**: lost on redeploy, and tarred nightly as an empty directory while the healthcheck stays green. **papra** (Campaign 10) and **gramps-web** + **wishlist** (the 53-template sweep) all convicted. | **READY (S)** — the class is detected; papra itself is open | — | **The gate SHIPPED**: `app-catalog-felhom.eu/scripts/check-volume-persistence.py` (runtime probe; `docker diff` + mount-occupancy + writability, canary self-test, fails closed). It convicts papra `/app/data`[vol,EMPTY] → `db.sqlite` in the writable layer. **gramps-web and wishlist were FIXED in the sweep; papra was NOT — it is referred**, because the fix needs either the app to use `/app/data` or the template to mount `/app/app-data`. **THE REFERRAL IS RESOLVED, 2026-08-02 — papra is deployed NOWHERE, so the template fix strands nothing and can be applied.** The referral existed because changing where the volume mounts moves live data: an installed papra writes `db.sqlite` into the container's writable layer, and a remount relocates the path out from under it. With no instance deployed there is no live data to move, so the cheaper leg — **the template mounts `/app/app-data`** — is takeable directly, without waiting on upstream to adopt `/app/data`. **Provenance, stated because it decides the row:** the observation is `docker ps -a` on demo-hp's **guest 9201** returning empty, supplied with the 2026-08-02 task; **this session did not re-measure** (documentation-only, every box fenced). **Scope of that evidence, honestly:** it covers guest 9201 — the guest papra was convicted on in Campaign 10 — and **no other customer's guest was enumerated**, so a re-check belongs in the task that edits the template, before it edits it. **Next action: apply the template fix in `app-catalog-felhom.eu` and re-run `scripts/catalog_gates.py`** (deliberately not done here — that repo was out of scope for this task). See R-161 (nothing runs the gate automatically) and R-159/R-160 | CC | +| **R-156** | **An app's data is neither persisted nor backed up, and it reports healthy.** A template mounts a volume at a path the application never writes, so the data sits in the container's **writable layer**: lost on redeploy, and tarred nightly as an empty directory while the healthcheck stays green. **papra** (Campaign 10) and **gramps-web** + **wishlist** (the 53-template sweep) all convicted. | **CLOSED — all three apps fixed** (papra template, 2026-08-03) | — | **The gate SHIPPED**: `app-catalog-felhom.eu/scripts/check-volume-persistence.py` (runtime probe; `docker diff` + mount-occupancy + writability, canary self-test, fails closed). It convicts papra `/app/data`[vol,EMPTY] → `db.sqlite` in the writable layer. **gramps-web and wishlist were FIXED in the sweep; papra was NOT — it is referred**, because the fix needs either the app to use `/app/data` or the template to mount `/app/app-data`. **THE REFERRAL IS RESOLVED, 2026-08-02 — papra is deployed NOWHERE, so the template fix strands nothing and can be applied.** The referral existed because changing where the volume mounts moves live data: an installed papra writes `db.sqlite` into the container's writable layer, and a remount relocates the path out from under it. With no instance deployed there is no live data to move, so the cheaper leg — **the template mounts `/app/app-data`** — is takeable directly, without waiting on upstream to adopt `/app/data`. **Provenance, stated because it decides the row:** the observation is `docker ps -a` on demo-hp's **guest 9201** returning empty, supplied with the 2026-08-02 task; **this session did not re-measure** (documentation-only, every box fenced). **Scope of that evidence, honestly:** it covers guest 9201 — the guest papra was convicted on in Campaign 10 — and **no other customer's guest was enumerated**, so a re-check belongs in the task that edits the template, before it edits it. **Next action: apply the template fix in `app-catalog-felhom.eu` and re-run `scripts/catalog_gates.py`** (deliberately not done here — that repo was out of scope for this task). See R-161 (nothing runs the gate automatically) and R-159/R-160. **CLOSED 2026-08-03 — papra fixed, and the precondition was CHECKED rather than inherited.** The 2026-08-02 evidence covered only guest 9201 and both demo boxes have been wiped since, so it was re-measured three ways: `docker ps -a` (INCLUDING stopped containers) on **both** demo guests → no papra; and the hub's `/hosts` fleet view → exactly **two** enrolled hosts (`demo-felhom-8363b5`, `demo-hp-bb76ea`), **zero** papra references. Deployed nowhere ⇒ the template fix strands nothing. **The fix decided from the IMAGE, not the README or the upstream docs:** `docker inspect` of `ghcr.io/papra-hq/papra:26.6.1-rootless` gives `WORKDIR=/app` and **all three** data paths under `./app-data` — `DATABASE_URL=file:./app-data/db/db.sqlite`, `DOCUMENT_STORAGE_FILESYSTEM_ROOT=./app-data/documents`, `PAPRA_CONFIG_DIR=./app-data` — and **`/app/data` does not exist in the image at all**, so the old mount pointed at a path nothing could ever write. **Departure from the task's stated preference order, recorded because it was deliberate.** Option (1) — reconfigure the app to write where the template already mounted — WAS available (all three paths are env-settable). It was not taken: it enumerates data paths, so a fourth one added upstream would silently escape to the writable layer again, which is R-156's exact failure mode re-armed and invisible. Mounting the app's own data ROOT (`papra_data:/app/app-data`) captures every current AND future path by construction. One line instead of three env vars with an ongoing coupling to upstream. **The runtime gate is the arbiter and it was run, in BOTH directions.** `check-volume-persistence.py papra` → **CLEAN**, and its mandatory self-test passed on that run (*'prober flags the R-156 signature and clears a correct template — trustworthy'*), so the verdict carries a live proof that the instrument discriminates. **Red-proof on the real template, not just the canary:** reverting the mount to `/app/data` and re-running → **BROKEN**, with the exact R-156 evidence — *'mount /app/data is NOT writable by the app's own uid=999'*, *'DATA in the writable layer at /app/app-data/db (db_signature=True, e.g. [db.sqlite])'*, *'declared volume /app/data is EMPTY'*. Restored, re-run, CLEAN. **Two operational notes for the next person to run this gate:** it needs **root** (it reads volume contents under `/var/lib/docker/volumes`, mode `drwx--x---`; as a normal user its own canary self-test fails UNDETERMINED and it correctly refuses to report), and it hardcodes a scratch path `/srv/felhom-gate`. Run it **scoped to the app you touched** — unscoped it deploys all 53 templates and takes far longer than a session allows | — | | **R-157** | ~~**`bootrecon`'s start-ONCE sweep misses the boot orphan it exists to recover — TWO mechanisms.**~~ | **CLOSED — SHIPPED + PROVEN-LIVE** (B: controller v0.189.0; A: v0.190.0, 2026-08-02) | — | **Both mechanisms closed. (B)** the container-count signal → recorded intent (R-166). **(A)** the sweep looked ONCE at T+5 s, deriving candidates from a fleet docker was still restoring — 3 of 6 hard resets. Now a **settle-then-sweep window**: sample the fleet every 5 s, settled after 3 identical samples, sweep ONCE at the end; ends on settled OR a 50 s budget, and the log says which. **The budget is 50 s because a test rejected 60 s**: settle+budget+one 30 s retry must stay under the 90 s `deadAppBootGrace` or a successful recovery stops being silent; 60 s gave 95 s. A window that genuinely overruns emits a `LATE RECOVERY` WARN naming the apps — the grace was NOT widened to hide it (§8.3). **A defect in the fix, found by live validation not review:** `GetStacks()` is the Manager's cache, refreshed by the scheduler every 10 s, so sampling it every 5 s without refreshing let "settled" mean "the cache did not update" — observed missing a container removed 5 s before the window closed. `sampleBootFleet` now refreshes first. **Live: 6/6 hard resets on the shipped build, every app back every time** (settle times 10/40/10/10/15/15 s — i.e. the window routinely waited 2–8× longer than the old fixed 5 s), plus a before/after on ONE app on ONE box: the pre-fix window logged `no boot-orphaned apps` for calibre-web at 18:08:35, the fixed one found and recovered it at 18:18:50 | — | | **R-170** | ~~**The drive-backed boot gate infers a customer's Stop from a container count.**~~ | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.190.0, 2026-08-02) | — | `shouldRecreateOnBoot` now reads `desired_state` with the SAME three-way table as `isBootOrphan`: `stopped` → never; `running` → recreate whatever the container count; **absent → exactly the pre-v0.190.0 `hasContainers` behaviour**. `presentStable` untouched and still load-bearing (an absent drive is never recreated here — the very term the boot sweep was missing, R-171). Its comment argued at length FOR the container count and was rewritten; a correct implementation under a comment arguing the opposite is worse than either alone. **The agreement is pinned from BOTH sides** against one fixture table (`TestBothBootGatesAgreeOnIntent` / `TestShouldRecreateOnBoot_AgreesWithBootrecon`) because the two gates cannot be called from one package without an import cycle. **Live on 9201, both halves in one reboot:** calibre-web (drive-backed, `running`, ZERO containers) → `recreating drive-backed app calibre-web`; immich (`stopped`) → `1 drive-backed app(s) left stopped on purpose` | — | | **R-171** | **The boot sweep started apps whose data drive was ABSENT — a regression introduced by v0.189.0, now FIXED.** Replacing `isBootOrphan`'s container-count term with recorded intent made a drive-gate-stopped app (`compose down` ⇒ zero containers, and the gate never touches `desired_state` because it is not the customer) read as a boot orphan | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.190.0, 2026-08-02) | — | **Reasoned from the diff, then CONFIRMED on hardware before any fix was written** (`audits/DIAG-bootrecon-drive-absent-2026-08-02.md`). The sweep found and started calibre-web with its drive unmounted, burned both attempts and handed it to the dead-app alarm — **a false alarm about an app the drive gate is deliberately holding**. The *write* hazard did NOT materialise: compose failed `mkdir …/userdata: permission denied` because the unbound mountpoint is host-root-owned and the guest is unprivileged — **an accidental protection no code owns, no test pins, and one `chown` or one privileged guest away from gone**. Fix: new consumer-side seam `bootrecon.StartGate`, **fail-safe (cannot determine ⇒ do not start)**, wired in `main.go`; `Manager.DriveLive` reuses the userdata belt's own `isMountPoint` seam so the two cannot drift. **The rule is not new** — the API's `startGatedByMissingDrive` already refused this to the customer; the sweep bypassed it by calling `Manager.StartStack` directly. Widening the window (R-157 A) made two more holders reachable, so the same seam also refuses an app held by a **quiesce** or an **in-flight app-data operation** (§8.2), reusing `quiesce.SuppressedStacks()` and a new read-only `AppStopGuard.HeldStacks()`. Held apps report as `HeldByDrive`, never `StillDown` — that is the alarm's bucket. **ID established free:** `grep -ro "R-171\b" documentation/ *.md` → 0 hits before minting | — | @@ -92,7 +92,8 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha | **R-174** | ~~**The app-stop guard's crash recovery started apps onto MISSING drives — a regression in v0.189.0 code.**~~ | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.191.0, 2026-08-02) | — | **Found by REVIEW on 2026-08-02, in code shipped 2026-08-01, and closed the same session — R-171 one path over.** `appStopGuard.SetStarter(stackMgr)` handed `Recover` the RAW stack manager, whose `StartStack` has no drive gate, and `Recover` runs **at startup** — exactly when an external drive may not have come back. So: a backup stops an app, the box loses power, the drive does not remount, and the app is started on a missing drive. The rule was not new — the API's own `startGatedByMissingDrive` already refused this to the customer; the guard bypassed it. **`bootDriveGate` could NOT be reused whole**, and the reason is recorded in the code: its holder #2 reads `bootAppStopGuard.HeldStacks()`, which during `Recover` is **the guard's own marker** — it would refuse every recovery it was meant to perform — and holders #1/#2 read package-level vars assigned AFTER `Recover()` runs, so a whole-gate reuse would be correct only by accident of nil-safety. Holder #3 is extracted into a shared `driveStartGate` with **two callers, one implementation**, and `TestBootDriveGateAndAppStopShareTheDrivePredicate` pins the delegation. **A REFUSAL IS NOT A FAILURE:** new `ErrStartRefused` + a `Refused` bucket — both keep the marker, only `Failed` alarms, because routing a deliberate hold into `NotifyBackupFailed` (customer-enabled by default) is the very R-171 false alarm this fixes. `main.go` guards on `Alarming()`, not `!= nil`, and the pre-existing seam test was TIGHTENED to require it. **Live on 9201, both directions:** drive held unmounted → `refusing to restart "calibre-web" … drive /mnt/felhom-drives/hdd_1 is not a live mountpoint`, marker retained byte-identical, zero containers started, `not alarming`; drive returned → `restarted calibre-web`, marker CLEARED. **ID established free:** `grep -ro "R-174\b" documentation/ *.md` → 0 hits | — | | **R-175** | ~~**`07-backup-architecture.md` §7.5 states ONE box's size bound as if it were the fleet's.**~~ | **CLOSED — FIXED 2026-08-03** (same pass as R-165) | — | **Measured, not inferred** (`audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1: `pct config 9201` on both hosts). Independent of the merge — the sentence is wrong today and will be wrong differently after R-165. **The fix is to state the bound as a FUNCTION of `mp1`, not a constant**, and to say which box any quoted figure came from. Same class as the comment-asserting-an-invariant rule: a doc stating a fleet-wide number that only one machine satisfies reads as settled and is not. **ID established free:** `grep -ro "R-175\b" documentation/ *.md` → 0 hits **FIXED.** §7.5 gained a **7.5.1** which (a) states plainly that the bound is a FUNCTION of `mp1` and applies only to a box still on the split layout, naming all three real shapes, and (b) records that the ceiling itself has been removed by R-165 for boxes built from golden ≥ 0.192.0. Fixed in the same pass as the merge rather than filed and forgotten, because the section would otherwise have been wrong in two ways at once | CC | | **R-176** | **Two prerequisites for the R-165 merge are UNMEASURED, and both are cheap.** (a) Whether a **pre-merge archive** (carrying `mp1`) restore-tests cleanly into a **merged-layout** guest — reading `mountParity` (`felhom-agent/internal/reconcile/restoretest.go:347`) says it should, because the restore recreates `mp1` from the archive so archive and restored guest agree; **that was reasoned from source and never executed.** (b) The in-place per-box migration (move `/felhom-data` onto `mp0`, drop the slot, verify) has **never been rehearsed even once**, so "is the box restorable at every point of it?" is currently unknown | **(a) ANSWERED 2026-08-03 (P1: PASS). (b) NOT REQUIRED — operator ruling: every node is reinstalled, none migrated** | blocks R-165 landing safely | **Filed because this project's own record is that FOUR production designs specced against unvalidated mechanisms were all wrong** — which is exactly why R-165's own spike refused to design. Both are one command on a **Tier-0** box (D-d: both demo boxes are disposable). (b) is only required work if Peti's box turns out to need migrating rather than reinstalling — the hub cannot answer that (M5: `peti-felhom` exists as a customer with **no host in the register**), so it is the operator's input. **ID established free:** `grep -ro "R-176\b" documentation/ *.md` → 0 hits **UPDATE 2026-08-03.** **(a) is measured and passed** — `audits/SPIKE-r165-phase0-2026-08-03.md` P1: a real pre-merge archive (`mp0+mp1`, confirmed from its own vzdump log) restore-tested on demo-hp, `pass: true`, `mount_parity: ok`, 84 s, with `mountParity` untouched. One limit stated rather than glossed: it ran with the pre-merge agent because the merged one did not exist yet, and the comparison is archive-vs-its-own-restore which never consults the host layout — **re-run it once against agent v0.120.0**, which is one command. **(b) is withdrawn, not deferred:** the operator ruled that every node is REINSTALLED rather than migrated in place (both demo boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed in a few weeks), so the in-place migration rehearsal has no consumer. Recorded explicitly rather than silently skipped | CC | -| **R-181** | **The capture floor guards the cheap leg and not the leg that fills the volume — and its refusal message asserts an invariant the code does not provide.** B2 (controller v0.192.0) is recorded on R-165 as the deliberate replacement for the bulkhead the `mp1` partition used to give. It is consulted in exactly one place — `m.unitFloorBlocked(stack.Name)` at `recovery_unit.go:328`, inside `captureAllRecoveryUnits`, which writes a manifest and a compose copy: **a few KB.** The leg that writes the bulk, `runVolumeDumps` (`backup.go:535`), has **no floor check at all** — its gates are protected-stack, volume-less, disconnected, decommissioned — and it runs FIRST, by design (*"MUST run before captureAllRecoveryUnits so the manifests enumerate the fresh tars"*, `backup.go:483`). So the write that fills the filesystem is unguarded, and the floor then refuses the write that would have cost almost nothing. **Second limb: the refusal message is false.** `recovery_unit.go:331` prints *"the previous unit is untouched and NOTHING was deleted"*. Nothing was deleted — true. Untouched — **measured false**: privatebin's `volume-dumps/privatebin_privatebin_data.tar` went `26c546c2…` → `b538ab89…` and opengist's went **182,272 B → 2,147,666,432 B**, both rewritten by the earlier leg, while each unit's `manifest.json` kept `created_at: 2026-08-03T06:34:26Z` and its `checksums` block covers only the three compose files — so a unit's payload can be swapped under a stale descriptor and **nothing in the unit can detect it** | **READY (M) — NEW 2026-08-03** | blocks **R-165** reaching PROVEN-LIVE | **FIRST LIVE FIRING OF B2, and it is why the runbook asked for one.** Proven on demo-hp 2026-08-03 06:40:03 on a box reinstalled from the merged golden (R-178). Method: a real 2 GiB file in opengist's data volume, then `fallocate` to bring the filesystem to 96 % used / 3.0 GiB free — both floor terms deliberately still clear, so the run started. **The `fallocate` instrument was proven before use** (5 GiB moved guest `df` 977M→6.0G while thin-pool `data_percent` stayed 29.03 → 29.03: zero blocks allocated), because demo-hp's thin pool is 53.93 GiB and a real fill to 97 % of a 70 G volume would have exhausted it and corrupted every guest on the box including the `drill-r50` fixture. Sequence observed: opengist's volume dump wrote **2.0 GB unguarded** → free fell to 1.0 GB → **both** apps' recovery-unit captures were then REFUSED on the `1.0 GiB free` term, each pushing `recovery_unit_capture_failed` (severity `error`) to the hub, accepted HTTP 200. **What DOES hold: it refuses per app rather than aborting the run, it never deletes, and the alert reaches the operator.** **Fix shape, not written this session by design (§7 of the runbook):** the floor belongs before the write in `runVolumeDumps` too, the message must stop claiming what the earlier leg has already falsified, and per `CLAUDE.md` *"a comment asserting an invariant needs a test pinning it"* the pinning test must assert the **consequence** (after a refusal, is the previous unit's payload byte-identical?) and not the mechanism. **Class:** the sixth entry in `CLAUDE.md`'s own table of shipped guarantees the code did not provide — found, as four of those were, only on live hardware | CC | +| **R-182** | **The reserve re-alerts on every status refresh, so a full disk pages the operator on a timer rather than once.** `captureAllRecoveryUnits` has TWO callers: the backup run (`runDBDumpsInternal`), which since R-181 holds a per-run admission memo so a refused app alerts exactly once — and `GetFullStatus` (`backup.go:967`), the **periodic status refresh**, which calls it with no run scope. Outside a run the verdict is decided fresh, so every refused app pushes another `recovery_unit_capture_failed` (severity `error`) on every refresh | **READY (S) — NEW 2026-08-03** | — | **Measured live, not inferred.** demo-hp 2026-08-03: the backup run at **08:59:46** produced exactly one alert per app (2 apps → 2 events, correct); a **second identical pair fired at 08:59:59**, thirteen seconds later, from the status refresh. Under a sustained full disk that is a repeating operator page for a condition already reported. **PRE-EXISTING, NOT INTRODUCED BY R-181 — stated because it decides the priority.** v0.192.0's `captureAllRecoveryUnits` called `unitFloorBlocked` → `unitNotify` on the same path with the same frequency; R-181 changed neither caller. It is filed now because R-181's live proof is the first time anyone watched the event stream closely enough to see it. **Deliberately NOT fixed in the R-181 task**, whose §9.2 mandates minimal changes and whose §15 routes observations here; the fix changes alerting semantics on a path that task did not own. **The mitigating claim should be checked before the fix is scoped:** `backup.go`'s own comment says *"NO CONTROLLER-SIDE COOLDOWN — the hub owns cooldown"*. If the hub genuinely dedupes `recovery_unit_capture_failed`, this is log noise and an S; if it does not, it is a real repeated page. **That is exactly the shape of a comment asserting an invariant with no test pinning it** (`CLAUDE.md`'s own table), so verify it at the hub rather than trusting it. **Fix shape (not implemented):** give the status-refresh sweep its own admission scope, or make the non-run path decide without alerting — the run path is already correct and must not be disturbed | CC | +| **R-181** | **The capture floor guards the cheap leg and not the leg that fills the volume — and its refusal message asserts an invariant the code does not provide.** B2 (controller v0.192.0) is recorded on R-165 as the deliberate replacement for the bulkhead the `mp1` partition used to give. It is consulted in exactly one place — `m.unitFloorBlocked(stack.Name)` at `recovery_unit.go:328`, inside `captureAllRecoveryUnits`, which writes a manifest and a compose copy: **a few KB.** The leg that writes the bulk, `runVolumeDumps` (`backup.go:535`), has **no floor check at all** — its gates are protected-stack, volume-less, disconnected, decommissioned — and it runs FIRST, by design (*"MUST run before captureAllRecoveryUnits so the manifests enumerate the fresh tars"*, `backup.go:483`). So the write that fills the filesystem is unguarded, and the floor then refuses the write that would have cost almost nothing. **Second limb: the refusal message is false.** `recovery_unit.go:331` prints *"the previous unit is untouched and NOTHING was deleted"*. Nothing was deleted — true. Untouched — **measured false**: privatebin's `volume-dumps/privatebin_privatebin_data.tar` went `26c546c2…` → `b538ab89…` and opengist's went **182,272 B → 2,147,666,432 B**, both rewritten by the earlier leg, while each unit's `manifest.json` kept `created_at: 2026-08-03T06:34:26Z` and its `checksums` block covers only the three compose files — so a unit's payload can be swapped under a stale descriptor and **nothing in the unit can detect it** | **CLOSED — SHIPPED** (controller v0.193.0 + v0.193.1, 2026-08-03) | unblocks **R-165** | **FIRST LIVE FIRING OF B2, and it is why the runbook asked for one.** Proven on demo-hp 2026-08-03 06:40:03 on a box reinstalled from the merged golden (R-178). Method: a real 2 GiB file in opengist's data volume, then `fallocate` to bring the filesystem to 96 % used / 3.0 GiB free — both floor terms deliberately still clear, so the run started. **The `fallocate` instrument was proven before use** (5 GiB moved guest `df` 977M→6.0G while thin-pool `data_percent` stayed 29.03 → 29.03: zero blocks allocated), because demo-hp's thin pool is 53.93 GiB and a real fill to 97 % of a 70 G volume would have exhausted it and corrupted every guest on the box including the `drill-r50` fixture. Sequence observed: opengist's volume dump wrote **2.0 GB unguarded** → free fell to 1.0 GB → **both** apps' recovery-unit captures were then REFUSED on the `1.0 GiB free` term, each pushing `recovery_unit_capture_failed` (severity `error`) to the hub, accepted HTTP 200. **What DOES hold: it refuses per app rather than aborting the run, it never deletes, and the alert reaches the operator.** **Fix shape, not written this session by design (§7 of the runbook):** the floor belongs before the write in `runVolumeDumps` too, the message must stop claiming what the earlier leg has already falsified, and per `CLAUDE.md` *"a comment asserting an invariant needs a test pinning it"* the pinning test must assert the **consequence** (after a refusal, is the previous unit's payload byte-identical?) and not the mechanism. **Class:** the sixth entry in `CLAUDE.md`'s own table of shipped guarantees the code did not provide — found, as four of those were, only on live hardware. **CLOSED 2026-08-03 — controller v0.193.0 (`fef07c3`) + v0.193.1 (`6c43bf6`), proven live on demo-hp.** **The fix is ONE admission verdict per app per run** (`internal/backup/admission.go`), taken before that app's FIRST write and consulted by all three legs — the three write under one per-app root (`appbackup.RecoveryUnitPath`), which is exactly why one verdict can honestly cover them. **Decided LAZILY at the app's first write, never once at run start**: app A's dump can put app B under the reserve, so a run-start verdict would read a disk that no longer exists — the same class of mistake one level up. **Never re-decided between an app's own legs** (that IS the split this closes) and **reset per run**. Placed ahead of `DumpAppVolumesSafe`, which stops the stack as its first act, so a refused app is never bounced; placed AFTER the volume-less check, which has no write to gate. Exactly ONE operator alert per refused app per run. Leg order unchanged. **The floor is now SIZE-AWARE**, which is the second half of the defect: it asks whether THIS app's write would cross the reserve, not only whether the filesystem is already below it — the term whose absence admitted an app at 96% and then let it write 2 GB. Estimate = the app's previous `.sql`+`.tar` on disk; **no history → headroom-only** deliberately, or the first backup becomes the one that can never happen, and the alert says so. **A container-based `du` was MEASURED and rejected, not assumed**: 66 timed runs on demo-hp guest 9201, **median ~355 ms/volume (341–404)** on volumes holding tens of KB — the cost is container start-up, not the walk. Decisive on top: `docker run` needs the writable layer, so the instrument can fail under exactly the pressure the reserve exists to handle; and the previous-dump estimate measures the ARTIFACT that will be written rather than the live volume. **THE MESSAGE WAS NOT WEAKENED — the behaviour moved so the wording became true**, and it is checked by sha256 tree fingerprint, not by reading the log line (which is what lied). **LIVE PROOF, demo-hp guest 9201, the same method that found it.** The `fallocate` instrument was RE-PROVEN on the rebuilt box before use (guest `df` 1.2G→6.2G on a 5 GiB step while thin-pool `data_percent` stayed **36.83 → 36.83**: zero blocks allocated), because a real fill of a 70 G volume would exhaust the 53.93 GiB pool. **Headroom term @ 08:59:46** — 906 MB free / 99%: both apps refused, **`TREE_SHA` 111d1760c18d3440f700634ab325f8b8 IDENTICAL before and after** (10 files, incl. opengist's tar still at 182,272 B — R-181's own 'before' figure), **no `Stopping for safe volume dump` line at all** (it is present in the 08:58 baseline, which is what makes its absence evidence), 0 volume dumps, one `recovery_unit_capture_failed` per app HTTP 200. **Freed and re-run @ 09:01:33** — both captured normally. **SIZE term proven separately @ 09:03:00**, reproducing the original sequence with a real 2 GiB file in opengist's volume (its previous tar then **2,147,666,432 B**, the exact live figure) and the filesystem at **91% used / 2.9 GB free — both headroom terms deliberately clear**: opengist refused `(size)` — *"this app's last backup was 2.0 GB and writing it again would cross the reserve"* — while **privatebin was ADMITTED and dumped normally**, proving the term is per-app and not a global halt. **Teardown complete**: fill removed, planted file removed, `pct fstrim 9201` returned 67.5 GiB, pool **29.43%** (below the 36.83% baseline), tree byte-identical to the pre-test fingerprint. **v0.193.1 shipped in the same session**, found by this very proof run: the estimate was rendered fixed to 2-decimal GiB, so opengist's real **178 KB** printed as `estimated 0.00 GiB write` — which reads as *no estimate was available* and is the opposite of what happened. Rendering moved to `humanizeBytes`; arithmetic unchanged. Re-verified live: `estimated 178.0 KB write`. **11 new tests + 4 red-proofs**, each demonstrated failing then restored: both dump-leg gates removed (= v0.192.0) → Scenario A red with the tree shown changing; the size term removed → Scenario D red; a prune injected into the refusal path → Scenario F red; the floor moved above the warning band → Scenario G red. **Recorded honestly: the specified Scenario-F mutation (remove the reserve entirely) did NOT turn F red** — removing it makes every app write, which overwrites and adds but deletes nothing, so a deletion-watching test correctly stays green; the prune mutation is the one that proves the assertion. The DB leg cannot run without Docker, so its gate is pinned by an **AST walk** of `backup.go` asserting `admitApp` precedes `DumpOne` — `strings.Contains` is insufficient, a commented-out call still contains the string. **§3's correction CONFIRMED in passing and not chased**: `restore_points.go:57-59` takes the manifest mtime then `newestArtifact` over `.sql` and `.tar`, so the newest of the three wins — the restore point does NOT show a stale timestamp. **New finding from the live run → R-182.** | — | | **R-180** | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** `felhom-host-install.sh` validates the archive storage EXISTS (`pvesm status --storage`, `:1583`) and that the golden volid RESOLVES on it (`:1661`), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default `local local-lvm felhom-pbs` (`--acl-storages`, which `runbooks/day0-install.md` tells the operator **not** to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step | **READY (S) — NEW 2026-08-03** | — | **Hit live on demo-hp 2026-08-03** during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on `felhom-backup` (the enrolled NVMe, where the box's vzdumps live) and `--archive-storage felhom-backup` passed. Pre-flight passed; steps 1–7 ran; step 8 returned `reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)`. **The cost is the ORDER, not the error** — by the time it fires, step 2 has minted the PVE token, step 4b has **rotated root@pam and vaulted it** (so the old console password is already dead), and step 5 has installed the agent. Recovery was `--resume` after moving the golden to `local`, which worked cleanly. **This is statically checkable in pre-flight**: `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a one-line assertion over two variables both known at `:1583`. Same class as R-29 — the checkable thing that nothing checks | CC | | **R-179** | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** The teardown's residue-diff provenance (`day0-install.md` Part E: *"a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers"*) is from **v1.9.1**, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps `/etc/systemd/system/mnt-felhom\x2ddrives-.mount` and `.automount` after a full uninstall | **READY (S) — NEW 2026-08-03** | — | **Observed on demo-hp 2026-08-03** after `--uninstall --vmid 9201`: `mnt-felhom\x2ddrives-Felhom\x2dShare.automount` **loaded failed failed**, its `.mount` `loaded inactive dead`, and `mnt-felhom\x2ddrives.mount` still `active mounted` — the uninstall's own output had warned `/mnt/felhom-drives/Felhom-Share is busy — NOT forcing` and `/mnt/felhom-drives root bind left mounted`, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, `daemon-reload`, unmount the autofs then the parent. **NEGATIVE CONTROL, same day:** demo-felhom's uninstall left **nothing** (`ls /etc/systemd/system | grep -i felhom` → only the unrelated `felhom-bootstrap.service`; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** `felhom-bootstrap.service` is NOT residue — it is the ISO first-boot unit, `disabled`+`inactive`, exactly-once and already fired | CC | | **R-178** | **The merged golden (0.192.0) is built and published but NO BOX HAS BEEN REINSTALLED FROM IT, and it is deliberately UNVOUCHED.** `build-golden.sh` v3.0.0 baked it with variant V-c and every retargeted assertion passed on the real bake (`including mount point mp0 ('/var/lib/felhom')`, no `mp1` line, `both paths are ONE filesystem`); it is in the registry (HTTP 200, sha `54e2a4c431daf580…`). What has NOT happened is Part 4: reinstall each demo box from it and prove claim → deploy an app → back up → restore | **CLOSED — BOTH BOXES REINSTALLED AND PROVEN (2026-08-03)** | blocks R-165 reaching PROVEN-LIVE; blocks the capability-map row | **The golden is UNVOUCHED ON PURPOSE and that is the safe state**, not an oversight: vouching is what makes a fresh install pick it up, so vouching a golden no box has been proven from would put an unproven disk layout in front of the next install anywhere. **Prove first, then vouch** — the bake script's own output treats the hub record as a separate deliberate step for this reason. **Everything else for the merge is shipped and green:** controller v0.192.0 (the B2 floor) is live on 9201, agent v0.120.0 is live on BOTH hosts, and `felhom-host-install.sh` computes the single grow from the thin pool. So a reinstall is now a self-contained piece of work with no code left to write. **Order matters: ONE box at a time**, demo-hp first, proven end to end, and only then demo-felhom — two in parallel leaves no working reference to compare against. Note demo-felhom carries the PBS-DR/offsite tier, so it is the one whose backup chain a reinstall actually disturbs. **ID established free:** `grep -ro "R-178\b" documentation/ *.md` → 0 hits. **CLOSED 2026-08-03 — both boxes reinstalled from the merged golden, by two DIFFERENT supply paths, and proven end to end** (`REPORT.md`). **demo-hp — the layout proof**, installed with `--golden local:backup/vzdump-lxc-9100-2026_08_03-07_33_00.tar.zst` (installer v1.22.0, sha `ed02acb2…`, byte-identical to the repo copy): `mp0 …mp=/var/lib/felhom,backup=1,size=70G`, **no `mp1`**; `/var/lib/docker` → `…disk--1[/docker]` and `/mnt/sys_drive` → `…disk--1[/sys_drive]`, both real mounts, both writable, both in `/etc/fstab`; ONE `df` figure (69G/65G) and `stat -c %d` = `64519` on all three paths; **reboots 3/3** (08:13:33 / 08:13:59 / 08:14:19, controller healthy in 12s/7s/7s, all three still mountpoints after each). **demo-felhom — the pipeline proof**, installed with `--force-gitea-golden` and NO local golden used (preflight logged *"golden: none local — will fetch + verify from Gitea in step 7/8"*, bypassing the 06:58 bake artifact sitting on the same box): **`verified sha256 54e2a4c431daf580… matches the hub manifest`** for the golden and **`verified sha256 a7763d31b55b5ce7…`** for the agent — the observable this second box exists to produce; `mp0 …size=250G`, `grep -c '^mp1:'` → 0, one `df` figure (246G/233G), reboots 3/3 (09:18:55 / 09:19:12 / 09:19:30). **Journey proven on BOTH**, endpoint-level (no browser on DooPlex — the exact endpoints the dashboard's own JS calls): claim (`POST /claim` with the pre-auth HMAC CSRF + `felhom_claim_csrf` cookie; gate discriminator flipped `dashboard not yet claimed` → `authentication required`) → deploy (`POST /api/stacks//deploy`) → capture (`POST /api/debug/backup/dbdump`, which runs the production `RunDBDumps`) → **restore** (`POST /backup/restore`): a planted marker deleted from the live volume came back with an **identical sha256** on each box (`ac1faae6…ae861` privatebin/demo-hp in 9.2s; `bc550798…b59e939` opengist/demo-felhom in 9.4s), recovery units on the single volume in both cases. **Ceiling gone, measured:** 65 GiB (demo-hp) and 233 GiB (demo-felhom) available to a recovery unit, against the 19 GiB and 45 GiB their pre-wipe `mp1` slices offered. **Two deviations, both the operator's call and both recorded:** the golden was ALREADY vouched when the session opened (hub log `2026/08/03 07:23:26 Artifact manifest set: agent=0.119.0 golden=0.192.0`, ~10 min before the first read of this session — so §7's prove-then-vouch order was already spent and the operator elected to accept it); and agent **0.120.0 had never been published**, so the vouched agent was 0.119.0 — published + vouched before the reinstalls (→ **R-115** third instance). **Three new findings: R-179, R-180, R-181** | — | diff --git a/documentation/backlog/ROADMAP.md b/documentation/backlog/ROADMAP.md index 55715848..9631888e 100644 --- a/documentation/backlog/ROADMAP.md +++ b/documentation/backlog/ROADMAP.md @@ -20,7 +20,7 @@ | ID | Item | Size | Status | Notes / map rows flipped | |----|------|------|--------|--------------------------| -| R-115 | **Publishing is a remembered step — forgotten within eight hours of being documented as forgettable** | M | idea — **WAITING-ON-OPERATOR**, 2026-07-29 | A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published, so **"deployed" and "installable" are independent states that drift silently**. **Instance 1 — R-111** (morning): 17 agent releases v0.97.0–v0.113.0 stranded; a new customer would have installed without the whole R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT. Found only because the E-2d Phase 0 gate happened to look. **Instance 2 — agent 0.114.0** (same afternoon): the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, never published — which blocked Session C, since a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix. **The finding is the RECURRENCE, not either instance** — both are fixed. R-111 named this leg in its own text (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; it recurred the same day, which is the evidence that **a note is not a mechanism**. **Class: → R-29, one layer up** (a control that exists and is never walked) — deliberately NOT given a second ID. **Filed as its own item rather than reopening R-111** because R-111's finding (the channel *was* stale) is closed and verified end-to-end by the E-2d install, while the process defect that caused it is a distinct problem with a distinct fix and a distinct owner. **Operator's decision, mechanisms first:** (a) publish as a step in the build/release path so deployed and installable cannot diverge; (b) a gate that refuses to deploy an unpublished+unvouched version — strongest, fails closed; (c) a session-end checklist entry; (d) accept manual + a pre-Session-C verification. **(a)/(b) are mechanisms, (c)/(d) are reminders — and R-29's whole finding is that reminders do not hold.** No code written when filed, by design | +| R-115 | **Publishing is a remembered step — forgotten within eight hours of being documented as forgettable** | M | **READY** — operator ruling 2026-08-03: **mechanism (b), a build-side gate that REFUSES to deploy or vouch an unpublished version.** THIRD instance the same day (agent v0.120.0 deployed to both boxes while unpublished; a documented-path reinstall would have silently downgraded them and *succeeded*). CC's to build | A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published, so **"deployed" and "installable" are independent states that drift silently**. **Instance 1 — R-111** (morning): 17 agent releases v0.97.0–v0.113.0 stranded; a new customer would have installed without the whole R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT. Found only because the E-2d Phase 0 gate happened to look. **Instance 2 — agent 0.114.0** (same afternoon): the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, never published — which blocked Session C, since a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix. **The finding is the RECURRENCE, not either instance** — both are fixed. R-111 named this leg in its own text (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; it recurred the same day, which is the evidence that **a note is not a mechanism**. **Class: → R-29, one layer up** (a control that exists and is never walked) — deliberately NOT given a second ID. **Filed as its own item rather than reopening R-111** because R-111's finding (the channel *was* stale) is closed and verified end-to-end by the E-2d install, while the process defect that caused it is a distinct problem with a distinct fix and a distinct owner. **Operator's decision, mechanisms first:** (a) publish as a step in the build/release path so deployed and installable cannot diverge; (b) a gate that refuses to deploy an unpublished+unvouched version — strongest, fails closed; (c) a session-end checklist entry; (d) accept manual + a pre-Session-C verification. **(a)/(b) are mechanisms, (c)/(d) are reminders — and R-29's whole finding is that reminders do not hold.** No code written when filed, by design | | R-116 | **The drive-absent alarm and its recovery are a mismatched pair — generic on the way out, specific on the way back** | S | idea — **PROVEN LIVE 2026-07-29** | Absent fires `storage_disconnected`; return fires `backup_target_restored`. `backup_target_absent` never fires at all (count 0 across a full Session-C run), so an operator gets an alarm they cannot match to its recovery — exactly what `notifyDriveReturned`'s own comment forbids. Root cause: `notifyDriveAbsent` (`intermediary.go:635-646`) branches on `isTarget[a.Path]` with `a.Path` the GUEST path, and `driveTargetByPath` (`:602-616`) builds it as `out[GuestPath] = d.BackupTarget` — but **the drive is TWO `/disks` rows and the flag and the guest path sit on different ones**: the `felhom-backup` storage row has `BackupTarget: true` (`felhom-agent/internal/localapi/disks.go:211`) and gets a guest path only while classified user-data, while the registry union row has the guest path and **never assigns `BackupTarget`** (`disks.go:265-267`). Absent ⇒ the flagged row loses its guest path ⇒ the union row writes `false` ⇒ generic. On return the rows rejoin ⇒ specific. v0.184.1 fixed the KEYING, not this. **Only reachable because R-113 made the gate fire at all.** Fix likely agent-side; decide the repo first. Blocks E-2's C5. Evidence: `audits/SESSION-C-2026-07-29.md` §5 | | R-113 | **The drive-absent gate cannot fire on device loss — E-2b's alarm is wired to an unreachable condition** | M | idea — **PROVEN LIVE 2026-07-29** | `planDriveGates` (`felhom-controller/internal/web/intermediary.go:216-262`) treats a path as present by OR-ing in `d.BoundUnderParent`, which the agent derives from `GuestSeesMount()` — *"is this path a mount target in the guest's `/proc//mountinfo`"* (`internal/localapi/disks.go:210`). The raw drive mount is a **device-bound systemd unit** and dies with the device; **the agent's own bind under the shared parent is not device-bound and its mountinfo entry outlives the device**, so the gate reads it as present and `notifyDriveAbsent` is never called. Live on a fresh box: target drive hot-detached, agent said `enrolled drive absent by UUID` every 20 s for 4½ min, controller logged **0** `[gate]` lines, hub received **zero** events — neither `backup_target_absent` nor the generic `storage_disconnected`. Not a virtualisation artefact (device-bound-mount vs manual-bind is the same on metal); caveat: SCSI hot-detach, physical unplug not staged. **Sixth instance of seam-built-but-never-wired — E-2b wired the seam to a condition that cannot occur.** Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.2 | | R-112 | **E-2's degraded banner and offer have no UI consumer — correct endpoint, invisible to the customer** | S | idea — **PROVEN LIVE 2026-07-29** | `GET /api/storage/backup-target` returns byte-exact Hungarian copy (verified on a live box), and nothing in the product asks for it: `grep 'backup-target'` across every `*.html`/`*.js`/`*.css` → **0 hits**; no template references `OfferPath`/`Degraded`/the copy; `resolveBackupTargetState` and `degradedMessageFor` are consumed **only** by the JSON handler, with **no page handler injecting the state**. Decisive contrast: the templates fetch **18 distinct `/api/storage/*` endpoints** — `backup-target` and `backup-target/assign` are the only two with zero references. The handler's own comment calls itself *"the dashboard's source for the degraded banner and the offer"*. v0.185.1 shipped as *"the offer endpoints were mounted where nothing routed to them"* and fixed the **mount**, stopping one layer short of the **render**; its test pins dispatch, not reachability. **Fifth instance of the class. Fix R-114 first** — wiring this alone starts showing customers a wrong message. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.1 | @@ -157,18 +157,18 @@ Self-resolves the moment the target answers (the storage read succeeds, sees the | R-97 | **The whole-guest backup tier has NO failure signal to the hub — `internal/quiesce` never notifies** | S | **SHIPPED (controller v0.177.0 + hub v0.78.0, 2026-07-27)** — **R-97a:** `quiesce.TierNotifier`, a seam (not an import) wired by an init-only setter, edge-triggered on the R-88 breaker ARMING so a failing tier is reported once per run rather than once per retry; recovery rides `recordSuccess`'s existing bool. **NEW operator-only event types** `whole_guest_backup_failed`/`_recovered` — deliberately NOT `backup_failed`, which carries a customer Hungarian template AND sits in demo-felhom's live `enabled_events`, so reusing it would have emailed the CUSTOMER about a backup they cannot act on while it was still retrying. The recovery joins `recoveredPairedDownTypes` because its `info` severity would otherwise be dropped by `severityNotifies` — the operator would hear it break and never hear it heal. **The hub's operator cooldown was keyed `customerID:eventType` alone**, so one tier would have masked the other for an hour; now narrowly extended with a `tier` suffix taken from the event details, leaving every other event type unchanged. **R-97b:** a suppression window keyed to the quiesce CYCLE (not a state test — v0.164.0's `!= StateStopped` filter cannot see an app caught MID-RESTART, which is exactly how BookStack alarmed), consumed at the same single derivation point `classifyRunStates`. Grace = **180 s**, derived from the deploy flow's 120 s health timeout and Mealie's 60 s `start_period`; it **expires**, so an app that genuinely fails to come back still alarms. **PROVEN LIVE end-to-end with a control:** the new type POSTs 200 from inside guest 9201 while a bogus type 400s, and `notification_log` shows **1 operator row, 0 customer rows**. The quiesce→notify link itself is unit-proven only. | On 2026-07-27 three whole-guest backups failed and three quiesce cycles stopped and restarted every customer app stack, and **not one `backup_failed` event reached the hub.** It is not the allowlist — the hub already carries `backup_failed` and `backup_completed` (they are emitted by the controller's *app-data* backup path). The cause is that **`internal/quiesce` does not import `internal/notify` at all**: the tier R-82 built has no route to the hub, so a whole-guest backup can fail indefinitely in silence. The loop's only trace was `app_start_failed` — **info** severity, **Hungarian**, on the **customer** channel — telling the customer BookStack was down (it had been caught mid-restart by the third cycle) without saying why, during an outage the system itself caused. So the one signal that did fire was both the wrong tier and the wrong story. **Shape:** emit `backup_failed`/`backup_completed` from `quiesceAndPollTiers` naming the TIER, operator-tier; and decide whether a quiesce-induced restart should suppress `app_start_failed` the way controller **v0.164.0**'s deliberate-stop filter does — an app the backup stopped on purpose is not a fault. R-88's breaker bounds the repetition but changes nothing about the silence | | R-95 | **The restic offsite tier's credential CAN DELETE — R-89's "parallel question", now ANSWERED** | M | idea — established read-only 2026-07-27 | **The exposure closed on the weekly PBS tier is fully open on the daily restic tier**, which holds the customer's actual documents and photos and is the only tier that survives losing the box. Established without mutating anything: **(1) Identity** — a per-customer *subaccount* on `storage-box-pool-1` (box 611714, bx11, `u629488`): `u629488-sub1` home `felhom-demo-felhom`, `sub2` peti-felhom, `sub3` demo-hp, each labelled `felhom-customer`. Auth is an **SSH key stored ON THE BOX** (`…/felhom-controller-data/_data/data/offbox/ssh_key`, 0600, beside `repo_password` + a pinned `known_hosts`) — customer-side, not hub-side, so a compromised guest holds it. **(2) Read-write: YES** — the API reports **`readonly=False` on all three subaccounts**, and it is not merely latent: the controller runs `restic forget --group-by host,tags --keep-daily 7 --keep-weekly … --prune` **from the box** (`backup/offbox.go:984`, also `:1070`). Delete rights are exercised on every run. **(3) Append-only: NO, and not expressible** — the repo is built as `sftp:` (`offbox.go:482`); restic's append-only mode requires the **REST server** backend, which plain SFTP cannot provide. **(4) A zero-code mitigation exists and is unused:** the box type carries `snapshot_limit=10` and the API reports `snapshot_plan=null` with **0 snapshots** and `size_snapshots=0`. Hetzner Storage Box snapshots are taken **server-side, outside the SFTP namespace** — an SFTP subaccount cannot delete them — so they are a genuine immutability layer at no extra cost and with no code change. **Rule once for both tiers, per R-89.** Options, cheapest first: enable a snapshot plan (operator click, immediate); split backup-write from prune so pruning runs somewhere the box cannot reach; or move the repo to restic's REST server with `--append-only`. Flips the capability-map row for offsite immutability | | R-94 | ~~A hand-synced version constant drifts, and the gate that would catch it is never run~~ | XS | **CLOSED — SHIPPED hub v0.87.0, 2026-08-02** | Closed by **deleting** the label rather than deriving it: the Setup command fetches the installer at run time from a 30 s-git-synced website (R-110), so no build-time value in the hub can be true. `hostinstall_gates.py` gate 1 inverted to pin the ABSENCE of a version literal; the tautological `render_test.go` assertion deleted (demonstrated passing at `9.9.9`). Detail: `OPEN-ITEMS.md` R-94 | -| R-110 | **`main` is the installer's publish channel — there is no staging** | S | idea — found 2026-07-29, **WAITING-ON-OPERATOR (a ruling, not a defect)** | `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a `--period=30s`, and nginx serves that working tree directly (`location /scripts/`, `root /usr/share/nginx/html/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`, `:1262`; `runbooks/day0-install.md` C.1) receives. **There is no tag, no pinned-version path, no staging copy and no rollback other than another push** — for the one artifact that runs as **root on a virgin box**, the most privileged thing Felhom ships. **Two consequences worth stating plainly:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29, so the proof run confirms what customers already receive rather than clearing it for release; and the precaution the old R-94 row recorded ("do not point every new box at an installer that has never run") **was never available to take**, because nothing points boxes at a version. **Open question for the operator, not a defect to fix blind:** should `/scripts/` serve a pinned release — a tag-tracked git-sync ref, or a versioned directory (`/scripts/1.22.0/…`) with the hub's generated command naming a version — or is `main`-tracking the accepted shape for a one-operator product where the alternative is a release ritual nobody performs? **Exposure today is zero** (no boxes are installing), which is exactly why it is cheap to decide now. Whichever way it goes, it decides whether R-94 leg (a) makes the label a *fact* (derived from the served script) or keeps it a *claim*. Flips no capability-map row — the map states what the platform does, and this changes nothing about that | +| R-110 | **`main` is the installer's publish channel — there is no staging** | S | **READY** — operator ruling 2026-08-03: **option (b), the channel moves to a TAG**, so publishing is moving the tag and rollback is moving it back. **Must cover BOTH channels** — the nginx `/scripts/` git-sync AND the nine files the installer fetches from `raw/branch/main` — or it only half-works. CC's to build | `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a `--period=30s`, and nginx serves that working tree directly (`location /scripts/`, `root /usr/share/nginx/html/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`, `:1262`; `runbooks/day0-install.md` C.1) receives. **There is no tag, no pinned-version path, no staging copy and no rollback other than another push** — for the one artifact that runs as **root on a virgin box**, the most privileged thing Felhom ships. **Two consequences worth stating plainly:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29, so the proof run confirms what customers already receive rather than clearing it for release; and the precaution the old R-94 row recorded ("do not point every new box at an installer that has never run") **was never available to take**, because nothing points boxes at a version. **Open question for the operator, not a defect to fix blind:** should `/scripts/` serve a pinned release — a tag-tracked git-sync ref, or a versioned directory (`/scripts/1.22.0/…`) with the hub's generated command naming a version — or is `main`-tracking the accepted shape for a one-operator product where the alternative is a release ritual nobody performs? **Exposure today is zero** (no boxes are installing), which is exactly why it is cheap to decide now. Whichever way it goes, it decides whether R-94 leg (a) makes the label a *fact* (derived from the served script) or keeps it a *claim*. Flips no capability-map row — the map states what the platform does, and this changes nothing about that | | R-128 | **`ISO_VERSION` "aligns with SCRIPT_VERSION" was a comment nothing evaluated** | XS | **CLOSED — iso v1.26.0, 2026-07-31** | Closed by **correcting the claim, not asserting it**: the ISO is frozen while `felhom-host-install.sh` is fetched at run time from `main` (R-94/R-110), so an assertion would invent a constraint. `build-felhom-iso.sh:45-52`. Full reasoning in `OPEN-ITEMS.md` | | R-154 | **`[first-boot]` is automated-install-only, and nothing in the tree said so** | XS | **CLOSED — iso v1.26.0, 2026-07-31** | A PVE property, measured with a same-image control (`audits/SPIKE-universal-iso-3-2026-07-31.md` §2); recorded at `scripts/iso/pkg/build-deb.sh:6-11`. Superseded in practice by the `.deb` delivery route | | R-155 | **`iso-repack.sh` refused any ISO without `auto-installer-mode.toml`** | XS | **CLOSED — iso v1.26.0, 2026-07-31** | **Narrowed, not deleted** — unchanged for `FELHOM_MENU=single` (`iso-repack.sh:121-128`), does not apply to `release` where the file's absence *is* gate G1. Do not remove it wholesale | -| R-156 | **An app's data is neither persisted nor backed up, and it reports healthy** — a template mounts a volume at a path the app never writes, so data sits in the container's writable layer: lost on redeploy, tarred nightly as an empty dir, healthcheck green | S | **DETECTED + GATE SHIPPED; papra REFERRED** (catalog sweep, 2026-08-02) | Found by Campaign 10 on **papra**; the 53-template sweep convicted **gramps-web** and **wishlist** too and **fixed both**. The gate is `app-catalog-felhom.eu/scripts/check-volume-persistence.py` — a runtime probe (`docker diff` + mount occupancy + writability, canary self-test, fails closed); the detector and the gate are one program. **papra is NOT fixed**: the fix needs the app to use `/app/data` or the template to mount `/app/app-data`, so it is referred. Sweep evidence: `app-catalog-felhom.eu/audits/persistence-sweep-2026-08-02/` (53 probe.json). Ranking in `OPEN-ITEMS.md` | +| R-156 | ~~**An app's data is neither persisted nor backed up, and it reports healthy**~~ | S | **CLOSED — all three apps fixed, 2026-08-03.** papra's template now mounts `papra_data:/app/app-data` (the app's own data ROOT, chosen over reconfiguring three env vars so a future upstream path cannot escape). Deployed nowhere — re-checked on both demo guests and the hub fleet view, not inherited. Gate run in BOTH directions: CLEAN fixed, BROKEN reverted | Found by Campaign 10 on **papra**; the 53-template sweep convicted **gramps-web** and **wishlist** too and **fixed both**. The gate is `app-catalog-felhom.eu/scripts/check-volume-persistence.py` — a runtime probe (`docker diff` + mount occupancy + writability, canary self-test, fails closed); the detector and the gate are one program. **papra is NOT fixed**: the fix needs the app to use `/app/data` or the template to mount `/app/app-data`, so it is referred. Sweep evidence: `app-catalog-felhom.eu/audits/persistence-sweep-2026-08-02/` (53 probe.json). Ranking in `OPEN-ITEMS.md` | | R-157 | **`bootrecon`'s start-ONCE sweep misses the boot orphan it exists to recover** | S | **CLOSED — SHIPPED + PROVEN-LIVE (v0.189.0 + v0.190.0, 2026-08-02)** | Both mechanisms closed. **B:** the container count → recorded intent (R-166). **A:** the single T+5 s observation → a settle-then-sweep window (sample every 5 s, settled after 3 identical samples, one sweep at the end, ends on settled OR a 50 s budget). Budget is 50 s because settle+budget+one retry must fit the 90 s `deadAppBootGrace`; a test rejected 60 s at 95 s. Late recoveries are REPORTED, not hidden by widening the grace. **The fix had its own defect, found live:** sampling the Manager's 10 s-refreshed cache let "settled" mean "the cache did not update" — `sampleBootFleet` now refreshes first. **6/6 hard resets clean on the shipped build**, plus a same-app before/after (missed 18:08:35, recovered 18:18:50) | | R-170 | **The drive-backed boot gate infers a Stop from a container count** | S | **CLOSED — SHIPPED + PROVEN-LIVE (v0.190.0, 2026-08-02)** | `shouldRecreateOnBoot` reads `desired_state` with the same three-way table as `isBootOrphan`; absent keeps the old `hasContainers` behaviour exactly; `presentStable` untouched. Agreement pinned from both sides against one fixture table. Live: calibre-web (`running`, zero containers) recreated and immich (`stopped`) left alone in the same reboot | | R-171 | **The boot sweep started apps whose data drive was ABSENT** — a regression from v0.189.0 | S | **CLOSED — SHIPPED + PROVEN-LIVE (v0.190.0, 2026-08-02)** | Reasoned from the diff, CONFIRMED on hardware first (`audits/DIAG-bootrecon-drive-absent-2026-08-02.md`). The write hazard was blocked only by an ACCIDENTAL filesystem permission (host-root-owned mountpoint + unprivileged guest); the false dead-app alarm was real on every box. New fail-safe `bootrecon.StartGate` seam, also covering quiesce and in-flight app-data operations (§8.2). The rule already existed on the API path (`startGatedByMissingDrive`); the sweep bypassed it | | R-166 | **App state gets a desired/observed model with its own store** (operator decision D-b) | M | **SHIPPED + PROVEN-LIVE — controller v0.189.0, 2026-08-02** | Tri-state `desired_state` in `app.yaml`, written ONLY by the customer's action; `bootrecon` reads intent instead of `len(Containers) > 0`; **absent means UNKNOWN** so legacy boxes keep byte-identical behaviour; running-only backfill; `backup.AppStopGuard` covers every stop→work→start window in its own marker file. **The durable fix for R-157 mechanism B.** Both blocking facts answered at source first: the crash-safe journal pattern DID already exist (quiesce, migrate) but covered **none** of the app-data path, and the SQLite store was rejected as a home because `metrics.db` is optional by design. `07`/`02` architecture docs carry the desired/in-flight/observed split. **Left for R-170:** the drive-backed boot gate still uses the container-count inference | | R-158 | ~~A local Tier-1 backup failure reaches no hub channel~~ | S | **SHIPPED — controller v0.191.0 + hub v0.89.0, 2026-08-02** | Collapsed per the lifecycle rule. Closed by **R-167** (D-c's operator half) — no second row was filed for the same wire. New `unitNotify` seam fired per app from `captureAllRecoveryUnits` with the loop continuing, carrying the target filesystem's used/free bytes. **Routed to a new OPERATOR-ONLY type `recovery_unit_capture_failed`, NOT the `backup_failed` this row proposed** — that type is customer-enabled by default and would email the customer about a failure they cannot act on; decision D-c overrides the proposal. **Flips:** the new capability-map row *"A failed per-app Tier-1 backup reaches the OPERATOR"* → **PROVEN-LIVE** (`customer | … | skipped | operator_only` observed in the hub's `notification_log`, 9201, 2026-08-02) | | R-167 | ~~Storage monitoring and backup alerts (decision D-c)~~ | M | **SHIPPED — controller v0.191.0/.1/.2 + hub v0.89.0, 2026-08-02** | Collapsed per the lifecycle rule. **Landed BEFORE the R-165 merge rather than with it** — D-a's condition (2) requires the same step and never after, and first is strictly better: the warnings were proven on hardware while the wall is still standing. Customer half = `internal/fillwatch`, per filesystem, two threshold terms (85% / 5 GiB; critical 95% / 2 GiB), edge-triggered with persisted state and a 75% / 7 GiB hysteresis return. **It emits the pre-existing `disk_warning`/`disk_critical` pair, which had NO producer in any repo — the sixth *built-but-never-wired* instance**; the hub's two generic `customerMessages` entries were deleted so the dynamic Hungarian survives. Operator half = R-158. **Flips:** two new capability-map rows → **PROVEN-LIVE**. **Follow-ups:** R-177 (no run-now path), R-175 (§7.5's bound is one box's) | -| R-165 | Merge `mp1` into `mp0` — the dedicated backup partition stops existing (decision D-a) | M | **SPIKED 2026-08-02 — WAITING-ON-OPERATOR** | `audits/SPIKE-r165-mp1-merge-2026-08-02.md`, M1-M5, **no layout touched**. Prerequisite R-167 is now SHIPPED. **Three findings the merge session must not re-derive:** (1) *"the layout" is not one thing* — demo-felhom `200G/50G`, demo-hp `50G/20G`, golden `16G/8G`, so any fixed pair is already wrong for one box (→ R-175); (2) **`mp1` is a BULKHEAD, not only a ceiling** — an overflow today cannot reach `/var/lib/docker`, and after the merge it can, which is the one place "the merge is cheap" stops being true; (3) the golden fails closed on the split in **four** places, not one. **M5: D-a's condition (1) is currently SATISFIED** — no external box is in the hub's host register, and both demo boxes are Tier 0 and reinstallable, so migration cost for the measurable population is zero. **Recommendation: S1 (one volume, two directories) + B2 (a refusal threshold in the capture path), fresh-install shape.** **Two prerequisites unmeasured → R-176.** Does NOT close R-163 until it lands | +| R-165 | ~~Merge `mp1` into `mp0` — the dedicated backup partition stops existing (decision D-a)~~ | M | **CLOSED — PROVEN-LIVE 2026-08-03.** Variant V-c shipped (golden v3.0.0 + agent v0.120.0 + controller v0.192.0); both demo boxes reinstalled and proven end to end (R-178). **Its stated replacement for the bulkhead — B2 — was completed by R-181 the same day** (controller v0.193.0/.1): the reserve became a per-app, per-run admission covering all three write legs and gained a size term, and both terms were re-proven by filling demo-hp. R-181 is collapsed into this row rather than carried separately | `audits/SPIKE-r165-mp1-merge-2026-08-02.md`, M1-M5, **no layout touched**. Prerequisite R-167 is now SHIPPED. **Three findings the merge session must not re-derive:** (1) *"the layout" is not one thing* — demo-felhom `200G/50G`, demo-hp `50G/20G`, golden `16G/8G`, so any fixed pair is already wrong for one box (→ R-175); (2) **`mp1` is a BULKHEAD, not only a ceiling** — an overflow today cannot reach `/var/lib/docker`, and after the merge it can, which is the one place "the merge is cheap" stops being true; (3) the golden fails closed on the split in **four** places, not one. **M5: D-a's condition (1) is currently SATISFIED** — no external box is in the hub's host register, and both demo boxes are Tier 0 and reinstallable, so migration cost for the measurable population is zero. **Recommendation: S1 (one volume, two directories) + B2 (a refusal threshold in the capture path), fresh-install shape.** **Two prerequisites unmeasured → R-176.** Does NOT close R-163 until it lands | | R-159 | **wishlist's data landed in an ANONYMOUS volume — never backed up, orphaned by a redeploy** | XS | **SHIPPED** (`templates/wishlist/`, 2026-08-02) — filed for the CLASS | Image declares `VOLUME /usr/src/app/data`; template mounted `wishlist_data:/data`, a path the app never writes. `ResolveDockerVolumeNames` returns names only for compose-declared volumes, so `DumpAppVolumes` never sees an anonymous one. **The class is open:** any image `VOLUME` at an unmounted path is silent unbacked-up storage — **`immich-server` has one today** at `/data`, empty when measured. Proposed `REUSE.md` rule: a template mounts every path in `Config.Volumes`, or says why not | | R-160 | **gramps-web persisted three paths and wrote to none of them** | XS | **SHIPPED** (`templates/gramps-web/`, 2026-08-02) | `/app/data` appears nowhere in the image's env; the accounts DB and **the family tree** (`GRAMPS_DATABASE_PATH=/root/.gramps/grampsdb`) both landed in the writable layer. Upstream persists **eight** paths, the template three, one a phantom. **Severity above papra's:** papra loses documents the customer may hold elsewhere; gramps-web loses the artefact built inside the app, of which no other copy exists by construction | | R-161 | **The volume-persistence gate is enforced by convention, not automatically** | M | **RULED + SHIPPED at reduced scope** (operator, 2026-08-02; `app-catalog` `fd7747d`) | Enforcement is **convention**: the catalog repo has no CI at all (`.gitea/workflows`, `.github`, drone/woodpecker — searched, none). **R-29's record: three orphaned gates, one enforced, and only the enforced one ever stopped anything — so the ruling copied the shape that works.** **Controller-side REJECTED with a measured reason:** a check at template load can only read the file, and a static audit of all 53 templates reports the catalog clean **including papra** — it would pass on the very defect it exists to catch (the property is runtime-only; see `check-volume-persistence.py`'s header). **CI REJECTED for now** — neither repo has any, no users yet. **Shipped:** `scripts/catalog_gates.py`, one entry point over all three gates, non-zero on any failure, mandated in the catalog's `CLAUDE.md` as `site_gates.py` is. **Open residue is only the automatic half**, sufficient while one person touches templates | From 8360f940bfb21cd3c7db2f2fde4a188ffb503cb8 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 11:36:47 +0200 Subject: [PATCH 09/13] docs: seventh row in the shipped-guarantees table (R-181), versioned workspace CLAUDE.md MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Syncs documentation/runbooks/workspace-CLAUDE.md with the workspace root file. Row 7: the B2 refusal claimed the previous unit was untouched; nothing-deleted held, untouched was measured false. Also records WHY it survived review: it passed a full green suite AND three of its own red-proofs, because every one of them asserted the mechanism inside captureAllRecoveryUnits and none asserted the consequence across the whole backup run. The test that would have caught it is the one the fix ships — fingerprint the tree before and after, and compare. --- documentation/runbooks/workspace-CLAUDE.md | 12 ++++++++---- 1 file changed, 8 insertions(+), 4 deletions(-) diff --git a/documentation/runbooks/workspace-CLAUDE.md b/documentation/runbooks/workspace-CLAUDE.md index 76ea98c2..c8b46d3a 100644 --- a/documentation/runbooks/workspace-CLAUDE.md +++ b/documentation/runbooks/workspace-CLAUDE.md @@ -178,7 +178,7 @@ turns a true alarm into one the operator dismisses. ### A comment asserting an invariant needs a test pinning it, or it is a wish -**Six instances in this project have shipped guarantees the code did not provide** — each survived +**Seven instances in this project have shipped guarantees the code did not provide** — each survived review because the comment read as settled: | # | Comment | What it claimed | What the code did | @@ -189,10 +189,14 @@ review because the comment read as settled: | 4 | `classifyRunStates` I1 | *"StateStopped means deliberately stopped by the user"* | quiesce stops stacks the same way — a failed restart was silent (F-CRIT-1) | | 5 | `inflight.go` | *"a caller that cannot acquire DEFERS"* | the backup caller recorded a failure and paged the operator (F-A1) | | 6 | `quiesce.go` | the agent's 409 *prevents* "a spurious failure" | on the start path it produced one (F-A1) | +| 7 | `recovery_unit.go` B2 refusal (R-181) | *"the previous unit is untouched and NOTHING was deleted"* | *nothing deleted* held; **untouched was measured false** — the floor was checked ONLY in `captureAllRecoveryUnits`, while the two dump legs wrote the bulk into the same tree first and unguarded, so a 182,272 B tar became 2,147,666,432 B under a manifest that had not moved | -Two of these (4 and 5/6) were found by Campaign 8 **on live hardware**, not by review or unit tests -— #4 had a green, red-proofed test suite over a production path that was broken two independent -ways. So: +Three of these (4, 5/6 and 7) were found **on live hardware**, not by review or unit tests — #4 had a +green, red-proofed test suite over a production path that was broken two independent ways, and #7 +survived a full green suite plus three of its own red-proofs, because every one of them asserted the +mechanism inside `captureAllRecoveryUnits` and none asserted the **consequence** across the whole +backup run. The test that would have caught it is the one #7's fix ships: fingerprint the tree before +and after, and compare. So: - If a comment states an invariant, **name the test that pins it**, or write one. - If an invariant has a stated dependency (*"if either invariant changes, revisit this"*), that is From bee68484588900cb638a597f74357a10c18f081a Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 12:08:37 +0200 Subject: [PATCH 10/13] =?UTF-8?q?installer=20v1.23.0=20=E2=80=94=20publish?= =?UTF-8?q?ing=20becomes=20an=20act,=20not=20a=20side-effect=20(R-110,=20R?= =?UTF-8?q?-183)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two channels moved off main in the same change, because either one left behind makes the other cosmetic. Channel 1 — the served script. webpage.yaml git-synced /scripts/ from --branch=main every 30s and nginx served that tree, so pushing this file WAS publishing it: within half a minute it was what every new machine downloaded and ran as root, with no staging and no rollback but another push. The sync is now SPLIT: the website keeps tracking main at the same cadence (a copy edit must never need a release) and /scripts/ tracks the tag installer-v. PROVEN before the manifest was touched: git-sync v4.4.0 follows a tag AND notices a MOVED one — measured on a throwaway sync against this repo, "update required ... local: remote:" -> "updated successfully", within one period. The moved-tag half is what the publish model rests on. Channel 2 — the sixteen files fetched at run time. fetch_raw pulled from $AGENT_REPO/raw/branch/main; it now pulls raw/tag/v$ART_AGENT_VER. That is a correctness fix, not only a channel one (R-183): a fresh install fetched the vouched agent BINARY while taking its unit file, sudoers and guarded wrappers from whatever main held. Two refs, one install, nothing compared them. Their correct ref was never SCRIPT_VERSION — they do not live in this repo. No fallback to a branch: a vouched version whose tag is missing fails loudly rather than quietly serving main. Channel 3 — the URL — needed no change, recorded rather than left silent: https://felhom.eu/scripts/felhom-host-install.sh never carried a ref, so both producers follow the tag with no edit. No hub change, no hub version bump. Gate 6 in hostinstall_gates.py pins all three structurally with no network, so it stays in --fast and runs in CI. It deliberately does NOT assert "a tag exists for the current SCRIPT_VERSION": that would go red on the very push that bumps the version, before publishing — and publishing being separate is the ruling. --- manifests/webpage.yaml | 86 ++++++++++++++++++++++++++++++++-- scripts/CHANGELOG.md | 48 +++++++++++++++++++ scripts/felhom-host-install.sh | 25 ++++++++-- scripts/hostinstall_gates.py | 56 ++++++++++++++++++++++ 4 files changed, 209 insertions(+), 6 deletions(-) diff --git a/manifests/webpage.yaml b/manifests/webpage.yaml index f995f9d2..af28fbf9 100644 --- a/manifests/webpage.yaml +++ b/manifests/webpage.yaml @@ -71,8 +71,12 @@ data: # Host-install script. It lives at the repo's /scripts (outside the website doc-root), # synced into .../current/scripts by git-sync (see the sparse-checkout ConfigMap). Served # as text/plain so operators can inspect it in a browser before download-then-run. + # R-110: served from the INSTALLER TAG's tree, not the website's. The URL is unchanged + # (https://felhom.eu/scripts/felhom-host-install.sh) — it never carried a ref, so every + # producer of it (the bootstrap script, the hub's day-0 command) follows the tag with no + # edit. What changed is which tree this root points at. location /scripts/ { - root /usr/share/nginx/html/current; + root /usr/share/nginx/scripts/current; default_type text/plain; } @@ -213,8 +217,12 @@ metadata: name: git-sync-sparse-checkout namespace: felhom-system data: + # R-110: TWO sparse-checkouts, because there are now two syncs with two different refs. + # The website tracks `main` (a copy edit must never need a release); /scripts/ tracks the + # installer TAG (pushing the installer must never publish it). sparse-checkout: | /website/ + sparse-checkout-scripts: | /scripts/ --- # =================== @@ -246,6 +254,9 @@ spec: - name: git-data mountPath: /usr/share/nginx/html readOnly: true + - name: git-data-scripts + mountPath: /usr/share/nginx/scripts + readOnly: true - name: nginx-config mountPath: /etc/nginx/conf.d/default.conf subPath: default.conf @@ -269,11 +280,14 @@ spec: initialDelaySeconds: 3 periodSeconds: 10 + # ── The WEBSITE sync — tracks `main`, unchanged cadence ────────────────────────────── + # Deliberately still a branch: the site is content, and a typo fix must reach felhom.eu in + # thirty seconds without cutting a release. Only /scripts/ moved to a tag (R-110). - name: git-sync image: registry.k8s.io/git-sync/git-sync:v4.4.0 args: - --repo=https://gitea.dooplex.hu/admin/felhom.eu.git - - --branch=main + - --ref=main - --root=/git - --link=current - --period=30s @@ -294,13 +308,52 @@ spec: securityContext: runAsUser: 65534 # nobody + # ── The INSTALLER sync — tracks a TAG (R-110, operator ruling 2026-08-03) ───────────── + # felhom-host-install.sh runs as root on a virgin machine. Before this it was served + # straight from `main`, so pushing it WAS publishing it: within thirty seconds it was what + # every new machine downloaded and ran, with no staging and no rollback but another push. + # + # Publishing is now moving this tag; rolling back is moving it back. PROVEN, not assumed: + # git-sync v4.4.0 follows a tag AND notices a moved one — measured 2026-08-03 on a + # throwaway sync against this very repo (`update required … local: remote:` → + # `updated successfully`, one period, ~20 s). + # + # Bump this ref when the installer's published version changes. `hostinstall_gates.py` + # gate 6 fails if this sync stops naming an `installer-v…` tag. + - name: git-sync-scripts + image: registry.k8s.io/git-sync/git-sync:v4.4.0 + args: + - --repo=https://gitea.dooplex.hu/admin/felhom.eu.git + - --ref=installer-v1.23.0 + - --root=/git-scripts + - --link=current + - --period=30s + - --sparse-checkout-file=/etc/git-sync-scripts/sparse-checkout + volumeMounts: + - name: git-data-scripts + mountPath: /git-scripts + - name: sparse-checkout-scripts + mountPath: /etc/git-sync-scripts + resources: + requests: + memory: "32Mi" + cpu: "10m" + limits: + memory: "128Mi" + cpu: "100m" + securityContext: + runAsUser: 65534 # nobody + # Init container: wait for first sync before nginx starts initContainers: + # BOTH trees are seeded before nginx accepts traffic. The second one is why /scripts/ has + # no 404 window across this change: a fresh pod does not become ready until the installer + # tag has been checked out, exactly as the website already worked. - name: git-sync-init image: registry.k8s.io/git-sync/git-sync:v4.4.0 args: - --repo=https://gitea.dooplex.hu/admin/felhom.eu.git - - --branch=main + - --ref=main - --root=/git - --link=current - --one-time @@ -312,16 +365,43 @@ spec: mountPath: /etc/git-sync securityContext: runAsUser: 65534 + - name: git-sync-scripts-init + image: registry.k8s.io/git-sync/git-sync:v4.4.0 + args: + - --repo=https://gitea.dooplex.hu/admin/felhom.eu.git + - --ref=installer-v1.23.0 + - --root=/git-scripts + - --link=current + - --one-time + - --sparse-checkout-file=/etc/git-sync-scripts/sparse-checkout + volumeMounts: + - name: git-data-scripts + mountPath: /git-scripts + - name: sparse-checkout-scripts + mountPath: /etc/git-sync-scripts + securityContext: + runAsUser: 65534 volumes: - name: git-data emptyDir: {} + - name: git-data-scripts + emptyDir: {} - name: nginx-config configMap: name: nginx-config - name: sparse-checkout configMap: name: git-sync-sparse-checkout + items: + - key: sparse-checkout + path: sparse-checkout + - name: sparse-checkout-scripts + configMap: + name: git-sync-sparse-checkout + items: + - key: sparse-checkout-scripts + path: sparse-checkout --- apiVersion: v1 kind: Service diff --git a/scripts/CHANGELOG.md b/scripts/CHANGELOG.md index 059e92fa..1811ac2e 100644 --- a/scripts/CHANGELOG.md +++ b/scripts/CHANGELOG.md @@ -1,3 +1,51 @@ +## v1.23.0 — the installer is published, not pushed (2026-08-03, R-110 + R-183) + +**Two channels moved off `main` in the same change, because either one left behind makes the other +cosmetic.** + +**Channel 1 — the served script.** `manifests/webpage.yaml` git-synced `/scripts/` from +`--branch=main` on a 30 s period and nginx served that working tree, so **pushing this file WAS +publishing it**: within half a minute it was what every new machine downloaded and ran as root, with +no staging and no rollback but another push. The sync is now **split in two**: the website keeps +tracking `main` at the same cadence (a copy edit must never need a release), and `/scripts/` tracks +the tag **`installer-v`**. Publishing is moving that tag; rolling back is moving it +back. + +**PROVEN, not assumed:** git-sync v4.4.0 follows a tag *and* notices a **moved** one — measured on a +throwaway sync against this repo, `update required … local: remote:` → `updated +successfully`, within one period (~20 s). The moved-tag half is what the whole publish model rests +on, so it was measured before the manifest was touched. + +**Channel 2 — the sixteen files the installer fetches while it runs.** `fetch_raw` pulled from +`$AGENT_REPO/raw/branch/main`. It now pulls from **`raw/tag/v$ART_AGENT_VER`** — the agent version the +hub has vouched and whose binary sha this script already verifies. + +**That is a correctness fix, not only a publish-channel one (→ R-183).** These are the AGENT's +configs — its systemd unit, its sudoers, its guarded wrappers — and a fresh install was fetching the +**vouched binary** while taking its configs from **whatever `main` held**. Two refs, one install, and +nothing compared them. The right ref for them was never this script's `SCRIPT_VERSION`: they do not +live in this repo and have no relationship to its version line. + +**No fallback to a branch.** A vouched version whose tag is missing fails loudly rather than quietly +serving `main` — a silent fallback is the appearance of control with none of it. `felhom-agent` +carries `v` tags from now on, `release-agent.sh` creates them, and `agent_gates.py` fails if +the vouched version is not downloadable. + +**Channel 3 — the URL — needed no change, and that is worth recording rather than leaving as a +silence.** `https://felhom.eu/scripts/felhom-host-install.sh` never carried a ref: the ref lives in +the manifest. So both producers of that URL (`scripts/iso/felhom-bootstrap.sh`, the hub's day-0 +command) follow the tag with no edit — **and no hub change, so no hub version bump.** + +**Gate 6 in `hostinstall_gates.py`** pins all three structurally, with no network so it stays in +`--fast` and runs in CI on every push: no `raw/branch/` ref anywhere in the installer; `fetch_raw` +still pins to `$ART_AGENT_VER`; the manifest still syncs `/scripts/` from an `installer-v…` tag and +the website still from `main`. + +**It deliberately does NOT assert "a tag exists for the current SCRIPT_VERSION".** That gate would go +red on the very push that bumps the version, before publishing — and publishing being a separate +deliberate act is the entire ruling. A gate that fails on the normal path is one people learn to +ignore. + ## docs — v1.22.0 exercised end to end on two real reinstalls (2026-08-03, R-178) — **no script change** **Nothing shipped.** `felhom-host-install.sh` stayed at **v1.22.0**; the published copy at diff --git a/scripts/felhom-host-install.sh b/scripts/felhom-host-install.sh index 3aed3872..df5193fa 100644 --- a/scripts/felhom-host-install.sh +++ b/scripts/felhom-host-install.sh @@ -184,7 +184,7 @@ set -euo pipefail -SCRIPT_VERSION="1.22.0" # the SINGLE version source (F-1): -h and the run banners follow it. +SCRIPT_VERSION="1.23.0" # the SINGLE version source (F-1): -h and the run banners follow it. # The hub used to carry a copy for its Setup tab; R-94 DELETED it # (2026-08-02) because the hub cannot know which version a box runs — # the Setup command fetches this script at run time. scripts/ @@ -492,12 +492,31 @@ fetch_verify() { # exists, else anonymous). These are non-executable text (not the integrity-checked binary); the # sudoers is `visudo -cf`-validated before install, which catches corruption/tampering that would # matter. $1=repo-path $2=dest +# +# R-110 / R-183: PINNED TO THE AGENT VERSION BEING INSTALLED, never to a branch. +# +# These sixteen files are the AGENT's configs — its systemd unit, its sudoers, its guarded wrappers — +# so the ref that is correct for them is the agent version this run is installing, which the hub has +# vouched and whose binary sha this script verifies. It is NOT the installer's own SCRIPT_VERSION: +# these files do not live in the installer's repo and have no relationship to its version line. +# +# Before this they came from `raw/branch/main`, which is a REAL SKEW and not only a publish-channel +# defect (R-183): a fresh install fetched the vouched agent BINARY while taking its unit file and +# sudoers from whatever `main` happened to hold — two refs, one install, and nothing compared them. +# +# NO FALLBACK TO A BRANCH. A vouched version whose tag is missing must fail loudly here rather than +# quietly serving `main`, because a silent fallback is exactly the "appearance of control with none of +# it" this change exists to remove. `agent_gates.py`'s published-version gate keeps the tag and the +# vouched version in step, so this die is a backstop and not the primary control. fetch_raw() { local path="$1" dest="$2" + # Late steps (mgmt-watchdog, OOB) can run without step 5 having resolved the manifest. + [[ -n "$ART_AGENT_VER" ]] || resolve_artifacts + [[ -n "$ART_AGENT_VER" ]] || die "cannot pin $path: no agent version resolved from the hub manifest" local -a _auth; _git_auth_args _auth curl -fsS "${_auth[@]}" -o "$dest" \ - "$GITEA_BASE/$GITEA_OWNER/$AGENT_REPO/raw/branch/main/$path" \ - || die "raw fetch failed: $path" + "$GITEA_BASE/$GITEA_OWNER/$AGENT_REPO/raw/tag/v$ART_AGENT_VER/$path" \ + || die "raw fetch failed: $path (agent tag v$ART_AGENT_VER — is that version tagged in $AGENT_REPO?)" [[ -s "$dest" ]] || die "raw fetch empty: $path" } diff --git a/scripts/hostinstall_gates.py b/scripts/hostinstall_gates.py index 00150dec..0f9e35d1 100644 --- a/scripts/hostinstall_gates.py +++ b/scripts/hostinstall_gates.py @@ -146,6 +146,62 @@ if re.search(r'^PVE_STORAGES=\([^)]*felhom-pbs[^)]*\)', src, re.M): else: fail("felhom-pbs missing from the default PVE_STORAGES — narrowing it 403s the PBS-DR apply-bridge") +# ── 6. the publish channel is pinned, not floating (R-110 / R-183) ────────────── +# +# WHAT THIS ASSERTS, AND WHAT IT DELIBERATELY DOES NOT. +# +# It does NOT assert "a tag exists for the current SCRIPT_VERSION". That gate would fail the very +# push that bumps SCRIPT_VERSION, before publishing has happened — and publishing being a SEPARATE +# deliberate act is the whole point of R-110's ruling. A gate that goes red on the normal path is a +# gate people learn to ignore, which is the reasoning the task's own §8.4 applies to the agent-side +# gate; it applies here identically. "Is the vouched version actually downloadable" is a real +# invariant and it lives where a missing artifact genuinely breaks day-0 — `felhom-agent`'s +# `agent_gates.py`, which has the network access to answer it. +# +# What it asserts instead are the two STRUCTURAL regressions that would silently return the +# installer to a floating channel, both answerable by reading files (no network, so this stays in +# `--fast` and therefore runs in CI on every push): +# +# 6a. no `raw/branch/` ref anywhere in the installer — one of the sixteen agent-config fetches +# slipping back to `main` is exactly how a channel stays floating unnoticed, and it is +# invisible in a diff that touches one line. +# 6b. `fetch_raw` still pins to the resolved agent version — the positive form, so the mechanism +# cannot be quietly deleted rather than regressed. +# 6c. the website manifest still syncs `/scripts/` from a TAG ref and the website from `main` — +# the split is the deploy-side half of the same channel, and reverting it is one word. +branch_refs = [l for l in lines if "raw/branch/" in l and not l.lstrip().startswith("#")] +if branch_refs: + fail("installer still fetches from a BRANCH ref — the run-time channel is floating again " + "(R-110/R-183). Offending line(s): %s" % "; ".join(l.strip()[:90] for l in branch_refs)) +else: + ok("no raw/branch/ ref in the installer — every run-time fetch is pinned") + +if re.search(r'raw/tag/v\$ART_AGENT_VER/', src): + ok("fetch_raw pins the agent configs to the vouched agent version") +else: + fail("fetch_raw no longer pins to $ART_AGENT_VER — the agent's configs and its binary can " + "again come from different refs in one install (R-183)") + +WEBPAGE = os.path.join(ROOT, "manifests", "webpage.yaml") +try: + with io.open(WEBPAGE, "r", encoding="utf-8") as f: + wp = f.read() +except IOError as e: + fail("cannot read manifests/webpage.yaml to check the publish channel: %s" % e) + wp = None +if wp is not None: + # The scripts sync must name a tag ref; the website sync must still track main. + if re.search(r'--ref=installer-v', wp): + ok("manifest: /scripts/ syncs from an installer tag") + else: + fail("manifests/webpage.yaml has no `--ref=installer-v…` sync — /scripts/ is not served " + "from a tag, so pushing the installer publishes it again (R-110)") + if re.search(r'--(branch|ref)=main', wp): + ok("manifest: the website still tracks main (a copy edit must not need a release)") + else: + fail("manifests/webpage.yaml no longer tracks main for the website — pinning the SITE to " + "the installer tag turns every copy edit into a release") + print() if fails: print("hostinstall gates: %d FAILURE(S)" % len(fails)) From 6a82719426b08d5e7fda73994b9ac2821c8efe56 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 12:24:22 +0200 Subject: [PATCH 11/13] test(R-110): Scenario A marker pushed to main WITHOUT moving the tag A comment only. It exists to be LOOKED FOR at the served URL: if it appears there, /scripts/ is still tracking main and the publish channel is still floating. It must not appear until installer-v1.23.0 is moved. --- scripts/felhom-host-install.sh | 2 ++ 1 file changed, 2 insertions(+) diff --git a/scripts/felhom-host-install.sh b/scripts/felhom-host-install.sh index df5193fa..ceacdec3 100644 --- a/scripts/felhom-host-install.sh +++ b/scripts/felhom-host-install.sh @@ -185,6 +185,8 @@ set -euo pipefail SCRIPT_VERSION="1.23.0" # the SINGLE version source (F-1): -h and the run banners follow it. +# R-110 Scenario A marker: this line was pushed to `main` WITHOUT moving installer-v1.23.0. +# If it appears at https://felhom.eu/scripts/felhom-host-install.sh, the channel is still floating. # The hub used to carry a copy for its Setup tab; R-94 DELETED it # (2026-08-02) because the hub cannot know which version a box runs — # the Setup command fetches this script at run time. scripts/ From e79a20bbedc35c0f4022849c4fa44c29552cb022 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 12:27:39 +0200 Subject: [PATCH 12/13] =?UTF-8?q?test(R-110):=20remove=20the=20Scenario=20?= =?UTF-8?q?A=20marker=20=E2=80=94=20proof=20complete?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Scenario A: pushed to main without moving the tag; the served bytes stayed byte-identical (sha 2f859555…) and the marker never appeared, while the website tree advanced to the new commit — both halves of the split in one observation. Scenario B: moving the tag published it (~40s, sha ea2b4aa9…, marker present); moving it back rolled it back to exactly the pre-publish sha with the marker gone. The website returned 200 throughout. --- scripts/felhom-host-install.sh | 2 -- 1 file changed, 2 deletions(-) diff --git a/scripts/felhom-host-install.sh b/scripts/felhom-host-install.sh index ceacdec3..df5193fa 100644 --- a/scripts/felhom-host-install.sh +++ b/scripts/felhom-host-install.sh @@ -185,8 +185,6 @@ set -euo pipefail SCRIPT_VERSION="1.23.0" # the SINGLE version source (F-1): -h and the run banners follow it. -# R-110 Scenario A marker: this line was pushed to `main` WITHOUT moving installer-v1.23.0. -# If it appears at https://felhom.eu/scripts/felhom-host-install.sh, the channel is still floating. # The hub used to carry a copy for its Setup tab; R-94 DELETED it # (2026-08-02) because the hub cannot know which version a box runs — # the Setup command fetches this script at run time. scripts/ From b0b269b28d2ea3c486faae5e2ccb70d17e822528 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Mon, 3 Aug 2026 12:44:08 +0200 Subject: [PATCH 13/13] docs: R-110 + R-115 closed, R-182 re-scoped by measurement, R-183/R-184 filed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit R-182's direction REVERSED by Part 0's measurement. Filed yesterday as "the reserve re-alerts on every status refresh" — too many alerts, seen at the sending end. Measured at the receiving end: 9 events received today, 2 operator emails sent. When two apps are refused in the same second the operator is told about ONE; the other is dropped before LogNotification, so it leaves no row on any channel and cannot be audited. The operator cooldown key is customerID:eventType(+tier) and the capture-failed event carries `app` but no `tier`, so the key has no app identifier. Same failure mode as R-97a, in a second event type that never opted into the narrow fix. Nothing changed — Part 0 was investigation only. Correction owed: yesterday's report said "one recovery_unit_capture_failed per app, HTTP 200". True of what the CONTROLLER pushed; a reader would take it as "the operator was told about each app", which is false. R-110 CLOSED (installer v1.23.0). Both channels moved. The spec's mechanism for channel 2 rested on a factual error — the run-time fetches are sixteen, not nine, and come from felhom-agent, not this repo — so no tag here could cover them; pinned to the agent version being installed instead, on the operator's ruling. Channel 3 needed no change: the URL never carried a ref, so no hub change and no hub bump. R-115 CLOSED. release-agent.sh builds, tags, publishes and verifies by an independent download; check-published-versions.py refuses a tag with no package; CI now runs the full gate set so it actually runs. R-183 NEW+CLOSED: a fresh install fetched the vouched agent binary and its sixteen config files from two different refs, and nothing compared them. R-184 NEW: nothing stops the hub vouching a version that was never released. The R-115 gate cannot see it — measured, the hub manifest and Gitea's package listing are both 401 anonymously. capability map: new PROVEN-LIVE row for the published installer channel. STATUS.md 138 -> 127 lines. --- CLAUDE.md | 16 + CONTEXT.md | 26 ++ REPORT.md | 358 +++++++++--------- STATUS.md | 90 +++-- .../architecture/00-capability-map.md | 1 + documentation/backlog/OPEN-ITEMS.md | 8 +- documentation/backlog/ROADMAP.md | 4 +- 7 files changed, 274 insertions(+), 229 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index eb2022a1..dd4816f5 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -213,6 +213,22 @@ local and skippable, and only CI is neither. - **Website** auto-deploys via git-sync; just push to `main` (live in 1–2 min). Website changes go through `repo_gates.py` above (it runs `site_gates.py`); new pages go into that gate's `PAGES` list. Emergency edits: https://files.felhom.eu. All `website/` HTML is **UTF-8 with BOM** — preserve it. +- **THE INSTALLER DOES NOT (R-110, 2026-08-03).** `manifests/webpage.yaml` runs **two** git-syncs: + the website from `main` as above, and `/scripts/` from the tag **`installer-v`**. + Pushing `scripts/felhom-host-install.sh` therefore changes nothing that any machine downloads — + which it used to, within thirty seconds, for the one artifact that runs as **root on a virgin box**. + - **To publish:** cut `installer-v`, bump the `--ref` in `webpage.yaml` + (both the sidecar and the init container), commit, and sync. `hostinstall_gates.py` gate 6 + fails if the manifest stops naming an `installer-v…` tag or if the website stops tracking `main`. + - **To roll back:** move the tag back to the previous commit and wait ~30 s. **No ArgoCD sync and + no deploy** — git-sync picks up a moved tag on its next period, measured live on 2026-08-03 in + both directions. That is the emergency lever; fix forward with a new version afterwards. + - **Do NOT pin the website to the tag.** The sparse-checkout used to cover `/website/` and + `/scripts/` in one sync, and pinning that would turn every copy edit into a release. + - The **URL never carries a ref** (`https://felhom.eu/scripts/felhom-host-install.sh`), so + `felhom-bootstrap.sh` and the hub's day-0 command follow the tag with no edit — do not add one. + - The installer's own sixteen run-time fetches are pinned separately, to `raw/tag/v$ART_AGENT_VER` + in the **agent** repo (R-183) — they are the agent's configs, not this repo's. - **Manifests** are GitOps via the `felhom` app — commit to `main`, then deliberate sync. ## Key patterns diff --git a/CONTEXT.md b/CONTEXT.md index 6fd56751..79d432f6 100644 --- a/CONTEXT.md +++ b/CONTEXT.md @@ -97,6 +97,32 @@ not provide, and the fourth of those found on live hardware rather than by revie no run scope, so a refused app re-alerts on every status refresh (measured: a second identical alert pair 13 s after the run's). Pre-existing in v0.192.0; R-181 changed neither caller. +**S-15 — publishing is an act, not a side-effect of pushing (2026-08-03, R-110 + R-115 + R-183).** +Two rulings, one shape: something became live because someone pushed, not because anyone decided. + +- **The installer.** `/scripts/` now git-syncs the tag `installer-v`; the **website + keeps tracking `main`** in a second sync, because pinning both would make every copy edit a + release. Publish = cut the next tag + bump the manifest `--ref` + sync. **Roll back = move the tag + back**, which takes ~30 s and needs no ArgoCD sync at all — git-sync v4.4.0 follows a moved tag, + and that half was measured before the manifest was touched because the whole model rests on it. +- **The sixteen run-time fetches were NOT what the spec described** — sixteen, not nine, and from + `felhom-agent`, not this repo — so no tag here could cover them. They are pinned to + `raw/tag/v$ART_AGENT_VER` instead, which is strictly better: the agent's configs now come from the + same ref as the agent binary being installed. That closed a real skew (**R-183**), not just a + channel. +- **The URL needed no change**, and that is worth knowing rather than re-deriving: it never carried + a ref, so both producers follow the tag automatically — and no hub change means no hub bump. +- **The agent.** `scripts/release-agent.sh` is THE release path: build → tag → publish → **verify by + an independent download**. It does not vouch. `check-published-versions.py` refuses a `v` + tag with no downloadable package, and **CI now runs the full gate set** rather than `--fast`, + without which that gate would have been registered and never run. +- **The gate's invariant is not the one specified, and P-C is why:** the hub manifest and Gitea's + package listing are both **401** anonymously; the package download and the tags api are not. So CI + can ask *is this installable* but not *what is vouched*. The residue is **R-184**. +- **Neither gate asserts "the newest version is published."** That would go red on the very push + that bumps a version, before publishing — and a gate that fails on the normal path is one people + learn to ignore. + **S-11 — D-c's routing, and why R-158's own proposal was overruled (2026-08-02, R-167 SHIPPED).** Decision D-c splits two signals by AUDIENCE, and the split is the ruling: **a fill warning is the CUSTOMER's** (they can free space, delete files, add a drive) and **a per-app backup capture failure diff --git a/REPORT.md b/REPORT.md index 196ed85d..d6991575 100644 --- a/REPORT.md +++ b/REPORT.md @@ -1,216 +1,200 @@ -# REPORT — R-181 (the reserve guards the write that fills the disk) + R-156 (papra) + two operator rulings +# REPORT — publishing becomes an act, not a side-effect (R-110, R-115) + R-182 measured, R-183/R-184 filed -**Date:** 2026-08-03 · **Repos:** `felhom-controller` (v0.192.0 → **v0.193.1**), `app-catalog-felhom.eu`, `felhom.eu` (docs only — **no hub change, no hub version bump**) +**Date:** 2026-08-03 · **Repos:** `felhom.eu` (installer **v1.22.0 → v1.23.0**), `felhom-agent` (**no bump**) +**Nothing was built** — no image, no binary, no golden. **Hub stays v0.89.0.** -## 1. Baselines — re-read on arrival, all matched §1 +## 1. Baselines — re-read on arrival, both matched §1 -| Repo | `main` @ arrival | Version | Shipped | +| Repo | @ arrival | Version | Result | |---|---|---|---| -| `felhom-controller` | `4be6467b501b` | v0.192.0 | **v0.193.0 `fef07c3`** → **v0.193.1 `6c43bf6`** | -| `app-catalog-felhom.eu` | `7cb58ecdf8e7` | n/a | `122bbee` | -| `felhom.eu` | `6b5d64c1fa73` | hub v0.89.0 | docs only, **no bump** | +| `felhom.eu` | `8360f940bfb2` | hub v0.89.0, `SCRIPT_VERSION="1.22.0"`, **0 tags** (confirmed) | installer **v1.23.0**, first tag `installer-v1.23.0` | +| `felhom-agent` | `9dfd89cb947e` | v0.120.0 | **unchanged** — scripts and gates only | -All three clean (`git status --porcelain` empty, `HEAD == origin/main`) before every build. +## 2. Part 0 — the R-182 measurement, and it REVERSED the row -## 2. The fix +Filed yesterday as *"the reserve re-alerts on every status refresh"* — **too many** alerts, observed +at the sending end. Measured at the **receiving end**, it is the opposite. -**One admission verdict per app per run** (`controller/internal/backup/admission.go`), taken before -that app's **first** write and consulted by all three legs — DB dump, volume dump, unit capture. The -three write under one per-app root (`appbackup.RecoveryUnitPath`), which is what makes one verdict -able to cover them honestly. +Method: the hub's SQLite copied **with its `-wal`** (4 MB and newer than the db — copying `hub.db` +alone would have read stale data, the exact trap this project recorded before), freshness confirmed by +the newest `notification_log` row post-dating the session. -- **Lazy, not run-wide.** App A's dump can put app B under the reserve; a run-start verdict reads a - disk that no longer exists. **Never re-decided between an app's own legs** — that is the split being - closed. **Reset per run.** -- **Ahead of `DumpAppVolumesSafe`**, which stops the stack as its first act, so a refused app is never - bounced. **After** the volume-less check, which has no write to gate. -- **Exactly one operator alert per refused app per run.** Leg order unchanged. -- **Size term added:** *would this app's write cross the reserve?* — estimated from its previous - `.sql` + `.tar`. **No history → headroom-only**, or the first backup becomes the one that can never - happen; the alert says so when that applies. +**9 `recovery_unit_capture_failed` events received today → 2 operator emails sent.** -## 3. Files +| time | apps refused (events in) | operator emails out | +|---|---|---| +| 06:40:03 | privatebin, opengist | **opengist only** | +| 08:59:46/47 | opengist, privatebin | **privatebin only** | +| 08:59:59 | privatebin, opengist | **none** | +| 09:03:00 | opengist | **none** | +| 09:07:06 | privatebin, opengist | **none** | -| File | | +**Cause, confirmed at source:** the operator cooldown key is +`customerID + ":" + eventType + cooldownTierSuffix(details)` (`dispatcher.go:268`, 1 hour hardcoded). +`RecoveryUnitFailureDetails` carries **`app`** and **no `tier`**, so the suffix is empty and the key +holds **no app identifier**. The first refused app takes the slot; every other app's refusal for the +next hour is dropped — and dropped **before `LogNotification`**, so it leaves **no row on any +channel** and cannot be audited afterwards. + +This is **R-97a's failure mode in a second event type**; that row's own comment states it +(*"`felhom-pbs` failing at 09:00 would swallow `local` failing at 09:20"*). `cooldownTierSuffix` was +written narrow on purpose; `recovery_unit_capture_failed` simply never opted in. + +**A correction I owe on yesterday's report.** It said *"one `recovery_unit_capture_failed` per app, +HTTP 200"*. That was true of what the **controller pushed**, and a reader would take it as *the +operator was told about each app* — which is false. The gap between an accepted event and a sent +email is the whole of this row. + +**Nothing was changed** (§8.5). R-182 is re-scoped with the evidence and the fix shape. + +## 3. Probes + +| | Question | Method | Verdict | +|---|---|---|---| +| **P-A** | does git-sync v4.4.0 follow a tag, and notice a **moved** one? | throwaway `docker run` git-sync against this repo, tag moved under it | **PASS both halves** — `update required … local:fb65202 remote:8360f94` → `updated successfully`, one period (~20 s) | +| **P-B** | does Gitea serve `raw/tag//`? | one fetch on a throwaway tag | **PASS** — HTTP 200, byte-identical to `raw/branch/main` | +| **P-C** | can CI read the package registry? | anonymous fetches | **PARTIAL, and it changed the gate's design** — package **download** 200 (and **404** for a fake version, so it discriminates), **tags** api 200; package **listing** api **401**, hub artifact manifest **401** | + +**Publish model P-A implies:** publishing is **moving the tag**; rollback is **moving it back**, in +~30 s with no ArgoCD sync and no deploy. Probe teardown: container, sync tree and probe tag all gone +(`git ls-remote --tags` → 0 at the time). + +## 4. §8.2's three channels — enumerated + +| Channel | Before | After | | +|---|---|---|---| +| 1. the served script | `main`, 30 s | **`installer-v1.23.0`** | **MOVED** — `webpage.yaml` split into two syncs | +| 2. the run-time fetches | `raw/branch/main` | **`raw/tag/v$ART_AGENT_VER`** | **MOVED** — but see below | +| 3. the URL producers | `main` | unchanged | **NO CHANGE NEEDED** — and that is a finding, not an omission | + +**Channel 2 was not what the spec described, and the spec's mechanism for it was unimplementable.** +There are **sixteen** fetches, not nine, and they come from **`felhom-agent`**, not `felhom.eu` — so +no tag on this repo could ever have covered them, and §8.1's *"derive the tag from `SCRIPT_VERSION`"* +was impossible for them. Raised before building; operator ruled to pin them to **the agent version +being installed**, which the installer already resolves from the hub manifest and already sha-verifies. +That is strictly better than any installer-derived tag: binary and configs now come from one ref. + +**Channel 3 needed no change because the URL never carried a ref** — +`https://felhom.eu/scripts/felhom-host-install.sh` is path-based; the ref lives in the manifest. So +`felhom-bootstrap.sh` and the hub's day-0 command follow the tag automatically. **No hub template +change ⇒ no hub bump**, so §1's rule was never in tension and the STOP it anticipated never arose. + +## 5. The tag convention + +- **Shape:** `installer-v` in `felhom.eu` (prefixed so it cannot be read as a hub, + agent, controller or golden version); `v` in `felhom-agent` (that repo versions one thing). + **No new constant in the installer** — channel 2 derives its ref from `$ART_AGENT_VER` at run time, + and channel 1's ref lives only in the manifest. +- **Publish:** cut `installer-v`, bump the `--ref` in `webpage.yaml` (sidecar *and* + init container), commit, sync. +- **Roll back:** move the tag back to the previous commit — takes ~30 s, **no ArgoCD sync, no deploy**. + +## 6. Scenario A — proven by HTTP + +A real commit was pushed to `main` (a marker comment in the installer) **without moving the tag**, and +three sync periods were allowed to pass so "unchanged" means "had every chance to change": + +``` +website tree (main): .worktrees/6a82719… <- ADVANCED to the new commit +scripts tree (tag): .worktrees/bee6848… <- STAYED +sha256 before push: 2f859555382c4c69c18c48dccd8d8b132ffd49b4dbe4e03e5dbb192e8d883555 +sha256 after push: 2f859555382c4c69c18c48dccd8d8b132ffd49b4dbe4e03e5dbb192e8d883555 +marker present at the served URL? 0 +https://felhom.eu/ -> HTTP 200 +``` + +Both halves of the split in one observation: the site still tracks `main`, the installer does not. + +## 7. Scenario B — publish and rollback, both directions + +| act | result | |---|---| -| `controller/internal/backup/admission.go` | **new** — the gate, the memo, the estimator | -| `controller/internal/backup/admission_test.go` | **new** — 11 tests | -| `controller/internal/backup/backup.go` | run scope + gates in the DB and volume legs | -| `controller/internal/backup/recovery_unit.go` | `floorVerdict` size-aware; capture leg via `admitApp` | -| `controller/internal/backup/capture_floor_test.go` | 3 call sites updated for the new signature | -| `controller/README.md`, `REUSE.md`, `CHANGELOG.md` | | -| `app-catalog-felhom.eu/templates/papra/docker-compose.yml` | mount moved to `/app/app-data` | +| tag moved `bee6848 → 6a82719` | scripts tree moved in **~40 s**; served `sha256 ea2b4aa9…`; **marker present** | +| tag moved back `→ bee6848` | scripts tree back in **~40 s**; served `sha256 2f859555…` — **exactly** the pre-publish sha; **marker gone** | -## 4. Tests — 28 packages `ok`, `rc=0` (read separately from any commit) +`https://felhom.eu/` returned 200 throughout. The marker commit was then reverted, and the tag moved +to `main`'s head — a **byte no-op**, verified by the served sha not changing. -All 11 new tests pass, plus the pre-existing floor suite. Refusal assertions are **sha256 tree -fingerprints before and after**, never log lines — the defect being fixed *is* a log line the tree -contradicted. +## 8. Files, commits, tags -The DB leg cannot run without Docker (`DiscoverDatabases` shells out), so its gate is pinned by an -**AST walk** of `backup.go` asserting `admitApp` precedes `DumpOne`. `strings.Contains` is -insufficient: a commented-out call still contains the string. +**`felhom.eu`** — `bee6848` (installer + gate + manifest), `6a82719` (Scenario A marker), `e79a20b` +(marker removed), plus the docs commit below. +`scripts/felhom-host-install.sh` · `scripts/hostinstall_gates.py` · `scripts/CHANGELOG.md` · +`manifests/webpage.yaml` · `CLAUDE.md` · `CONTEXT.md` · `STATUS.md` · `REPORT.md` · +`documentation/backlog/{OPEN-ITEMS,ROADMAP}.md` · `documentation/architecture/00-capability-map.md` -### Red-proofs — each demonstrated failing, then restored +**`felhom-agent`** — `dd2d1fe` (release path + gate + CI), `0db7766` (REPORT). +`scripts/release-agent.sh` **(new)** · `scripts/check-published-versions.py` **(new)** · +`scripts/agent_gates.py` · `.gitea/workflows/gates.yml` · `CLAUDE.md` · `CHANGELOG.md` · `REPORT.md` + +**Tags created:** `felhom.eu/installer-v1.23.0` (the first tag this repo has ever had) and +`felhom-agent/v0.120.0` (retroactive, at `cd6e267` — the commit the published binary was built from; +`configs/` is byte-identical there and at `main`, so nothing depended on the choice). + +## 9. Tests and red-proofs + +| Check | Result | +|---|---| +| `felhom.eu` `repo_gates.py --fast` | all 5 gates OK | +| `felhom-agent` `go build ./... && go vet ./...` | OK | +| `felhom-agent` `go test ./...` | **29 packages ok, rc=0** (read separately from any commit) | +| `agent_gates.py --fast` | `published` correctly **SKIPPED** (hook must not fail on a network blip) | +| `agent_gates.py` (full) | both OK | + +**Red-proofs, each demonstrated failing then restored:** | # | Mutation | Result | |---|---|---| -| 1 | **Both** dump-leg `admitApp` gates removed (= exactly v0.192.0) | Scenario A **RED** — *"the VOLUME leg ran for a refused app"*; with the leg assertions temporarily made non-fatal, the **tree fingerprint changed** too. Also red: Scenario C, Scenario D, and the AST wiring test (which named the DB leg specifically) | -| 2 | The entire size term removed from `floorVerdict` (both its thresholds) | Scenario D **RED** — 0 alerts where 1 was required | -| 3a | The reserve removed entirely | Scenario F **PASSED — recorded honestly.** The specified mutation does not exercise the assertion: removing the reserve makes every app write, which overwrites and adds but **deletes nothing**, so a deletion-watching test correctly stays green | -| 3b | A prune injected into the refusal path | Scenario F **RED** — this is the mutation that proves the test watches deletion | -| 4 | Floor moved above the warning band (90% / 6 GiB) | `TestFloorSitsBelowTheCriticalWarningBand` **RED** | +| C | one of the sixteen fetches reverted to `raw/branch/main` | **RED** — gate 6a *and* 6b both fired | +| D | assertions 6a **and** 6b removed (every guard the test covers), same bad installer | **zero** mentions of the regression — the guards are what catch it | +| 6c | the manifest before the split | **RED** on its own, before I fixed it — the gate was demonstrated red by the real pre-change state | +| F | `v9.9.9` tagged and not published | **RED**, `binary NOT downloadable (HTTP 404 …)`, rc=1 | +| F′ | the gate **deregistered** from `agent_gates.py`, same bad state | **rc=0, "all agent gates OK"** — restored → `CONVICTED: published`, rc=1 | -Every mutation removed **every** guard its test covers (#1 removed both dump-leg gates, not one). +**Scenario F measured on real CI, not inferred.** Runs **69** and **70** are on the *same commit* +`0db7766`: **success** before `v9.9.9` existed, **failure** after pushing it. One variable. This also +retrospectively explains runs 67/68. **One deliberate CI failure email reached the operator — that was +this proof, not an incident.** I could not read CI's own step log: the jobs endpoint needs a Gitea API +token, and the only credential available (`~/.docker/config.json`) is a registry password that the API +rejects — so the controlled before/after replaced the log rather than an assumption standing in for it. -## 5. Live validation — demo-hp guest 9201 (Tier 0), the method that found the defect +## 10. No version bumps, nothing built -**Method:** endpoint-level — `POST /api/debug/backup/dbdump`, the exact endpoint the debug UI button -calls, which runs the production `RunDBDumps`. No browser on DooPlex. +`felhom-agent` **v0.120.0** unchanged (no Go code changed). Hub **v0.89.0** unchanged (no hub file +touched). The installer's `SCRIPT_VERSION` **did** go 1.22.0 → 1.23.0 — the installer is not in §12's +no-bump list, its behaviour changed materially, and the tag derives from it. No image, binary or +golden was built. -**The instrument was re-proven before use.** demo-hp's thin pool is 53.93 GiB, so a real fill of a -70 G volume would exhaust it and corrupt every guest. A 5 GiB `fallocate` step moved guest `df` -1.2G → 6.2G while thin-pool `data_percent` held **36.83 → 36.83** — zero blocks allocated. Re-checked -at every step of the fill. +## 11. Register -### Headroom term — 08:59:46, 906 MB free / 99% used - -| Observable | Result | +| ID | Outcome | |---|---| -| Tree fingerprint before | `TREE_SHA=111d1760c18d3440f700634ab325f8b8` (10 files; opengist's tar **182,272 B** — R-181's own "before" figure) | -| Tree fingerprint after | **`111d1760c18d3440f700634ab325f8b8` — identical** | -| Volume dumps written | **0** (baseline run at 08:58 wrote 2) | -| `Stopping for safe volume dump` | **absent** — and this is evidence, not an absence, because that line **is** present in the 08:58 baseline | -| Operator alerts | one `recovery_unit_capture_failed` per app, severity `error`, HTTP 200 | +| **R-110** | **CLOSED — SHIPPED** (installer v1.23.0), both-channels condition honoured, though not in the shape the ruling assumed | +| **R-115** | **CLOSED — SHIPPED** (`release-agent.sh` + `check-published-versions.py`, no bump) | +| **R-182** | **RE-SCOPED — the direction reversed** by Part 0's measurement; still open, now correctly described | +| **R-183** | **NEW, and CLOSED the same session** — binary and configs came from two different refs | +| **R-184** | **NEW, open** — nothing stops the hub vouching a version that was never released | -Free space restored → re-run at **09:01:33**: both apps captured normally. +**IDs established free:** `^| \*\*R-183\*\*` / `^| \*\*R-184\*\*` in `OPEN-ITEMS.md` → **0 rows** each; +all other hits are this session's own code and changelogs (forward references I wrote). `R-185` → 0 +hits anywhere and remains free. -### Size term — 09:03:00, proven separately +## 12. Observations — noticed, documented, NOT acted on -Reproducing the original sequence: a real 2 GiB file planted in opengist's volume, backed up so its -**previous** tar became **2,147,666,432 B** (the exact live figure), then the filesystem set to -**91% used / 2.9 GB free — both headroom terms deliberately clear**. +1. **The gate cannot see what is vouched** — filed as R-184 rather than papered over. Closing it needs + either a hub credential in CI (operator's call) or a check at vouch time in the hub (better: fails + closed where the mistake is made, needs no new credential). +2. **A suppressed operator alert leaves no row at all.** The cooldown returns before `LogNotification`, + so the hub's own records cannot distinguish "never happened" from "held back". Recorded inside + R-182 because it is what made that row take a day to get the right way round. +3. **`on: [push]` fires CI for tag pushes too.** Useful (it is how Scenario F was measured), but it + means a tag push runs the full gate set — worth knowing before anyone adds an expensive gate. +4. **`felhom.eu` CI still runs `--fast`.** Correct today, since all its gates are network-free; if a + network gate is ever added there, that workflow needs the same change the agent's just got. -- **opengist refused `(size)`** — *"this app's last backup was 2.0 GB and writing it again would cross the reserve"* -- **privatebin ADMITTED and dumped normally** — the term is per-app, not a global halt -- Tree unchanged; 1 volume dump instead of 2 +## 13. Teardown -### One honest correction to the "app not stopped" claim - -`StartedAt` on both apps *did* move, 26 s **after** the refusal. It was the **quiesce loop** for the -whole-guest PBS backup, which my fill had broken — not the app-data path. Its own backoff logic then -behaved correctly (*"deferring its next quiesce by 15m so the apps are not stopped again for a backup -that cannot succeed"*). The app-data claim rests on the **absence of the `Stopping … for safe volume -dump` line**, which is the line that appears when that leg bounces an app. - -## 6. The `du` measurement (§Part 1.3) — measured, then rejected - -**66 timed runs** on demo-hp guest 9201, `docker run --rm -v :/v alpine du -sb /v`: -**median ~355 ms per volume, range 341–404 ms** — on volumes holding **tens of KB**. The cost is -container start-up, not the walk, so it does not shrink for small apps and only grows for real ones. - -**Rejected**, on two grounds beyond the number: `docker run` needs the writable layer, so the -measurement mechanism can fail under exactly the disk pressure the reserve exists to handle; and the -previous-dump estimate measures the **artifact that will be written** rather than the live volume, -which is the truer predictor. The previous-dump estimate stands. - -## 7. The refusal message as shipped, and what it guarantees - -``` -[WARN] [backup] App backup REFUSED for opengist (headroom) — refused: backing up this app would -leave the filesystem below the reserve (reserve: 97% used or 1.0 GiB free; the filesystem is already -below it, before this app's estimated 178.0 KB write) — /mnt/sys_drive: 64.3/68.7 GB used (94%), -0.9 GB free; NO database dump, NO volume dump and NO recovery-unit capture was written for it, the -previous unit is untouched and NOTHING was deleted -``` - -**It guarantees, for that app in that run:** no DB dump, no volume dump and no capture were written; -every file under `backups/primary/` is byte-identical; the app was not stopped; nothing anywhere -was deleted; exactly one operator alert was sent. All five verified by fingerprint above. - -**The wording was not weakened to fit the behaviour** — the behaviour moved so the wording became -true. What was *added* is the bound term (`headroom` / `size`) and the estimate. - -**v0.193.1 — found by this very proof run.** The estimate was rendered fixed to two-decimal GiB, so -opengist's real **178 KB** printed as `estimated 0.00 GiB write`, which reads as *no estimate was -available* — the opposite of what happened. Shipped the same session because it is the same defect -class the whole task is about. Re-verified live after redeploy: `estimated 178.0 KB write`. - -## 8. papra (R-156, last leg) - -**Precondition checked, not inherited** — both boxes were wiped and rebuilt today, so the 2 August -evidence was re-measured: `docker ps -a` (**including stopped**) on **both** demo guests → no papra; -hub `/hosts` → exactly two enrolled hosts (`demo-felhom-8363b5`, `demo-hp-bb76ea`), **zero** papra. - -**Decided from the image, not the README:** `WORKDIR=/app`, `DATABASE_URL=file:./app-data/db/db.sqlite`, -`DOCUMENT_STORAGE_FILESYSTEM_ROOT=./app-data/documents`, `PAPRA_CONFIG_DIR=./app-data` — and -**`/app/data` does not exist in the image at all**. - -**Departure from the task's stated preference order, stated because it was deliberate.** Option (1) -(reconfigure the app to write to `/app/data`) *was* available — all three paths are env-settable. Not -taken: it enumerates data paths, so a fourth added upstream would silently escape to the writable -layer again — this defect re-armed and invisible. Mounting the app's own data **root** captures every -current and future path by construction. - -**Gate output — the arbiter, run in both directions:** - -- fixed → `papra CLEAN`, with the self-test passing on that run: *"prober flags the R-156 signature and clears a correct template — trustworthy"* -- reverted to `/app/data` (red-proof on the **real template**, not just the canary) → `BROKEN`: *"mount /app/data is NOT writable by the app's own uid=999"*, *"DATA in the writable layer at /app/app-data/db (db_signature=True, e.g. ['db.sqlite'])"*, *"declared volume /app/data is EMPTY"* -- `catalog_gates.py papra` (full, not `--fast`) → **rc=0**, all three gates OK - -**Two operational findings about the gate:** it needs **root** (it reads `/var/lib/docker/volumes`, -mode `drwx--x---`; as a normal user its own canary fails UNDETERMINED and it correctly refuses a -verdict — fail-closed working as designed), and it hardcodes scratch path `/srv/felhom-gate`, created -on DooPlex. Unscoped it deploys all 53 templates; that run was aborted after 10 minutes and its -`volgate-*` scratch projects were cleaned up. - -## 9. §3's correction — confirmed in passing, not chased - -`restore_points.go:57-59` takes the manifest's mtime and then `newestArtifact` over the `.sql` and -`.tar` files, so **the newest of the three wins**. The restore point does **not** show a stale -timestamp. Confirmed and dropped, as instructed. - -## 10. Register - -| ID | Change | -|---|---| -| **R-181** | **CLOSED — SHIPPED** (v0.193.0 + v0.193.1), with the live evidence above | -| **R-156** | **CLOSED** — all three apps fixed | -| **R-110** | WAITING-ON-OPERATOR → **READY**, ruling attached: **option (b), tag-tracked**, and it must cover **both** channels (the `/scripts/` git-sync *and* the nine files fetched from `raw/branch/main`) or it only half-works | -| **R-115** | WAITING-ON-OPERATOR → **READY**, ruling attached: **mechanism (b)**, a build-side gate refusing to deploy or vouch an unpublished version; the third instance (agent v0.120.0) would have silently downgraded both demo boxes while reporting success | -| **R-182** | **NEW.** ID established free: `grep -ro "R-182\b"` over `documentation/` and `*.md` → 2 hits, both prose in `REPORT.md` recording it as *"checked and left unused"*; `R-183` → 0 hits and remains free | - -**R-165** is collapsed to CLOSED/PROVEN-LIVE in `ROADMAP.md`; the capability map's local-backup row -moves to **PROVEN-LIVE, both halves**, because the live fill proved the fixed behaviour for **both** -reserve terms. - -## 11. Observations — noticed, documented, NOT acted on - -1. **R-182 (filed).** The periodic status refresh (`GetFullStatus` → `captureAllRecoveryUnits`) runs - with no admission scope, so a refused app re-alerts on every poll — measured live: a second - identical alert pair 13 s after the run's. **Pre-existing in v0.192.0**; R-181 changed neither - caller. Its mitigation is a *comment* claiming the hub owns cooldown — which is exactly the - "invariant asserted in a comment with no test pinning it" shape, so verify at the hub before - scoping. -2. **A reserve refusal does not make the run fail.** The DB and volume legs record `SKIP`, not `FAIL`, - so `lastDBDump.Success` stays true and the customer-facing status does not turn red. Deliberate and - consistent with v0.192.0 (the capture refusal never set it either), and the operator alert is the - signal — but it means "backup succeeded" and "every app was backed up" are not the same statement. -3. **`UnitSpace.UsedPercent` and `df` disagree** — `df` reported 99% where the alert said 94%, because - `df`'s figure accounts for ext4 reserved blocks and the floor's does not. Harmless here (the - free-byte term bound), but a percent-term threshold is being compared against a number the operator - cannot reproduce with `df`. -4. **The whole-guest PBS backup fails when the volume is near-full**, pushing - `whole_guest_backup_failed` (severity `error`). Expected under a deliberate fill, and its backoff - behaved correctly; noted because it is collateral any future fill test will also produce. - -## 12. Teardown - -Fill file removed; the planted 2 GiB file removed; a final backup regenerated a correct 178 KB tar; -`pct fstrim 9201` returned 67.5 GiB and the thin pool settled at **29.43%**, *below* its 36.83% -baseline. The backups tree is byte-identical to the pre-test fingerprint. Guest helper scripts and the -credential file `shred`-ed. `volgate-*` scratch compose projects removed; the unrelated 9-day-old -`jarr-*` containers on DooPlex were left untouched. papra is **not** left deployed. - -No `--no-verify` was used on any push; the `felhom-controller` pre-push hook ran and reported -`gates OK` on both pushes. +Probe container, probe sync tree and probe tag (`probe-r110-delete-me`) removed; the red-proof tag +`v9.9.9` deleted (`git ls-remote --tags` → only `v0.120.0`); the Scenario A marker reverted from +`main` and the installer confirmed byte-identical to the published tag; the throwaway in-cluster curl +pod removed; the hub DB copy is scratch-only and holds no secret material in any committed file. diff --git a/STATUS.md b/STATUS.md index 87883e1d..7da1922c 100644 --- a/STATUS.md +++ b/STATUS.md @@ -27,34 +27,51 @@ time, and an app switched off deliberately stayed off every time. also delete it. A daily snapshot is armed as a stopgap, and we have never restored from that copy. *(R-95, R-87)* -**A full disk emails you repeatedly instead of once.** When the reserve refuses an app's backup you -are told once by the backup run — correctly — but the page showing backup status re-checks on a timer -and sends the same message again each time. Not new: as old as the reserve itself, and seen only -because we watched the alerts closely while proving the fix below. Harmless if the hub already -collapses repeats — which a comment claims and nobody has checked. *(R-182)* +**A full disk tells you about ONE app and silently swallows the rest.** Yesterday this was written +down the wrong way round — as *too many* emails. Measuring the receiving end reversed it: of nine +refusals the machine reported today, **two emails were sent**. When two apps are refused in the same +second you are told about one of them, and the other leaves no trace anywhere — not an email, not +even a line in the log saying it was held back. So a second app can be going unbacked-up while you +have already been told the problem is handled. It is the same fault we fixed once before for +whole-machine backups, in a second place that never opted into the fix. *(R-182)* ## What shipped recently -**The backup partition is gone, and both demo machines run on the new shape.** Wiped and rebuilt on -3 August and taken through the whole customer journey — set up, install an app, back it up, restore -it. One storage area instead of two; the space a backup can use went from 19 GB to 65 GB on the small -machine and 45 GB to 233 GB on the big one. Three reboots each, correct every time. The two were -rebuilt deliberately differently — one from a local copy of the image, one by the ordinary customer -route with the published fingerprint checked — so the disk shape and the delivery route are both -proven, rather than one proven twice. Their previous demo apps and data are gone; that was the point -of a wipe, and you approved it. *(R-165, R-178)* +**Pushing the installer no longer publishes it.** The script that runs as root on a brand-new +machine was copied from the main branch and served within thirty seconds, so pushing it *was* +publishing it, with no staging and no way back but another push. It now comes from a **labelled** +version: publishing is moving the label, and undoing it is moving the label back — about half a +minute, no deploy. The website is untouched by this and still updates in thirty seconds, because a +typo fix must never need a release. Proven by actually doing it: a real push changed nothing that +anyone downloads, moving the label published it, moving it back restored the previous bytes exactly. -**What replaced the wall — and it now watches the right moment.** The wall was quietly doing a second -job: keeping a runaway backup from eating the space the machine needs to keep running. That job is now -explicit, and as first built it was checked too late — the big write happened first, unchecked, and -only the small write after it was refused, while the message still promised your last good copy was -untouched. **Fixed and proven on 3 August.** The machine now decides once, per app, **before it writes -anything at all**, and that one answer covers all three steps: a refused app writes nothing, is not -restarted, and the promise is now literally true — checked by fingerprinting every file before and -after. It also stopped being blind to size, so an app is no longer waved through at 96% full and then -allowed to write two gigabytes. Proven by deliberately filling a demo machine, once for each way it -can refuse. Nothing is ever deleted to make room: every app has only one local copy, so "delete the -oldest" would always mean destroying some other app's only copy. *(R-181)* +**The catch that would have made it cosmetic was found and covered.** While it runs, the installer +fetches sixteen more files — not nine, and from the *agent's* repository, not the website's. They now +come from the same version of the agent the machine is installing. That closed a real fault nobody +had noticed: a new machine was getting the agent's tested program and its untested settings files, in +one install, from two different places. *(R-110, R-183)* + +**Releasing the agent now publishes it, in one command.** Putting a built agent where a new machine +can download it was a step someone had to remember, and it was forgotten three times in five days — +the last time leaving both demo machines running a version nobody could download, so a rebuild would +have quietly installed the *older* one and reported success. There is now one command that builds, +labels, publishes and then **downloads it back to check** — and a check that refuses to stay quiet if +a released version cannot actually be fetched. Proven by making CI fail on purpose and then go green +again on the same code. *(R-115)* + +**The backup partition is gone and both demo machines run on the new shape** — wiped, rebuilt and +taken through the whole customer journey on 3 August, by two deliberately different routes so the disk +shape and the delivery route are both proven. The space a backup can use went from 19 GB to 65 GB on +the small machine and 45 GB to 233 GB on the big one. Their previous demo apps and data are gone; that +was the point of a wipe, and you approved it. *(R-165, R-178)* + +**What replaced the wall now watches the right moment.** The wall was quietly keeping a runaway +backup from eating the space the machine needs to run. As first built, that replacement was checked +too late — the big write happened first, unchecked — while still promising your last good copy was +untouched. Fixed and proven on 3 August: the machine decides once, per app, **before it writes +anything**, and that one answer covers all three steps, so a refused app writes nothing, is not +restarted, and the promise is now literally true. It also stopped being blind to size. Nothing is ever +deleted to make room. *(R-181)* **The last of the three apps that never saved their data is fixed.** Installed nowhere, so nothing was stranded — checked on both demo machines and in the fleet list rather than assumed. Proven by the check @@ -75,24 +92,27 @@ step — but it notices quickly and tells you. *(R-29, R-161, R-168, R-169)* ## What we're working on -- **Now:** both of today's items are done — the reserve and the last unsaved app. Your two decisions - are written down and are ours to build. -- **Next:** building those two — moving the installer onto a labelled version so publishing is one - step you can undo, and a check that refuses to install a version nobody can download *(R-110, R-115)*. +- **Now:** nothing outstanding from today — the reserve, the last unsaved app, and both of your + decisions are all built and proven. +- **Next:** the alert that tells you about one app and swallows the second *(R-182)*. - **After:** the off-site copy that the machine making it can still erase *(R-95, R-87)*. ## Waiting on you - **A job, not a decision: the hub password needs changing.** A diagnostic command printed it into a session log; nothing suggests anyone else saw it. *(R-132)* -- **Nothing else.** You settled both open questions on 3 August — the installer moves onto a labelled - version, and a check will refuse to install a version nobody can download. Both are written down and - are ours to build. *(R-110, R-115)* +- **Nothing else.** Both decisions you took on 3 August are now built and proven. One small question + will come back later: the automatic check cannot see which version you have told machines to + install, only which ones exist — closing that either needs a password given to the build server or + a check inside the hub itself. Filed, not urgent. *(R-184)* ## Changed since last update -- **2026-08-03** — The reserve now guards the step that fills the disk, and its promise is true; the - last app whose data was never saved is fixed. Both proven on a demo machine, not just in tests. +- **2026-08-03** — Publishing became something you do rather than something that happens: the + installer and the agent both moved onto labelled versions with a way back, and a check now refuses + a release nobody can download. Earlier the same day: the reserve now guards the step that fills the + disk and its promise is true, and the last app whose data was never saved is fixed. All proven on + real machines, not just in tests. Earlier the same day: both demo machines wiped and rebuilt from the new base image and taken through set-up → install an app → back it up → restore it, with the backup space ceiling gone and measured. @@ -102,10 +122,6 @@ step — but it notices quickly and tells you. *(R-29, R-161, R-168, R-169)* the hub's own database is in no automatic backup** — it holds every machine's emergency password. Filed, not yet fixed. -- **2026-08-02** — Boot recovery finished; six hard resets, everything back every time. Two instances - of the same hole — starting an app whose external drive was missing — were found by reading the code - and fixed the same day. - - **2026-08-02** — Thirteen mechanical checks had built up and nothing ran most of them; two were failing quietly. Fixed. Decided the same day: the 20 GB backup partition goes away; and only this machine and the tester's box are protected, every other box may be broken or reinstalled freely. diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 89fc4a9f..c0a9f7ee 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -31,6 +31,7 @@ |---|---|---|---|---| | Appliance day-0 install: golden image → first boot → auto-confirm (zero clicks) → claimable box | installer, agent, hub, golden | **PROVEN-LIVE** (nested VM) | `DRILL-day0-vm-2026-07-12`, `DRILL-day0-take2-2026-07-12` | First firing on real customer hardware pending → R-1 | | BYO install: `--mode byo`, mandatory caps, host-mutation disclosure, coexistence guards | installer v1.15+, agent | **PARTIAL** | `DRILL-GL6-2026-07-08` (demo box); GL-8 coexistence fixes | Peti clean-slate reinstall on proxmox2 is the first real BYO run of the current path → R-1 | +| **The installer is PUBLISHED, not pushed — the artifact that runs as root on a virgin box is served from a version-controlled ref, and rolling back is one act** | scripts **v1.23.0** + `manifests/webpage.yaml` (R-110, operator ruling option (b)) | **PROVEN-LIVE (2026-08-03)** | `scripts/CHANGELOG.md` v1.23.0 + `REPORT.md`. **Proven by HTTP against the real URL, not from a pod's filesystem.** *Scenario A:* a real push to `main` without moving the tag left the served script **byte-identical** (`sha256 2f859555…`), and a marker comment planted in that very commit was **absent** from the served bytes, while the website tree advanced to the new commit in the same observation — both halves of the split in one measurement. *Scenario B:* moving the tag published in **~40 s** (`sha → ea2b4aa9…`, marker present) and moving it back restored **exactly** the pre-publish sha. `https://felhom.eu/` returned 200 throughout. *P-A, measured BEFORE the manifest was touched because the model rests on it:* git-sync v4.4.0 follows a tag **and notices a moved one** (`update required … local: remote:` → `updated successfully`) | **Two syncs, deliberately: the WEBSITE still tracks `main`.** Pinning both would turn every copy edit into a release, which makes the release meaningless and the site slow to fix. **Publish** = cut `installer-v` + bump the manifest `--ref` + sync; **roll back** = move the tag back, which needs **no ArgoCD sync and no deploy**. Continuity is structural rather than lucky: both trees are seeded by init containers so a fresh pod is not Ready until the tag is checked out, and `maxUnavailable` rounds to 0 on one replica, so a failed scripts-init leaves the OLD pod serving — the failure direction is *no update*, never *no `/scripts/`*. **The URL never carried a ref**, so the bootstrap script and the hub's day-0 command follow the tag with no edit and **no hub change**. The installer's own sixteen run-time fetches are a separate channel pinned to the AGENT's version (**R-183**), because they are the agent's configs and not this repo's — leaving them on `main` would have made the whole change cosmetic | | Bare-metal Felhom ISO (blank hardware → zero-touch auto-install → first-boot `host-install`); selectable UEFI loader; **universal secret-free / operator-bind** mode | scripts v1.19.0 (`scripts/iso/`) + hub v0.62.0 + assistant container | **PROVEN-LIVE on TWO different boards** (N100 2026-07-18; HP t740 2026-07-21) | **`tests/VALIDATION-n100-rehearsal-2026-07-18.md` — the full chain on real metal in a single pass:** the generic reusable pairing ISO (v1.20.0, `--loader mkimage`, SB off) booted the cheap AMI board that F1 had blocked, installed unattended, and the box **self-registered as an unclaimed appliance at 16:17:14 — the same second it first booted** (`appliance_registrations` id=3), then bound → credential-delivered → day-0 SUCCESS 16:32:32 → floor-lifted to current. **F1 is closed on physical hardware.** Prior nested legs: slice A `SPIKE-baremetal-iso-2026-07-16` (build gate, disk-filter fail-safe, stub→host-install fetch); slice B RUNBOOK-B (shim boots+installs OVMF SB-enforcing + SeaBIOS; `--loader mkimage` boots+installs SB-off; mkimage SB-enforcing **FAILS** `Access Denied`; surgery byte-identical); **slice C (2026-07-17): the GENERIC secret-free ISO** — box self-registers as an unclaimed appliance (`POST /api/v1/appliance/register`, one-shot poll delivery, 404-no-oracle — all live-verified through the public ingress), operator binds on the Hosts page, hub delivers credentials once; bootstrap harness proves direct(zero-appliance-calls)/pairing/delivery; artifact proven secret-free (baked env = hub URL only) | **F1 loader caveat:** `--loader mkimage` fixes cheap AMI firmware that can't USB-boot the stock GRUB — UNSIGNED → **Secure Boot must be OFF**; default `shim` keeps SB. **Slice C bind is operator-password-gated** (CC stages, Viktor binds) → the live boot→register→bind→day-0 composition + physical N100 boot fold into the supervised rehearsal (R-1). Customer-facing **self-bind page = R-27 slice 1 SHIPPED (hub v0.66.0, 2026-07-17)** — see the dedicated self-bind row | **Second board, 2026-07-21 (demo-hp, HP t740 / Ryzen V1756B / AMI M42):** the whole chain ran on virgin hardware in one pass — armed install → self-registration as an unclaimed appliance → operator bind → day-0 → running guest 9201 + agent 0.92.1 as `demo-hp-bb76ea`. **The shim loader booted with Secure Boot ENABLED**, which retires the assumption that Felhom installs need SB off — that was an N100-firmware workaround. The exact-serial disk filter took the system SSD and left the box's 1TB NVMe untouched/unenrolled on hardware it had never seen. Two failures filed rather than smoothed over: **R-59** (no DHCP → the installer baked a static fallback instead of aborting) and **R-61** (baked root password unknowable → no console access). | Box survives a wrong-NIC install: hub-unreachable first boot → legible Hungarian console screen (NIC table + remedy) + NIC sweep self-heal (bounded DHCP + hub probe per NIC, success-only persist), and the baked root password is operator-knowable (`.rootpw.txt`) | scripts v1.24.0 (`scripts/iso/felhom-bootstrap.sh` `network_gate`/`sweep_nics`, `build-felhom-iso.sh` rootpw emission) | **PROVEN-LIVE (nested drill — nested ≠ metal: metal proof rides the next real multi-NIC install)** | `audits/SPIKE-firstboot-nic-sweep-2026-07-22.md` — dead-NIC install from the virgin v1.24.0 ISO baked the 192.168.100.2 fallback (WITH a dead default gateway), the R-59 screen painted on the console (screendump captured), and after the cable move the box swept to the working NIC, re-leased and **self-registered at the hub unaided in under a minute**; the drill also caught + fixed the stale-fallback-route trap (flush before the bounded dhclient) and verified the emitted rootpw against the installed box's shadow hash | R-59 ships as a first-boot gate, not an install-time abort (recorded deviation — the fallback is the auto-installer's own, initrd hook out of scope); sweep is structurally first-boot-only (`state.json` gate + unit done-flag condition); a box past install-start gets the screen but its interfaces are never touched | | Customer claim: one-time emailed code → customer sets own password (bcrypt, operator never sees it) | controller v0.122, hub v0.50 | **PROVEN-LIVE** (drill VM) | `DRILL-day0-vm-2026-07-12` §10/F-4 (gate ON via real edge; claimed, code consumed) | Never executed by a non-Viktor human → R-3. **Deliverability (R-4), gmail half DONE 2026-07-18:** the rehearsal's claim email was the first sent under the tightened DMARC `p=quarantine` and **landed in the gmail Inbox, not spam** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`). **freemail.hu remains Viktor's open half.** (Dropped mis-cited `CAMPAIGN-4` F-C — that is the escrow-claim 502, not password claim) | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index e045ace8..7338c509 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -13,9 +13,9 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha | **R-88b** | ~~`/backup/due` cannot say *unknown*~~ | **SHIPPED + PROVEN-LIVE** (agent v0.105.0 + controller v0.178.0, 2026-07-27) | — | `age_state=unknown` captured on real hardware during a deliberate ep0 outage; controller deferred, **zero app stacks stopped** | — | | **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC | | **R-94** | ~~A hand-synced version constant drifts, and the gate that would catch it is never run~~ | **CLOSED — SHIPPED** (hub v0.87.0, 2026-08-02) | — | **All three legs closed.** **(a) closed by DELETION, not derivation** — deriving is not achievable honestly: the Setup command fetches `felhom-host-install.sh` at RUN TIME from a website that git-syncs `main` every 30 s (R-110), so no build-time value in the hub can be true, and a number that is wrong carries a version number's authority while being a guess. The const, the `pageData.ScriptVersion` field, its assignment and the rendered label are gone; a NOTE stands where the const was so it is not helpfully re-added. **(b)** `hostinstall_gates.py` gate 1 INVERTED — it now asserts the hub carries **no** host-install version literal, in six code shapes across every `.go`/`.html` under `hub/`; and the gate is now invoked, by `scripts/repo_gates.py` and the pre-push hook (→ R-29). **(c)** the tautological `render_test.go:219` assertion is deleted, not replaced — there is no version to assert. It was demonstrated PASSING with the const at `9.9.9` while the script was 1.22.0. The label had been wrong for 19 days (since 2026-07-14) | — | -| **R-110** | **`main` is the installer's publish channel — there is no staging.** `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a 30 s period and nginx serves that working tree directly (`location /scripts/`, `root …/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`) receives. There is no tag, no pinned-version path, no staging copy and no rollback other than another push — for the artifact that runs as **root on a virgin box**, the single most privileged thing Felhom ships | **READY (S)** — operator ruling taken 2026-08-03: **option (b), tag-tracked** | — | **Two consequences worth stating:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29 — and the precaution recorded on the old R-94 row as "do not point every new box at an installer that has never run" **was never available to take**. **Open question for the operator, not a defect to fix blind:** whether `/scripts/` should serve a pinned release (tag-tracked path, or a versioned directory with the customer command naming a version) or whether `main`-tracking is the accepted shape for a one-operator product. Exposure today is zero — there are no boxes installing — which is exactly why it is cheap to decide now. **SECOND INSTANCE, found 2026-07-29 by the E-2d run and filed here rather than as a new ID:** `felhom-host-install.sh` fetches **nine** files from `raw/branch/main` (`:2072`–`:2206`) and the hub manifest vouches a sha for exactly **one** (`wrapper_sha256` → `felhom-pbs-apply`; re-checked this run, no drift). E-2a's `felhom-backup-target-apply` (`:2116`) is installed **0755 to `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n` — a root-executed artifact taken from `main` with no pinned integrity, which is this row's class exactly. **OPERATOR RULING 2026-08-03 — option (b) chosen: the publish channel moves from `main`-tracking to a TAG.** Publishing becomes *moving the tag*, and rollback becomes *moving it back* — the property `main`-tracking cannot have at any price. Recorded here, **not built this session, by instruction**. **The ruling carries a condition that decides whether the fix works at all: it must cover BOTH channels.** (i) the nginx-served `/scripts/` git-sync (`manifests/webpage.yaml`, `--branch=main`, 30 s period) that every `felhom-bootstrap.sh` fetch and every operator day-0 command reads, AND (ii) **the nine files `felhom-host-install.sh` fetches from `raw/branch/main`** (`:2072`–`:2206`), of which the hub vouches a sha for exactly ONE. Fixing only (i) leaves a tagged installer pulling nine untagged files from `main` at run time — a staging story that is false in the place it matters most, since one of those nine (`felhom-backup-target-apply`) is installed **0755 into `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n`. **Exposure is still zero** (no boxes installing), which is exactly why it stays cheap. Now CC's to build | CC | +| **R-110** | **`main` is the installer's publish channel — there is no staging.** `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a 30 s period and nginx serves that working tree directly (`location /scripts/`, `root …/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`) receives. There is no tag, no pinned-version path, no staging copy and no rollback other than another push — for the artifact that runs as **root on a virgin box**, the single most privileged thing Felhom ships | **CLOSED — SHIPPED** (installer v1.23.0, 2026-08-03) | — | **Two consequences worth stating:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29 — and the precaution recorded on the old R-94 row as "do not point every new box at an installer that has never run" **was never available to take**. **Open question for the operator, not a defect to fix blind:** whether `/scripts/` should serve a pinned release (tag-tracked path, or a versioned directory with the customer command naming a version) or whether `main`-tracking is the accepted shape for a one-operator product. Exposure today is zero — there are no boxes installing — which is exactly why it is cheap to decide now. **SECOND INSTANCE, found 2026-07-29 by the E-2d run and filed here rather than as a new ID:** `felhom-host-install.sh` fetches **nine** files from `raw/branch/main` (`:2072`–`:2206`) and the hub manifest vouches a sha for exactly **one** (`wrapper_sha256` → `felhom-pbs-apply`; re-checked this run, no drift). E-2a's `felhom-backup-target-apply` (`:2116`) is installed **0755 to `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n` — a root-executed artifact taken from `main` with no pinned integrity, which is this row's class exactly. **OPERATOR RULING 2026-08-03 — option (b) chosen: the publish channel moves from `main`-tracking to a TAG.** Publishing becomes *moving the tag*, and rollback becomes *moving it back* — the property `main`-tracking cannot have at any price. Recorded here, **not built this session, by instruction**. **The ruling carries a condition that decides whether the fix works at all: it must cover BOTH channels.** (i) the nginx-served `/scripts/` git-sync (`manifests/webpage.yaml`, `--branch=main`, 30 s period) that every `felhom-bootstrap.sh` fetch and every operator day-0 command reads, AND (ii) **the nine files `felhom-host-install.sh` fetches from `raw/branch/main`** (`:2072`–`:2206`), of which the hub vouches a sha for exactly ONE. Fixing only (i) leaves a tagged installer pulling nine untagged files from `main` at run time — a staging story that is false in the place it matters most, since one of those nine (`felhom-backup-target-apply`) is installed **0755 into `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n`. **Exposure is still zero** (no boxes installing), which is exactly why it stays cheap. **CLOSED 2026-08-03 — installer v1.23.0, and the both-channels condition was HONOURED, but not in the shape the ruling assumed.** **The spec's mechanism for channel 2 rested on a factual error, found by reading the code:** the run-time fetches are **sixteen, not nine**, and they come from the **`felhom-agent`** repo, not from `felhom.eu` — so no tag on this repo could ever have covered them, and `§8.1`'s *"derive the tag from `SCRIPT_VERSION`"* was unimplementable for them. Operator ruled on the alternative: pin them to **the agent version being installed**, which the installer already resolves from the hub manifest and already sha-verifies. `fetch_raw` now fetches `raw/tag/v$ART_AGENT_VER/`, with **no fallback to a branch** — a vouched version whose tag is missing fails loudly, because a silent fallback is the appearance of control with none of it. That also fixed a latent skew → **R-183**. **Channel 1:** `manifests/webpage.yaml` split into TWO git-syncs — the website still tracks `main` at 30 s (a copy edit must never need a release), `/scripts/` tracks **`installer-v1.23.0`**. Both trees are seeded by init containers, so a fresh pod is not Ready until the tag is checked out and there is no 404 window; `maxUnavailable` rounds to 0 on one replica, so a failed scripts-init leaves the OLD pod serving — the failure direction is *no update*, never *no /scripts/*. **Channel 3 needed no change, and that is recorded rather than left as a silence:** `https://felhom.eu/scripts/felhom-host-install.sh` never carried a ref — the ref lives in the manifest — so `felhom-bootstrap.sh` and the hub's day-0 command follow the tag with **no edit and no hub version bump**, which is why §1's no-bump rule was never in tension. **PROVEN LIVE, both scenarios, by HTTP against the real URL.** *P-A (measured before the manifest was touched):* git-sync v4.4.0 follows a tag **and notices a MOVED one** — `update required … local:fb65202 remote:8360f94` → `updated successfully`, one period. *Scenario A:* a real push to `main` without moving the tag — the website tree advanced to the new commit while the scripts tree stayed put, the served sha stayed **byte-identical (`2f859555…`)** and a marker comment deliberately planted in that commit was **absent** from the served URL. *Scenario B:* moving the tag published it in ~40 s (sha → `ea2b4aa9…`, marker present), and moving it back rolled it back to **exactly** the pre-publish sha with the marker gone; `https://felhom.eu/` returned 200 throughout. **Gate 6 in `hostinstall_gates.py`** pins all of it structurally with no network, so it stays in `--fast` and runs in CI: no `raw/branch/` ref in the installer, `fetch_raw` still pinned, the manifest still splitting tag-vs-main. **It deliberately does NOT assert that a tag exists for the current `SCRIPT_VERSION`** — that would go red on the very push that bumps the version, before publishing, and publishing being a separate act is the whole ruling; the same reasoning §8.4 applies to the agent gate | — | | **R-111** | ~~**The Day-0 artifact channel is 17 agent releases stale — a box installed today gets agent `0.96.0`, not `0.113.0`.**~~ `felhom-host-install.sh` does not use `main`: it reads the hub-vouched manifest (`:423-436`) and fetches Gitea generic packages (agent `:1945`, golden `:2573`). Gitea's newest are **agent 0.96.0** and **golden 0.161.0**, and the hub's manifest selects exactly those — so a fresh box lands on **agent 0.96.0 + controller 0.161.0** (global floor `v0.156.0` < the golden's 0.161.0, so no self-update) against `main`'s 0.113.0 / 0.185.1. Agent 0.113.0 reached both demo boxes by **direct deploy and was never published** | **SHIPPED 2026-07-29 — the channel now serves agent 0.113.0 + golden 0.185.1** | — | **FIXED the same day it was found.** Agent **0.113.0** built from the clean tree @ `58b598b` and published (`scripts/publish-agent.sh`), sha `5f3247f756cb658e…`, round-trip GET verified. Golden **0.185.1** baked on the nested drill VM embedding controller `0.185.1`, published, sha `dba00f3e845c415e…` — bake clean: `Result=success`, overlay2, **all 3 mounts included** (rootfs+mp0+mp1), 0 FATAL/exclusions, upload HTTP 201, token-leak grep 0; log `drill/bake-0.185.1.log`; GL-1 teardown done (guest 9100 purged, secrets shredded, disk restored to `virgin`). Hub Day-0 manifest moved **both together in one POST** so it never vouched a new agent against an old golden; `min_agent` **0.93.0 → 0.113.0**, which is what controller v0.185.0 declares (`felhom-controller/CHANGELOG.md:15`) — **zero fleet impact, verified: all three enrolled hosts already run agent 0.113.0, so no box is held.** `wrapper_sha256` preserved verbatim (re-checked against `configs/felhom-pbs-apply` — no drift). **The global controller floor was deliberately NOT raised**: the golden now bakes 0.185.1, so a fresh box needs no self-update, and raising it would have been an unnecessary fleet-wide write. Original finding follows. **Found 2026-07-29 by the E-2d Phase 0 gate, which stopped the run before a VM was created.** 17 unpublished releases (v0.97.0–v0.113.0) strand the **entire R-82 tiered-backup arc** plus **F-CRIT-2** (a failed backup looking fresh — 7 days silent) and **F-REBOOT** (a guest rebooted mid-backup never returns): a new customer's box would install without them. **Blocks E-2d's C3/C4/C5** — those test endpoints and events that do not exist in 0.96.0/0.161.0. The **controller is fine** (registry has 0.185.1, floor-driven self-update), so the gap is specific to the two Gitea-generic artifacts. **Mirror of R-110, not a duplicate:** R-110 = the installer publishes instantly with no staging; R-111 = the agent/golden publish gate exists and was never walked. Fix should decide whether publishing joins the release train rather than staying a remembered step (R-29's shape, one layer up). Evidence: `audits/E2D-fresh-vm-2026-07-29.md` **DEFERRED LEG, AND IT RECURRED → R-115.** This row's shipped half stands and is not reopened: the bump happened, was verified, and was proven end-to-end by the E-2d install. But its own closing line — *decide whether publishing joins the release train rather than staying a remembered step* — was never acted on, and agent 0.114.0 reproduced the exact condition the same afternoon. The recurrence is filed as **R-115**, not as a reopen, because the stale-channel finding is closed while the process defect that caused it is a distinct problem with a distinct fix. | CC | -| **R-115** | **Publishing is a remembered step, and it was forgotten within eight hours of being documented as forgettable.** A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published — so "deployed" and "installable" are independent states that drift silently. **Two instances, both real:** **R-111** (2026-07-29 morning) — 17 agent releases v0.97.0–v0.113.0 stranded, so a new customer would have installed without the entire R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT; found only because the E-2d Phase 0 gate happened to look. **Agent 0.114.0** (same afternoon) — the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, and **unpublished until this task**, which blocked Session C: a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix | **READY (M)** — operator ruling taken 2026-08-03: **mechanism (b), build-side gate** | — | **The finding is the RECURRENCE, not either instance** — both instances are fixed. R-111's own text already named this leg (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; the leg then recurred the same day, which is the evidence that a note is not a mechanism. **Class: → R-29, one layer up** — a control that exists and is never walked; deliberately NOT given its own ID. **The decision is the operator's; the options, mechanisms first:** (a) **publish as a step in the build/release path**, so deployed and installable cannot diverge; (b) **a gate that refuses to deploy a version that is not published+vouched** — the strongest, and it fails closed; (c) a session-end checklist entry; (d) accept it as manual and add a pre-Session-C verification. **(a) and (b) are mechanisms; (c) and (d) are reminders — and R-29's whole finding is that reminders do not hold.** No code this session by design. **THIRD INSTANCE, 2026-08-03 — and it was found by a runbook that had been told there was nothing left to do.** Agent **v0.120.0** — the agent half of the R-165 merge — was built, committed at `cd6e267`, and deployed to BOTH demo hosts, and was **never published**: `GET …/generic/felhom-agent/0.120.0/felhom-agent` → **HTTP 404** (0.119.0 → 200), and the hub manifest accordingly vouched **0.119.0**. The consequence is the sharpest yet, because installer step 5's idempotent skip requires `installed == vouched` EXACTLY: a documented-path reinstall would have **downgraded both boxes** from the merge-aware 0.120.0 to the pre-merge 0.119.0 — silently, since the current `step_grows` sets `SYSDATA_GROW=0` so 0.119.0's `mp1` resize (`bringup.go` 4c, fatal on error) never fires and the install would have *succeeded* while proving a stack nobody ships. **R-178's own row asserted `agent v0.120.0 is live on BOTH hosts` and `no code left to write`; both were true and both were beside the point** — the gap was publication, which no one checks. Fixed in-session on the operator's ruling: `scripts/publish-agent.sh 0.120.0` (sha `a7763d31b55b5ce75457b4dba7b06aa300325811834b0be78af4587b47110b9d`, round-trip GET verified) then vouched, and both reinstalls then fetched and sha-verified it from Gitea. **This is the third instance of a row that has been WAITING-ON-OPERATOR since 2026-07-29; option (b) — a gate that refuses to deploy or vouch an unpublished version — would have caught all three.** **OPERATOR RULING 2026-08-03 — mechanism (b), build-side: a gate that REFUSES to deploy or vouch a version that is not published.** It is the strongest of the four options and the only one that fails closed; (c) and (d) were reminders, and R-29's whole finding is that reminders do not hold. Recorded here, **not built this session, by instruction**; it is now CC's to build. **The third instance is the argument for the ruling and belongs inside it:** agent **v0.120.0** (the agent half of the R-165 merge) was built, committed and deployed to BOTH demo hosts while `GET …/generic/felhom-agent/0.120.0/felhom-agent` returned **HTTP 404**, so the hub vouched 0.119.0. Because installer step 5's idempotent skip requires `installed == vouched` EXACTLY, a documented-path reinstall would have **silently downgraded both boxes** to the pre-merge agent — and would have *succeeded* while doing it, since the current `step_grows` sets `SYSDATA_GROW=0` so 0.119.0's fatal `mp1` resize never fires. A gate at deploy/vouch time is the only one of the four standing between that and the operator | CC | +| **R-115** | **Publishing is a remembered step, and it was forgotten within eight hours of being documented as forgettable.** A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published — so "deployed" and "installable" are independent states that drift silently. **Two instances, both real:** **R-111** (2026-07-29 morning) — 17 agent releases v0.97.0–v0.113.0 stranded, so a new customer would have installed without the entire R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT; found only because the E-2d Phase 0 gate happened to look. **Agent 0.114.0** (same afternoon) — the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, and **unpublished until this task**, which blocked Session C: a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix | **CLOSED — SHIPPED** (`release-agent.sh` + `check-published-versions.py`, 2026-08-03) | — | **The finding is the RECURRENCE, not either instance** — both instances are fixed. R-111's own text already named this leg (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; the leg then recurred the same day, which is the evidence that a note is not a mechanism. **Class: → R-29, one layer up** — a control that exists and is never walked; deliberately NOT given its own ID. **The decision is the operator's; the options, mechanisms first:** (a) **publish as a step in the build/release path**, so deployed and installable cannot diverge; (b) **a gate that refuses to deploy a version that is not published+vouched** — the strongest, and it fails closed; (c) a session-end checklist entry; (d) accept it as manual and add a pre-Session-C verification. **(a) and (b) are mechanisms; (c) and (d) are reminders — and R-29's whole finding is that reminders do not hold.** No code this session by design. **THIRD INSTANCE, 2026-08-03 — and it was found by a runbook that had been told there was nothing left to do.** Agent **v0.120.0** — the agent half of the R-165 merge — was built, committed at `cd6e267`, and deployed to BOTH demo hosts, and was **never published**: `GET …/generic/felhom-agent/0.120.0/felhom-agent` → **HTTP 404** (0.119.0 → 200), and the hub manifest accordingly vouched **0.119.0**. The consequence is the sharpest yet, because installer step 5's idempotent skip requires `installed == vouched` EXACTLY: a documented-path reinstall would have **downgraded both boxes** from the merge-aware 0.120.0 to the pre-merge 0.119.0 — silently, since the current `step_grows` sets `SYSDATA_GROW=0` so 0.119.0's `mp1` resize (`bringup.go` 4c, fatal on error) never fires and the install would have *succeeded* while proving a stack nobody ships. **R-178's own row asserted `agent v0.120.0 is live on BOTH hosts` and `no code left to write`; both were true and both were beside the point** — the gap was publication, which no one checks. Fixed in-session on the operator's ruling: `scripts/publish-agent.sh 0.120.0` (sha `a7763d31b55b5ce75457b4dba7b06aa300325811834b0be78af4587b47110b9d`, round-trip GET verified) then vouched, and both reinstalls then fetched and sha-verified it from Gitea. **This is the third instance of a row that has been WAITING-ON-OPERATOR since 2026-07-29; option (b) — a gate that refuses to deploy or vouch an unpublished version — would have caught all three.** **OPERATOR RULING 2026-08-03 — mechanism (b), build-side: a gate that REFUSES to deploy or vouch a version that is not published.** It is the strongest of the four options and the only one that fails closed; (c) and (d) were reminders, and R-29's whole finding is that reminders do not hold. Recorded here, **not built this session, by instruction**; it is now CC's to build. **The third instance is the argument for the ruling and belongs inside it:** agent **v0.120.0** (the agent half of the R-165 merge) was built, committed and deployed to BOTH demo hosts while `GET …/generic/felhom-agent/0.120.0/felhom-agent` returned **HTTP 404**, so the hub vouched 0.119.0. Because installer step 5's idempotent skip requires `installed == vouched` EXACTLY, a documented-path reinstall would have **silently downgraded both boxes** to the pre-merge agent — and would have *succeeded* while doing it, since the current `step_grows` sets `SYSDATA_GROW=0` so 0.119.0's fatal `mp1` resize never fires. A gate at deploy/vouch time is the only one of the four standing between that and the operator. **CLOSED 2026-08-03 — both halves, no version bump (no Go code changed).** **(1) `scripts/release-agent.sh` is now THE release path**: build → **tag** → publish → **verify by an INDEPENDENT download**. It calls the existing `publish-agent.sh` rather than reimplementing it, refuses a dirty or unpushed tree, refuses to re-release an existing version (one version name must never mean two binaries), and **deliberately does not vouch** — vouching points machines at a version and stays the operator's act. `CLAUDE.md`'s raw `go build` line is replaced by it, so the documented way to release cannot complete without publishing. It also **tags**, because R-183 made the tag part of the released artifact. **(2) `scripts/check-published-versions.py`**, registered in `agent_gates.py` as **not `--fast`** — it needs network, and a push must not fail because Gitea blinked. **The CI workflow now runs the FULL set instead of `--fast`**, without which the gate would have been registered and never run: the built-but-never-wired failure this project has shipped four times. **THE INVARIANT IS NOT THE ONE THE TASK SPECIFIED, and the reason was measured (P-C), not argued.** §8.4 asked for *"the version the hub tells machines to install must be downloadable"* — the better invariant, and **CI cannot see it**: the hub's artifact manifest is **401** without a per-customer passphrase and Gitea's package **listing** api is **401** without a token, while the package **download** url and the **tags** api are anonymous. Adding an operator credential to CI is the operator's call, not a gate author's. The implemented invariant — **every `v` tag must have a downloadable package and a tag tree serving the agent's configs** — needs no credential and **catches all three recorded instances**, because the release script creates the tag and publishes in one act. **What it does NOT catch is stated rather than assumed away: the hub vouching a version that was never released at all → R-184.** **Red-proof F, MEASURED ON REAL CI and not inferred:** runs **69** and **70** are on the *same commit* `0db7766` — **success** before a tagged-but-unpublished `v9.9.9` existed, **failure** after pushing it. Same code, same workflow, one variable. Locally: the gate exits 1 naming the 404; **deregistered from the entry point** the same bad state reports `all agent gates OK` **rc=0**; restored → `CONVICTED: published` rc=1. `v9.9.9` deleted afterwards (`git ls-remote --tags` → only `v0.120.0`). **One deliberate CI failure e-mail reached the operator at ~12:5x CEST — that was this proof, not an incident** | — | | **R-116** | ~~**The drive-absent alarm and its recovery were a MISMATCHED PAIR — absent fired the GENERIC `storage_disconnected`, return the SPECIFIC `backup_target_restored`; `backup_target_absent` never fired at all**~~ | **SHIPPED + PROVEN-LIVE** (agent v0.116.0, 2026-07-30) | — | **CLOSED. The full four-event sequence, on the wire, on a fresh box** (`audits/R116-v0116-2026-07-30.md`): `backup_target_absent (error)` on detach → `backup_target_restored (info)` on return for the TARGET, and `storage_disconnected (error)` → `storage_reconnected (info)` for a NON-target drive on the same box four minutes apart. **Two matched pairs, correctly discriminated — and discrimination is proven NON-trivially for the first time**, since both prior runs had the target itself emit the generic event. Gate fired in **3 s**; all four events reached the hub, so the specific alarm, its severity, its Hungarian copy and the hub routing are now exercised end-to-end. **Over-correction PASSES with a positive observable** (0 ABSENT lines / 0 drive events over 2m14s with both drives present, target `degraded:false`, while 2 `RETURNED` lines prove the gate was ticking). **NARROWED by the R-117 spike (`audits/SPIKE-r117-bind-liveness-2026-07-30.md` §12), and it stands as written:** the 2 `RETURNED` lines are a genuine positive observable, so rule 3 is satisfied — but `degraded:false` over that window was read off a drive whose bind was **dead** (R-117), so the window evidences **"the gate did not over-fire"** and **NOT** **"the drive was healthy."** No other part of this row changes: every input to the pairing fix is configuration-derived (`storage.cfg`'s `path` vs the `.mount` unit's `Where`), which R-117 does not touch. Ran on a nested PVE on **demo-hp** per `runbooks/target-selection.md` — through the **real day-0** from the v1.25.0 ISO, with the agent **installed unaided from the vouched Day-0 manifest** (published sha `b47c5c4dab641ee5…`, independent registry GET verified, manifest read back), drives enrolled through the real endpoints, device loss a real hot-detach. **THE FIX, and the ruling is the substantive part:** the mechanism was first isolated from the captured payload (`DIAG-r116-disks-payload-2026-07-30.md`) after two fixes aimed at shapes that do not occur. **Both smaller-looking options were REJECTED because they regress R-114** — `backup_target_offer.go:79` reads `BackupTarget && MountPath != ""` as *"a real drive with its own mountpoint — healthy"* and returns before its `TargetAbsent` branch, so back-filling `MountPath` on the Observe row **or** flagging the registry row (whose `MountPath` is the stale unit-file value) would have told the customer the backup target is fine while its drive was gone. **R-114's correctness was resting on R-116's bug** — a coupling invisible until the payload existed. Taken instead: the Observe row gets the **guest path only** (`mount_path` stays `""`, which is true) from a new `ConfigPath` (`json:"-"`, so the cross-repo golden + key-set contract is untouched), and the union row is deduped **on guest path** — the join being CONFIGURATION (`storage.cfg`'s `path` vs the `.mount` unit's `Where`), the only identity that survives the device. Tests 845→849; 4 red-proofs each asserted to land, and red-proof 1 replays v0.115.0's code and fails, which is the empirical proof it was inert. Its green test had supplied a `MountPath` production never supplies AND left `DriveTargets` nil so the union loop never ran — both corrected. v0.115.0 left in place (inert, harmless). Teardown all 3 layers; hub layer gate-blocked on ONLINE with the command recorded. **Caveat: the drill's controller was 0.185.1 from the golden, which PREDATES R-114**, so its absent-state banner showed the old false copy — the golden being a release behind, not a regression → **R-120** | — | | **R-120** | ~~**The golden baked a controller that predated R-114 + R-112, so a FRESH box showed the customer the WRONG absent-target message**~~ | **CLOSED — golden rebaked + PROVEN-LIVE, and the class now has an ENFORCED gate** (golden 0.186.0 + hub v0.82.0, 2026-07-30) | — | **`audits/R120-golden-rebake-2026-07-30.md`.** **Half 1 — the artifact.** Golden **0.186.0** baked from `main`'s controller in the DooPlex bake fixture (overlay2 OK, **3 mounts**, FATAL 0, exclusions 0, 618 MB, upload **201**, `GOLDEN_SHA256=b760ac6a33e70700…`, token-leak grep 0, GL-1 teardown, `drill.qcow2` back to `virgin`). Three observables: **published** — anonymous GET (what the installer does) 200 / 648930639 bytes / sha identical to the bake; **vouched** — manifest read BACK; **resolved** — `Artifact manifest served for customer sess-f (agent=0.116.0 golden=0.186.0)`. Floor **untouched** per publish-train rule 2 (`min_controller_version` still 0.156.0; it is a separate form); MinAgent left 0.113.0 as 0.186.0 declares. **Proven on a REAL day-0, not the fixture** (per the Part-1 rule now in `runbooks/target-selection.md`): VM 9402 on demo-hp from the v1.25.0 ISO → `Controller elindult (0.186.0)`. With the target detached the endpoint returned the **`TargetAbsent`** copy — *„A rendszermentés meghajtója nem érhető el — amíg vissza nem csatlakoztatod…"* — **and `offer_path` absent entirely**; the day-old read on the 0.185.1 golden had returned the false system-disk message **plus** an offer of the other drive. **Half 2 — the mechanism, operator ruling REFUSE.** hub **v0.82.0**: the gate sits in `hub/internal/web/configs.go` `handleSetArtifacts` immediately before the only write — the sole UI path to `SetArtifactManifest` — so it runs on every vouch without anyone choosing to, and it **refuses** rather than warning. Signal: `store.NewestReportedControllerVersion()` over `reports.controller_version`, **semver-compared in Go** (`MAX()` in SQL ranks 0.99.0 above 0.186.0 — a pair this fleet has shipped). Fail-open in exactly two deliberate cases: empty golden field, unknown fleet version. **NEAR-MISS RECORDED: the first draft read `guests.controller_version`, a column that exists and that NOTHING writes** — it would always have seen `""` and failed open, i.e. inert, this gate's own failure shape, one grep from shipping. 4 tests through the **production handler** over httptest (never a seam), the refusal asserting **both** the flash **and** that the manifest was not written; red-proof: deleting the block makes the stale golden vouchable again. **PROVEN LIVE on the deployed hub by re-attempting the original mistake:** vouching 0.185.1 → `HTTP 303 …flash=golden_behind_fleet` + `[WARN] artifact vouch REFUSED: golden 0.185.1 is older than the newest controller the fleet reports (0.186.0)`, and the manifest read back **unchanged at 0.186.0**. Recorded on **R-29's audit list** (`ROADMAP.md`) as the **first enforced gate** beside its three orphans, so the contrast is kept — the orphans are unchanged. Teardown all 3 layers; hub layer gate-blocked on ONLINE with the command recorded, exactly as `sess-e` was (and `sess-e` was deleted this run) | — | | **R-117** | ~~**A drive's guest bind becomes a DEAD MOUNT while every signal reads healthy — and it happens in TWO ways, only one of which the original framing covered.** (a) *after a detach/return*: the host raw mount heals onto the NEW device via its fs-UUID-keyed unit while the bind still names the OLD one, so the gate takes its `Return` branch and restarts the customer's apps onto a namespace that `EIO`s on every call; (b) *in STEADY STATE, no cycle at all* — a device that errors without disappearing leaves the raw mount `active`, `BoundUnderParent` `true` and the drive never `Disconnected`, so **the gate produces no action and NOTHING is emitted on any channel**~~ | **SHIPPED + PROVEN-LIVE** (agent **v0.117.0**, 2026-07-30) | — | **CLOSED. `audits/R117-v0117-2026-07-30.md`.** `BoundUnderParent` gains a THIRD term at both /disks sites: `bindLiveness` reads `/proc` only and requires (a) **the bind names the same device as the raw mount** and (b) **the filesystem has not aborted** (`shutdown` **or** `emergency_ro`, both measured). **BOTH CHECKS ARE LOAD-BEARING and this is the substantive part:** R-117 was filed as a detach/return defect, but a device that fails WITHOUT disappearing gives the identical all-signals-healthy state with the **devnos EQUAL** and the drive never `Disconnected`, so the gate emits nothing at all, indefinitely (R-117a) — the device comparison alone cannot see it, and a P1-only fix passes every payload test (red-proof RP3 exists for exactly that). **THREE states, never a bool:** `{Unknown, Live, StaleDevice, Aborted}`, `Unknown` is the zero value, and every caller reads `Usable()` where unknown counts **PRESENT** (absent stops a customer's apps — the `newestArchiveOn` trap). **NO NEW RECOVERY PATH:** `AttachDrive`'s normalize leg already did the repair and three call sites already invoked it (20 s ticker, agent startup, and **the controller's `Return` branch BEFORE `restartStacks`**); all three were defeated by `if n == 1 && GuestSeesMount(...)` logging *"fully live, no-op"* about an EIO namespace. **RULING (asked for, given, flagged for overrule):** `StaleDevice` ⇒ **re-bind** (the raw mount is a healthy new superblock; repairs live, guest never restarts — init PID identical); `Aborted` ⇒ **quiet no-op and SURFACE**, because a re-bind lands on the SAME dead superblock and this runs every 20 s = an infinite silent retry that masks the state. No operator decision required: it routes an already-broken state into the **existing** gate, event types and Hungarian copy — no new customer-facing concept — and the alternative is apps writing documents into a filesystem that rejects every write. **ORDERING TRAP caught by a test:** abort-first classifies the real return state as aborted (its stale bind carries `shutdown` too) and refuses the repair **while still reporting correctly**, so the abort flag is read off the RAW mount in the stale case. **LIVE on demo-hp** (brought 0.113.0 → 0.117.0 first — see R-121): RETURN `raw 8:32 / bind 8:16 shutdown` ⇒ `stale-device`, usable **false**; IN-PLACE `both 252:11 emergency_ro`, raw unit still `active` ⇒ `filesystem-aborted`, usable **false**; healthy ⇒ `live`; **340–497 µs**. **No block I/O proven by strace** (only `/proc/self/mountinfo`, **0** statfs) — the Part 1 `CLAUDE.md` fence applied to its own first consumer. **No regression through the REAL pipeline:** `GET /disks` with the controller's own credential shows the live backup-target drive `bound_under_parent=True`, with 32 gate lines in 3 min as the positive observable and zero spurious transitions. Tests **849→863**, 29/29 green, **6 red-proofs each verified to land** — and **RP1 failing to fail exposed a HOLLOW test**: the aborted fixture used a `/dev/mapper` device, for which `RoleForStorage` derives `role=system`, and a system row never runs the conjunction, so it reported false by DEFAULT and no mutation could fail it. Fixtures now assert the production row shape first. Teardown all 3 layers; hub layer = the vouched manifest, **retained** (it is the product, not scratch). **NOT covered:** the stale-bind repair on hardware — `StablePathForRaw` hardcodes the live parent, so it would write into guest 9201's namespace (R-117h); and sustained-load behaviour, still unmeasured. Follow-ups **R-117g** (no guided recovery for an aborted fs), **R-117h** (parent dir not test-seamable), **R-121** | CC | @@ -92,7 +92,9 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha | **R-174** | ~~**The app-stop guard's crash recovery started apps onto MISSING drives — a regression in v0.189.0 code.**~~ | **CLOSED — SHIPPED + PROVEN-LIVE** (controller v0.191.0, 2026-08-02) | — | **Found by REVIEW on 2026-08-02, in code shipped 2026-08-01, and closed the same session — R-171 one path over.** `appStopGuard.SetStarter(stackMgr)` handed `Recover` the RAW stack manager, whose `StartStack` has no drive gate, and `Recover` runs **at startup** — exactly when an external drive may not have come back. So: a backup stops an app, the box loses power, the drive does not remount, and the app is started on a missing drive. The rule was not new — the API's own `startGatedByMissingDrive` already refused this to the customer; the guard bypassed it. **`bootDriveGate` could NOT be reused whole**, and the reason is recorded in the code: its holder #2 reads `bootAppStopGuard.HeldStacks()`, which during `Recover` is **the guard's own marker** — it would refuse every recovery it was meant to perform — and holders #1/#2 read package-level vars assigned AFTER `Recover()` runs, so a whole-gate reuse would be correct only by accident of nil-safety. Holder #3 is extracted into a shared `driveStartGate` with **two callers, one implementation**, and `TestBootDriveGateAndAppStopShareTheDrivePredicate` pins the delegation. **A REFUSAL IS NOT A FAILURE:** new `ErrStartRefused` + a `Refused` bucket — both keep the marker, only `Failed` alarms, because routing a deliberate hold into `NotifyBackupFailed` (customer-enabled by default) is the very R-171 false alarm this fixes. `main.go` guards on `Alarming()`, not `!= nil`, and the pre-existing seam test was TIGHTENED to require it. **Live on 9201, both directions:** drive held unmounted → `refusing to restart "calibre-web" … drive /mnt/felhom-drives/hdd_1 is not a live mountpoint`, marker retained byte-identical, zero containers started, `not alarming`; drive returned → `restarted calibre-web`, marker CLEARED. **ID established free:** `grep -ro "R-174\b" documentation/ *.md` → 0 hits | — | | **R-175** | ~~**`07-backup-architecture.md` §7.5 states ONE box's size bound as if it were the fleet's.**~~ | **CLOSED — FIXED 2026-08-03** (same pass as R-165) | — | **Measured, not inferred** (`audits/SPIKE-r165-mp1-merge-2026-08-02.md` M1: `pct config 9201` on both hosts). Independent of the merge — the sentence is wrong today and will be wrong differently after R-165. **The fix is to state the bound as a FUNCTION of `mp1`, not a constant**, and to say which box any quoted figure came from. Same class as the comment-asserting-an-invariant rule: a doc stating a fleet-wide number that only one machine satisfies reads as settled and is not. **ID established free:** `grep -ro "R-175\b" documentation/ *.md` → 0 hits **FIXED.** §7.5 gained a **7.5.1** which (a) states plainly that the bound is a FUNCTION of `mp1` and applies only to a box still on the split layout, naming all three real shapes, and (b) records that the ceiling itself has been removed by R-165 for boxes built from golden ≥ 0.192.0. Fixed in the same pass as the merge rather than filed and forgotten, because the section would otherwise have been wrong in two ways at once | CC | | **R-176** | **Two prerequisites for the R-165 merge are UNMEASURED, and both are cheap.** (a) Whether a **pre-merge archive** (carrying `mp1`) restore-tests cleanly into a **merged-layout** guest — reading `mountParity` (`felhom-agent/internal/reconcile/restoretest.go:347`) says it should, because the restore recreates `mp1` from the archive so archive and restored guest agree; **that was reasoned from source and never executed.** (b) The in-place per-box migration (move `/felhom-data` onto `mp0`, drop the slot, verify) has **never been rehearsed even once**, so "is the box restorable at every point of it?" is currently unknown | **(a) ANSWERED 2026-08-03 (P1: PASS). (b) NOT REQUIRED — operator ruling: every node is reinstalled, none migrated** | blocks R-165 landing safely | **Filed because this project's own record is that FOUR production designs specced against unvalidated mechanisms were all wrong** — which is exactly why R-165's own spike refused to design. Both are one command on a **Tier-0** box (D-d: both demo boxes are disposable). (b) is only required work if Peti's box turns out to need migrating rather than reinstalling — the hub cannot answer that (M5: `peti-felhom` exists as a customer with **no host in the register**), so it is the operator's input. **ID established free:** `grep -ro "R-176\b" documentation/ *.md` → 0 hits **UPDATE 2026-08-03.** **(a) is measured and passed** — `audits/SPIKE-r165-phase0-2026-08-03.md` P1: a real pre-merge archive (`mp0+mp1`, confirmed from its own vzdump log) restore-tested on demo-hp, `pass: true`, `mount_parity: ok`, 84 s, with `mountParity` untouched. One limit stated rather than glossed: it ran with the pre-merge agent because the merged one did not exist yet, and the comparison is archive-vs-its-own-restore which never consults the host layout — **re-run it once against agent v0.120.0**, which is one command. **(b) is withdrawn, not deferred:** the operator ruled that every node is REINSTALLED rather than migrated in place (both demo boxes are Tier 0; the colleague's box carries none of our customer data and is clean-installed in a few weeks), so the in-place migration rehearsal has no consumer. Recorded explicitly rather than silently skipped | CC | -| **R-182** | **The reserve re-alerts on every status refresh, so a full disk pages the operator on a timer rather than once.** `captureAllRecoveryUnits` has TWO callers: the backup run (`runDBDumpsInternal`), which since R-181 holds a per-run admission memo so a refused app alerts exactly once — and `GetFullStatus` (`backup.go:967`), the **periodic status refresh**, which calls it with no run scope. Outside a run the verdict is decided fresh, so every refused app pushes another `recovery_unit_capture_failed` (severity `error`) on every refresh | **READY (S) — NEW 2026-08-03** | — | **Measured live, not inferred.** demo-hp 2026-08-03: the backup run at **08:59:46** produced exactly one alert per app (2 apps → 2 events, correct); a **second identical pair fired at 08:59:59**, thirteen seconds later, from the status refresh. Under a sustained full disk that is a repeating operator page for a condition already reported. **PRE-EXISTING, NOT INTRODUCED BY R-181 — stated because it decides the priority.** v0.192.0's `captureAllRecoveryUnits` called `unitFloorBlocked` → `unitNotify` on the same path with the same frequency; R-181 changed neither caller. It is filed now because R-181's live proof is the first time anyone watched the event stream closely enough to see it. **Deliberately NOT fixed in the R-181 task**, whose §9.2 mandates minimal changes and whose §15 routes observations here; the fix changes alerting semantics on a path that task did not own. **The mitigating claim should be checked before the fix is scoped:** `backup.go`'s own comment says *"NO CONTROLLER-SIDE COOLDOWN — the hub owns cooldown"*. If the hub genuinely dedupes `recovery_unit_capture_failed`, this is log noise and an S; if it does not, it is a real repeated page. **That is exactly the shape of a comment asserting an invariant with no test pinning it** (`CLAUDE.md`'s own table), so verify it at the hub rather than trusting it. **Fix shape (not implemented):** give the status-refresh sweep its own admission scope, or make the non-run path decide without alerting — the run path is already correct and must not be disturbed | CC | +| **R-183** | **A fresh install fetched the vouched agent BINARY and its sixteen CONFIG files from two different refs, and nothing compared them.** `felhom-host-install.sh` resolved the agent version from the hub manifest and sha-verified the binary — then took `felhom-agent.service`, `felhom-agent.sudoers` and fourteen more from `raw/branch/main`, i.e. whatever the agent repo's tip happened to hold at that second. One install, two refs, no comparison | **CLOSED — SHIPPED** (installer v1.23.0, 2026-08-03) | — | **Found while implementing R-110, by reading `fetch_raw`'s call sites rather than the spec's description of them** — the task said nine files from `felhom.eu`; they are **sixteen** and they come from **`felhom-agent`**. **Why it is a defect and not only untidiness:** these files are the agent's own operating surface — its systemd unit, its sudoers, its guarded wrappers — and `configs/felhom-backup-target-apply` is installed **0755 into `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n`. A config newer than the binary is a root-executed artifact the vouched version was never tested against. **Not hypothetical in shape:** the fleet has shipped exactly this class before, where a config and the code that reads it moved independently. **Fixed by pinning to the agent version the install is already committed to**, on the operator's ruling: `raw/tag/v$ART_AGENT_VER/`, resolved from the hub manifest that the binary's sha is already checked against — so binary and configs now come from ONE ref by construction. **No fallback to a branch**, deliberately: a missing tag dies loudly rather than quietly serving `main`. Pinned by `hostinstall_gates.py` gate 6 (no `raw/branch/` anywhere; the `$ART_AGENT_VER` pin still present), red-proofed by reverting one of the sixteen and by removing both assertions. `felhom-agent` now carries `v` tags (`v0.120.0` created retroactively at `cd6e267`, the commit the published binary was built from; `configs/` is byte-identical there and at `main`, so nothing depended on the choice) and `release-agent.sh` creates them as part of releasing | — | +| **R-184** | **Nothing prevents the hub from vouching an agent version that was never released.** The R-115 gate proves every RELEASED version is installable, but it works from git tags — so a hub artifact-manifest entry naming a version with no tag and no package is invisible to it. The installer would then die at step 5 on a virgin machine, as root | **READY (S) — NEW 2026-08-03** | — | **Filed BECAUSE the R-115 gate deliberately does not cover it, rather than leaving the gap unstated.** CI cannot check it: the hub's `/api/v1/artifacts/` answers **401** without a per-customer retrieval passphrase and Gitea's package **listing** api answers **401** without a token (both measured 2026-08-03, P-C), so a credential-free gate can ask *"is this version installable"* but never *"which version is vouched"*. **Two shapes, and the second is better:** (a) give CI a hub credential — expands what CI can reach, and is the operator's call not a gate author's; (b) **validate at vouch time, in the hub**: the operator UI's Day-0 artifact form refuses a version whose package is not downloadable. (b) fails closed at the moment of the decision, needs no new credential anywhere, and puts the check where the mistake is actually made. **Exposure is low and should be said so:** vouching is a deliberate operator action against a version they have just released, and R-115's release path now makes released-but-unpublished nearly impossible. This is the residue, not the main risk | CC | +| **R-182** | **A full disk tells the operator about ONE app and silently swallows every other app's refusal for an hour.** The hub's operator cooldown key is `customerID + ":" + eventType + cooldownTierSuffix(details)` (`hub/internal/notify/dispatcher.go:268`). `recovery_unit_capture_failed` carries **`app`** in its details and **no `tier`**, so the suffix is empty and the key contains **no app identifier**: the first refused app's alert takes the 1-hour slot and the second app's is dropped — and dropped BEFORE `LogNotification`, so it leaves **no row on any channel**. It cannot even be audited after the fact | **READY (M) — RE-SCOPED 2026-08-03, and the direction REVERSED** | — | **FILED AS THE OPPOSITE DEFECT AND THE MEASUREMENT OVERTURNED IT.** It was filed 2026-08-03 as *"the reserve re-alerts on every status refresh"* — too MANY alerts — from the controller-side observation that a second push followed 13 s after the first. That was the sending end. **Measured at the receiving end** (hub `notification_log` + `events`, read from a copy taken WITH its `-wal`, freshness confirmed by the newest row post-dating the session): **9 events received today → 2 operator e-mails sent.** **06:40:03** privatebin AND opengist both refused → **opengist e-mailed, privatebin's alert has no row at all**. **08:59:46/47** opengist AND privatebin both refused → **privatebin e-mailed, opengist's absent**. **08:59:59, 09:03:00, 09:07:06** → **no operator row whatsoever**, all inside the 1-hour cooldown opened at 08:59:47. So the controller pushing repeatedly is not the defect; the hub emitting at most one operator e-mail per customer per hour is, and the loser is silent. **This is R-97a's failure mode exactly, in a second event type.** That row's own comment states it: *"`felhom-pbs` failing at 09:00 would swallow `local` failing at 09:20 for the whole hour"*. `cooldownTierSuffix` was written NARROW on purpose — empty unless the producer sends a `tier` — so no existing type's behaviour changed; `recovery_unit_capture_failed` simply never opted in. **CORRECTION OWED, and it is the reason this was worth measuring:** the 2026-08-03 R-181 report said *"one `recovery_unit_capture_failed` per app, HTTP 200"*. That was **true of what the CONTROLLER pushed** and would be read as *the operator was told about each app* — which is **false**. The distinction between an accepted event and a sent e-mail is the whole of this row. **Fix shape (NOT implemented — Part 0 was investigation only, by instruction):** let the producer opt into a per-app cooldown key, the way R-97a let the whole-guest producer opt into a per-tier one — the narrow mechanism already exists and needs no widening. **And a suppressed operator alert should leave a `skipped` row rather than nothing**, or this class stays undiagnosable from the hub's own records | CC | | **R-181** | **The capture floor guards the cheap leg and not the leg that fills the volume — and its refusal message asserts an invariant the code does not provide.** B2 (controller v0.192.0) is recorded on R-165 as the deliberate replacement for the bulkhead the `mp1` partition used to give. It is consulted in exactly one place — `m.unitFloorBlocked(stack.Name)` at `recovery_unit.go:328`, inside `captureAllRecoveryUnits`, which writes a manifest and a compose copy: **a few KB.** The leg that writes the bulk, `runVolumeDumps` (`backup.go:535`), has **no floor check at all** — its gates are protected-stack, volume-less, disconnected, decommissioned — and it runs FIRST, by design (*"MUST run before captureAllRecoveryUnits so the manifests enumerate the fresh tars"*, `backup.go:483`). So the write that fills the filesystem is unguarded, and the floor then refuses the write that would have cost almost nothing. **Second limb: the refusal message is false.** `recovery_unit.go:331` prints *"the previous unit is untouched and NOTHING was deleted"*. Nothing was deleted — true. Untouched — **measured false**: privatebin's `volume-dumps/privatebin_privatebin_data.tar` went `26c546c2…` → `b538ab89…` and opengist's went **182,272 B → 2,147,666,432 B**, both rewritten by the earlier leg, while each unit's `manifest.json` kept `created_at: 2026-08-03T06:34:26Z` and its `checksums` block covers only the three compose files — so a unit's payload can be swapped under a stale descriptor and **nothing in the unit can detect it** | **CLOSED — SHIPPED** (controller v0.193.0 + v0.193.1, 2026-08-03) | unblocks **R-165** | **FIRST LIVE FIRING OF B2, and it is why the runbook asked for one.** Proven on demo-hp 2026-08-03 06:40:03 on a box reinstalled from the merged golden (R-178). Method: a real 2 GiB file in opengist's data volume, then `fallocate` to bring the filesystem to 96 % used / 3.0 GiB free — both floor terms deliberately still clear, so the run started. **The `fallocate` instrument was proven before use** (5 GiB moved guest `df` 977M→6.0G while thin-pool `data_percent` stayed 29.03 → 29.03: zero blocks allocated), because demo-hp's thin pool is 53.93 GiB and a real fill to 97 % of a 70 G volume would have exhausted it and corrupted every guest on the box including the `drill-r50` fixture. Sequence observed: opengist's volume dump wrote **2.0 GB unguarded** → free fell to 1.0 GB → **both** apps' recovery-unit captures were then REFUSED on the `1.0 GiB free` term, each pushing `recovery_unit_capture_failed` (severity `error`) to the hub, accepted HTTP 200. **What DOES hold: it refuses per app rather than aborting the run, it never deletes, and the alert reaches the operator.** **Fix shape, not written this session by design (§7 of the runbook):** the floor belongs before the write in `runVolumeDumps` too, the message must stop claiming what the earlier leg has already falsified, and per `CLAUDE.md` *"a comment asserting an invariant needs a test pinning it"* the pinning test must assert the **consequence** (after a refusal, is the previous unit's payload byte-identical?) and not the mechanism. **Class:** the sixth entry in `CLAUDE.md`'s own table of shipped guarantees the code did not provide — found, as four of those were, only on live hardware. **CLOSED 2026-08-03 — controller v0.193.0 (`fef07c3`) + v0.193.1 (`6c43bf6`), proven live on demo-hp.** **The fix is ONE admission verdict per app per run** (`internal/backup/admission.go`), taken before that app's FIRST write and consulted by all three legs — the three write under one per-app root (`appbackup.RecoveryUnitPath`), which is exactly why one verdict can honestly cover them. **Decided LAZILY at the app's first write, never once at run start**: app A's dump can put app B under the reserve, so a run-start verdict would read a disk that no longer exists — the same class of mistake one level up. **Never re-decided between an app's own legs** (that IS the split this closes) and **reset per run**. Placed ahead of `DumpAppVolumesSafe`, which stops the stack as its first act, so a refused app is never bounced; placed AFTER the volume-less check, which has no write to gate. Exactly ONE operator alert per refused app per run. Leg order unchanged. **The floor is now SIZE-AWARE**, which is the second half of the defect: it asks whether THIS app's write would cross the reserve, not only whether the filesystem is already below it — the term whose absence admitted an app at 96% and then let it write 2 GB. Estimate = the app's previous `.sql`+`.tar` on disk; **no history → headroom-only** deliberately, or the first backup becomes the one that can never happen, and the alert says so. **A container-based `du` was MEASURED and rejected, not assumed**: 66 timed runs on demo-hp guest 9201, **median ~355 ms/volume (341–404)** on volumes holding tens of KB — the cost is container start-up, not the walk. Decisive on top: `docker run` needs the writable layer, so the instrument can fail under exactly the pressure the reserve exists to handle; and the previous-dump estimate measures the ARTIFACT that will be written rather than the live volume. **THE MESSAGE WAS NOT WEAKENED — the behaviour moved so the wording became true**, and it is checked by sha256 tree fingerprint, not by reading the log line (which is what lied). **LIVE PROOF, demo-hp guest 9201, the same method that found it.** The `fallocate` instrument was RE-PROVEN on the rebuilt box before use (guest `df` 1.2G→6.2G on a 5 GiB step while thin-pool `data_percent` stayed **36.83 → 36.83**: zero blocks allocated), because a real fill of a 70 G volume would exhaust the 53.93 GiB pool. **Headroom term @ 08:59:46** — 906 MB free / 99%: both apps refused, **`TREE_SHA` 111d1760c18d3440f700634ab325f8b8 IDENTICAL before and after** (10 files, incl. opengist's tar still at 182,272 B — R-181's own 'before' figure), **no `Stopping for safe volume dump` line at all** (it is present in the 08:58 baseline, which is what makes its absence evidence), 0 volume dumps, one `recovery_unit_capture_failed` per app HTTP 200. **Freed and re-run @ 09:01:33** — both captured normally. **SIZE term proven separately @ 09:03:00**, reproducing the original sequence with a real 2 GiB file in opengist's volume (its previous tar then **2,147,666,432 B**, the exact live figure) and the filesystem at **91% used / 2.9 GB free — both headroom terms deliberately clear**: opengist refused `(size)` — *"this app's last backup was 2.0 GB and writing it again would cross the reserve"* — while **privatebin was ADMITTED and dumped normally**, proving the term is per-app and not a global halt. **Teardown complete**: fill removed, planted file removed, `pct fstrim 9201` returned 67.5 GiB, pool **29.43%** (below the 36.83% baseline), tree byte-identical to the pre-test fingerprint. **v0.193.1 shipped in the same session**, found by this very proof run: the estimate was rendered fixed to 2-decimal GiB, so opengist's real **178 KB** printed as `estimated 0.00 GiB write` — which reads as *no estimate was available* and is the opposite of what happened. Rendering moved to `humanizeBytes`; arithmetic unchanged. Re-verified live: `estimated 178.0 KB write`. **11 new tests + 4 red-proofs**, each demonstrated failing then restored: both dump-leg gates removed (= v0.192.0) → Scenario A red with the tree shown changing; the size term removed → Scenario D red; a prune injected into the refusal path → Scenario F red; the floor moved above the warning band → Scenario G red. **Recorded honestly: the specified Scenario-F mutation (remove the reserve entirely) did NOT turn F red** — removing it makes every app write, which overwrites and adds but deletes nothing, so a deletion-watching test correctly stays green; the prune mutation is the one that proves the assertion. The DB leg cannot run without Docker, so its gate is pinned by an **AST walk** of `backup.go` asserting `admitApp` precedes `DumpOne` — `strings.Contains` is insufficient, a commented-out call still contains the string. **§3's correction CONFIRMED in passing and not chased**: `restore_points.go:57-59` takes the manifest mtime then `newestArtifact` over `.sql` and `.tar`, so the newest of the three wins — the restore point does NOT show a stale timestamp. **New finding from the live run → R-182.** | — | | **R-180** | **`--archive-storage` is accepted without checking the agent's token will ever be granted on it, and the failure lands at step 8/8 — after the root@pam password has already been rotated.** `felhom-host-install.sh` validates the archive storage EXISTS (`pvesm status --storage`, `:1583`) and that the golden volid RESOLVES on it (`:1661`), both in pre-flight. It never checks that storage against the ACL set it is about to grant, which is the fixed default `local local-lvm felhom-pbs` (`--acl-storages`, which `runbooks/day0-install.md` tells the operator **not** to pass). A storage outside that set therefore passes every pre-flight gate and dies at the last step | **READY (S) — NEW 2026-08-03** | — | **Hit live on demo-hp 2026-08-03** during R-178 Phase A, self-inflicted and therefore a clean demonstration: the golden was staged on `felhom-backup` (the enrolled NVMe, where the box's vzdumps live) and `--archive-storage felhom-backup` passed. Pre-flight passed; steps 1–7 ran; step 8 returned `reconcile: bring-up restore: proxmox: POST /nodes/felhom-host/lxc -> HTTP 403: permission denied at /storage/felhom-backup (missing privilege Datastore.AllocateSpace)`. **The cost is the ORDER, not the error** — by the time it fires, step 2 has minted the PVE token, step 4b has **rotated root@pam and vaulted it** (so the old console password is already dead), and step 5 has installed the agent. Recovery was `--resume` after moving the golden to `local`, which worked cleanly. **This is statically checkable in pre-flight**: `ARCHIVE_STORAGE ∈ PVE_STORAGES` is a one-line assertion over two variables both known at `:1583`. Same class as R-29 — the checkable thing that nothing checks | CC | | **R-179** | **`--uninstall` leaves the NAS network-storage systemd units behind, with the automount in `failed` state and the parent bind still mounted.** The teardown's residue-diff provenance (`day0-install.md` Part E: *"a full-filesystem diff against the pre-install baseline showed zero Felhom-named leftovers"*) is from **v1.9.1**, which predates the NAS network-storage feature. A box that has ever had a network share configured keeps `/etc/systemd/system/mnt-felhom\x2ddrives-.mount` and `.automount` after a full uninstall | **READY (S) — NEW 2026-08-03** | — | **Observed on demo-hp 2026-08-03** after `--uninstall --vmid 9201`: `mnt-felhom\x2ddrives-Felhom\x2dShare.automount` **loaded failed failed**, its `.mount` `loaded inactive dead`, and `mnt-felhom\x2ddrives.mount` still `active mounted` — the uninstall's own output had warned `/mnt/felhom-drives/Felhom-Share is busy — NOT forcing` and `/mnt/felhom-drives root bind left mounted`, which is correct behaviour (it never forces an unmount) but is not teardown. Cleared by hand before the reinstall: stop both units, remove both unit files, `daemon-reload`, unmount the autofs then the parent. **NEGATIVE CONTROL, same day:** demo-felhom's uninstall left **nothing** (`ls /etc/systemd/system | grep -i felhom` → only the unrelated `felhom-bootstrap.service`; no felhom mounts) — because that box had no network share configured. **So the residue is conditional on the feature having been used, which is exactly why a diff taken on a box that never used it reported clean.** `felhom-bootstrap.service` is NOT residue — it is the ISO first-boot unit, `disabled`+`inactive`, exactly-once and already fired | CC | diff --git a/documentation/backlog/ROADMAP.md b/documentation/backlog/ROADMAP.md index 9631888e..799c979e 100644 --- a/documentation/backlog/ROADMAP.md +++ b/documentation/backlog/ROADMAP.md @@ -20,7 +20,7 @@ | ID | Item | Size | Status | Notes / map rows flipped | |----|------|------|--------|--------------------------| -| R-115 | **Publishing is a remembered step — forgotten within eight hours of being documented as forgettable** | M | **READY** — operator ruling 2026-08-03: **mechanism (b), a build-side gate that REFUSES to deploy or vouch an unpublished version.** THIRD instance the same day (agent v0.120.0 deployed to both boxes while unpublished; a documented-path reinstall would have silently downgraded them and *succeeded*). CC's to build | A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published, so **"deployed" and "installable" are independent states that drift silently**. **Instance 1 — R-111** (morning): 17 agent releases v0.97.0–v0.113.0 stranded; a new customer would have installed without the whole R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT. Found only because the E-2d Phase 0 gate happened to look. **Instance 2 — agent 0.114.0** (same afternoon): the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, never published — which blocked Session C, since a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix. **The finding is the RECURRENCE, not either instance** — both are fixed. R-111 named this leg in its own text (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; it recurred the same day, which is the evidence that **a note is not a mechanism**. **Class: → R-29, one layer up** (a control that exists and is never walked) — deliberately NOT given a second ID. **Filed as its own item rather than reopening R-111** because R-111's finding (the channel *was* stale) is closed and verified end-to-end by the E-2d install, while the process defect that caused it is a distinct problem with a distinct fix and a distinct owner. **Operator's decision, mechanisms first:** (a) publish as a step in the build/release path so deployed and installable cannot diverge; (b) a gate that refuses to deploy an unpublished+unvouched version — strongest, fails closed; (c) a session-end checklist entry; (d) accept manual + a pre-Session-C verification. **(a)/(b) are mechanisms, (c)/(d) are reminders — and R-29's whole finding is that reminders do not hold.** No code written when filed, by design | +| R-115 | ~~**Publishing is a remembered step — forgotten within eight hours of being documented as forgettable**~~ | M | **CLOSED — SHIPPED 2026-08-03** (`release-agent.sh` + `check-published-versions.py`, no version bump). Releasing now builds, tags, publishes and **verifies by an independent download** in one act; a `v` tag with no package fails the gate, and **CI runs the full gate set** so it actually runs. Red-proof measured on real CI: same commit, green before a tagged-unpublished version existed, red after. Residue → **R-184** | A box installs the agent from a Gitea generic package the hub explicitly vouches, never from git. Nothing in the build, deploy or session-end path publishes or checks that a version was published, so **"deployed" and "installable" are independent states that drift silently**. **Instance 1 — R-111** (morning): 17 agent releases v0.97.0–v0.113.0 stranded; a new customer would have installed without the whole R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT. Found only because the E-2d Phase 0 gate happened to look. **Instance 2 — agent 0.114.0** (same afternoon): the R-113 fix, built and pushed at `b58d7bc`, deployed to felhom-pve, never published — which blocked Session C, since a fresh drill box would have installed 0.113.0 and proven the bug rather than the fix. **The finding is the RECURRENCE, not either instance** — both are fixed. R-111 named this leg in its own text (*"decide whether publishing joins the release train rather than staying a remembered step"*) and closed SHIPPED without it; it recurred the same day, which is the evidence that **a note is not a mechanism**. **Class: → R-29, one layer up** (a control that exists and is never walked) — deliberately NOT given a second ID. **Filed as its own item rather than reopening R-111** because R-111's finding (the channel *was* stale) is closed and verified end-to-end by the E-2d install, while the process defect that caused it is a distinct problem with a distinct fix and a distinct owner. **Operator's decision, mechanisms first:** (a) publish as a step in the build/release path so deployed and installable cannot diverge; (b) a gate that refuses to deploy an unpublished+unvouched version — strongest, fails closed; (c) a session-end checklist entry; (d) accept manual + a pre-Session-C verification. **(a)/(b) are mechanisms, (c)/(d) are reminders — and R-29's whole finding is that reminders do not hold.** No code written when filed, by design | | R-116 | **The drive-absent alarm and its recovery are a mismatched pair — generic on the way out, specific on the way back** | S | idea — **PROVEN LIVE 2026-07-29** | Absent fires `storage_disconnected`; return fires `backup_target_restored`. `backup_target_absent` never fires at all (count 0 across a full Session-C run), so an operator gets an alarm they cannot match to its recovery — exactly what `notifyDriveReturned`'s own comment forbids. Root cause: `notifyDriveAbsent` (`intermediary.go:635-646`) branches on `isTarget[a.Path]` with `a.Path` the GUEST path, and `driveTargetByPath` (`:602-616`) builds it as `out[GuestPath] = d.BackupTarget` — but **the drive is TWO `/disks` rows and the flag and the guest path sit on different ones**: the `felhom-backup` storage row has `BackupTarget: true` (`felhom-agent/internal/localapi/disks.go:211`) and gets a guest path only while classified user-data, while the registry union row has the guest path and **never assigns `BackupTarget`** (`disks.go:265-267`). Absent ⇒ the flagged row loses its guest path ⇒ the union row writes `false` ⇒ generic. On return the rows rejoin ⇒ specific. v0.184.1 fixed the KEYING, not this. **Only reachable because R-113 made the gate fire at all.** Fix likely agent-side; decide the repo first. Blocks E-2's C5. Evidence: `audits/SESSION-C-2026-07-29.md` §5 | | R-113 | **The drive-absent gate cannot fire on device loss — E-2b's alarm is wired to an unreachable condition** | M | idea — **PROVEN LIVE 2026-07-29** | `planDriveGates` (`felhom-controller/internal/web/intermediary.go:216-262`) treats a path as present by OR-ing in `d.BoundUnderParent`, which the agent derives from `GuestSeesMount()` — *"is this path a mount target in the guest's `/proc//mountinfo`"* (`internal/localapi/disks.go:210`). The raw drive mount is a **device-bound systemd unit** and dies with the device; **the agent's own bind under the shared parent is not device-bound and its mountinfo entry outlives the device**, so the gate reads it as present and `notifyDriveAbsent` is never called. Live on a fresh box: target drive hot-detached, agent said `enrolled drive absent by UUID` every 20 s for 4½ min, controller logged **0** `[gate]` lines, hub received **zero** events — neither `backup_target_absent` nor the generic `storage_disconnected`. Not a virtualisation artefact (device-bound-mount vs manual-bind is the same on metal); caveat: SCSI hot-detach, physical unplug not staged. **Sixth instance of seam-built-but-never-wired — E-2b wired the seam to a condition that cannot occur.** Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.2 | | R-112 | **E-2's degraded banner and offer have no UI consumer — correct endpoint, invisible to the customer** | S | idea — **PROVEN LIVE 2026-07-29** | `GET /api/storage/backup-target` returns byte-exact Hungarian copy (verified on a live box), and nothing in the product asks for it: `grep 'backup-target'` across every `*.html`/`*.js`/`*.css` → **0 hits**; no template references `OfferPath`/`Degraded`/the copy; `resolveBackupTargetState` and `degradedMessageFor` are consumed **only** by the JSON handler, with **no page handler injecting the state**. Decisive contrast: the templates fetch **18 distinct `/api/storage/*` endpoints** — `backup-target` and `backup-target/assign` are the only two with zero references. The handler's own comment calls itself *"the dashboard's source for the degraded banner and the offer"*. v0.185.1 shipped as *"the offer endpoints were mounted where nothing routed to them"* and fixed the **mount**, stopping one layer short of the **render**; its test pins dispatch, not reachability. **Fifth instance of the class. Fix R-114 first** — wiring this alone starts showing customers a wrong message. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §5.1 | @@ -157,7 +157,7 @@ Self-resolves the moment the target answers (the storage read succeeds, sees the | R-97 | **The whole-guest backup tier has NO failure signal to the hub — `internal/quiesce` never notifies** | S | **SHIPPED (controller v0.177.0 + hub v0.78.0, 2026-07-27)** — **R-97a:** `quiesce.TierNotifier`, a seam (not an import) wired by an init-only setter, edge-triggered on the R-88 breaker ARMING so a failing tier is reported once per run rather than once per retry; recovery rides `recordSuccess`'s existing bool. **NEW operator-only event types** `whole_guest_backup_failed`/`_recovered` — deliberately NOT `backup_failed`, which carries a customer Hungarian template AND sits in demo-felhom's live `enabled_events`, so reusing it would have emailed the CUSTOMER about a backup they cannot act on while it was still retrying. The recovery joins `recoveredPairedDownTypes` because its `info` severity would otherwise be dropped by `severityNotifies` — the operator would hear it break and never hear it heal. **The hub's operator cooldown was keyed `customerID:eventType` alone**, so one tier would have masked the other for an hour; now narrowly extended with a `tier` suffix taken from the event details, leaving every other event type unchanged. **R-97b:** a suppression window keyed to the quiesce CYCLE (not a state test — v0.164.0's `!= StateStopped` filter cannot see an app caught MID-RESTART, which is exactly how BookStack alarmed), consumed at the same single derivation point `classifyRunStates`. Grace = **180 s**, derived from the deploy flow's 120 s health timeout and Mealie's 60 s `start_period`; it **expires**, so an app that genuinely fails to come back still alarms. **PROVEN LIVE end-to-end with a control:** the new type POSTs 200 from inside guest 9201 while a bogus type 400s, and `notification_log` shows **1 operator row, 0 customer rows**. The quiesce→notify link itself is unit-proven only. | On 2026-07-27 three whole-guest backups failed and three quiesce cycles stopped and restarted every customer app stack, and **not one `backup_failed` event reached the hub.** It is not the allowlist — the hub already carries `backup_failed` and `backup_completed` (they are emitted by the controller's *app-data* backup path). The cause is that **`internal/quiesce` does not import `internal/notify` at all**: the tier R-82 built has no route to the hub, so a whole-guest backup can fail indefinitely in silence. The loop's only trace was `app_start_failed` — **info** severity, **Hungarian**, on the **customer** channel — telling the customer BookStack was down (it had been caught mid-restart by the third cycle) without saying why, during an outage the system itself caused. So the one signal that did fire was both the wrong tier and the wrong story. **Shape:** emit `backup_failed`/`backup_completed` from `quiesceAndPollTiers` naming the TIER, operator-tier; and decide whether a quiesce-induced restart should suppress `app_start_failed` the way controller **v0.164.0**'s deliberate-stop filter does — an app the backup stopped on purpose is not a fault. R-88's breaker bounds the repetition but changes nothing about the silence | | R-95 | **The restic offsite tier's credential CAN DELETE — R-89's "parallel question", now ANSWERED** | M | idea — established read-only 2026-07-27 | **The exposure closed on the weekly PBS tier is fully open on the daily restic tier**, which holds the customer's actual documents and photos and is the only tier that survives losing the box. Established without mutating anything: **(1) Identity** — a per-customer *subaccount* on `storage-box-pool-1` (box 611714, bx11, `u629488`): `u629488-sub1` home `felhom-demo-felhom`, `sub2` peti-felhom, `sub3` demo-hp, each labelled `felhom-customer`. Auth is an **SSH key stored ON THE BOX** (`…/felhom-controller-data/_data/data/offbox/ssh_key`, 0600, beside `repo_password` + a pinned `known_hosts`) — customer-side, not hub-side, so a compromised guest holds it. **(2) Read-write: YES** — the API reports **`readonly=False` on all three subaccounts**, and it is not merely latent: the controller runs `restic forget --group-by host,tags --keep-daily 7 --keep-weekly … --prune` **from the box** (`backup/offbox.go:984`, also `:1070`). Delete rights are exercised on every run. **(3) Append-only: NO, and not expressible** — the repo is built as `sftp:` (`offbox.go:482`); restic's append-only mode requires the **REST server** backend, which plain SFTP cannot provide. **(4) A zero-code mitigation exists and is unused:** the box type carries `snapshot_limit=10` and the API reports `snapshot_plan=null` with **0 snapshots** and `size_snapshots=0`. Hetzner Storage Box snapshots are taken **server-side, outside the SFTP namespace** — an SFTP subaccount cannot delete them — so they are a genuine immutability layer at no extra cost and with no code change. **Rule once for both tiers, per R-89.** Options, cheapest first: enable a snapshot plan (operator click, immediate); split backup-write from prune so pruning runs somewhere the box cannot reach; or move the repo to restic's REST server with `--append-only`. Flips the capability-map row for offsite immutability | | R-94 | ~~A hand-synced version constant drifts, and the gate that would catch it is never run~~ | XS | **CLOSED — SHIPPED hub v0.87.0, 2026-08-02** | Closed by **deleting** the label rather than deriving it: the Setup command fetches the installer at run time from a 30 s-git-synced website (R-110), so no build-time value in the hub can be true. `hostinstall_gates.py` gate 1 inverted to pin the ABSENCE of a version literal; the tautological `render_test.go` assertion deleted (demonstrated passing at `9.9.9`). Detail: `OPEN-ITEMS.md` R-94 | -| R-110 | **`main` is the installer's publish channel — there is no staging** | S | **READY** — operator ruling 2026-08-03: **option (b), the channel moves to a TAG**, so publishing is moving the tag and rollback is moving it back. **Must cover BOTH channels** — the nginx `/scripts/` git-sync AND the nine files the installer fetches from `raw/branch/main` — or it only half-works. CC's to build | `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a `--period=30s`, and nginx serves that working tree directly (`location /scripts/`, `root /usr/share/nginx/html/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`, `:1262`; `runbooks/day0-install.md` C.1) receives. **There is no tag, no pinned-version path, no staging copy and no rollback other than another push** — for the one artifact that runs as **root on a virgin box**, the most privileged thing Felhom ships. **Two consequences worth stating plainly:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29, so the proof run confirms what customers already receive rather than clearing it for release; and the precaution the old R-94 row recorded ("do not point every new box at an installer that has never run") **was never available to take**, because nothing points boxes at a version. **Open question for the operator, not a defect to fix blind:** should `/scripts/` serve a pinned release — a tag-tracked git-sync ref, or a versioned directory (`/scripts/1.22.0/…`) with the hub's generated command naming a version — or is `main`-tracking the accepted shape for a one-operator product where the alternative is a release ritual nobody performs? **Exposure today is zero** (no boxes are installing), which is exactly why it is cheap to decide now. Whichever way it goes, it decides whether R-94 leg (a) makes the label a *fact* (derived from the served script) or keeps it a *claim*. Flips no capability-map row — the map states what the platform does, and this changes nothing about that | +| R-110 | ~~**`main` is the installer's publish channel — there is no staging**~~ | S | **CLOSED — SHIPPED 2026-08-03** (installer v1.23.0). `/scripts/` syncs `installer-v1.23.0`; the website still tracks `main`. Proven by HTTP: a push to `main` left the served bytes byte-identical, moving the tag published in ~40 s, moving it back restored the exact prior sha. The run-time fetches turned out to be **sixteen from the agent repo**, not nine from here — pinned to the agent version instead → **R-183** | `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a `--period=30s`, and nginx serves that working tree directly (`location /scripts/`, `root /usr/share/nginx/html/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`, `:1262`; `runbooks/day0-install.md` C.1) receives. **There is no tag, no pinned-version path, no staging copy and no rollback other than another push** — for the one artifact that runs as **root on a virgin box**, the most privileged thing Felhom ships. **Two consequences worth stating plainly:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29, so the proof run confirms what customers already receive rather than clearing it for release; and the precaution the old R-94 row recorded ("do not point every new box at an installer that has never run") **was never available to take**, because nothing points boxes at a version. **Open question for the operator, not a defect to fix blind:** should `/scripts/` serve a pinned release — a tag-tracked git-sync ref, or a versioned directory (`/scripts/1.22.0/…`) with the hub's generated command naming a version — or is `main`-tracking the accepted shape for a one-operator product where the alternative is a release ritual nobody performs? **Exposure today is zero** (no boxes are installing), which is exactly why it is cheap to decide now. Whichever way it goes, it decides whether R-94 leg (a) makes the label a *fact* (derived from the served script) or keeps it a *claim*. Flips no capability-map row — the map states what the platform does, and this changes nothing about that | | R-128 | **`ISO_VERSION` "aligns with SCRIPT_VERSION" was a comment nothing evaluated** | XS | **CLOSED — iso v1.26.0, 2026-07-31** | Closed by **correcting the claim, not asserting it**: the ISO is frozen while `felhom-host-install.sh` is fetched at run time from `main` (R-94/R-110), so an assertion would invent a constraint. `build-felhom-iso.sh:45-52`. Full reasoning in `OPEN-ITEMS.md` | | R-154 | **`[first-boot]` is automated-install-only, and nothing in the tree said so** | XS | **CLOSED — iso v1.26.0, 2026-07-31** | A PVE property, measured with a same-image control (`audits/SPIKE-universal-iso-3-2026-07-31.md` §2); recorded at `scripts/iso/pkg/build-deb.sh:6-11`. Superseded in practice by the `.deb` delivery route | | R-155 | **`iso-repack.sh` refused any ISO without `auto-installer-mode.toml`** | XS | **CLOSED — iso v1.26.0, 2026-07-31** | **Narrowed, not deleted** — unchanged for `FELHOM_MENU=single` (`iso-repack.sh:121-128`), does not apply to `release` where the file's absence *is* gate G1. Do not remove it wholesale |