From 952ebf486249bf4fee815693f07cce8057b221e9 Mon Sep 17 00:00:00 2001 From: kisfenyo Date: Wed, 29 Jul 2026 23:34:06 +0200 Subject: [PATCH] Record work, banked first: shrink the E-2d row, create the missing capability-map rows MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Unconditional and three sessions overdue, so it commits before any code is touched — E-2d itself stopped at Phase 0 and banked nothing. E-2d row: 822 words -> 121, and the contradiction resolved. Its State read CLOSED — PARTIALLY PROVEN while the cell's final sentence read "This row stays OPEN only for the residue"; a reader could not tell which. It is CLOSED, with R-116 the single named open leg. Nothing unique was binned. Three facts existed ONLY in that cell and are moved into audits/E2D-fresh-vm-2026-07-29.md as a new §1a: the local-lvm fence figures with the 888 GB nvme alternative, the exactMount subdirectory caveat and why the subdirectory is nonetheless the safe placement (no durable_id collision), and the ISO/PAIRING -> DIRECT fall-through derived at source with its line citations. drill-r50's blocked status was already in both audits. Capability map: it had ZERO rows for the backup-target work — grep gives 0 hits for backup_target and one for "E-2" that is a campaign date string. Three scenario rows added, at today's honest status, not the value hoped for later: C. Protection & recovery — installer Case A/B, DEGRADED recorded not hidden PROVEN-LIVE, cites E2D-fresh-vm C1+C2 D. Storage & devices — the offer, and that registration confers no role PROVEN-LIVE, cites SESSION-C C4 + the decline path F. Notifications & monitoring — the absent-target alarm and its pairing PARTIAL, cites SESSION-C C5, leg named, -> R-116 Row F is PARTIAL today per the doc's own strict enum (a leg not exercised live is PARTIAL with the leg named, never PROVEN-LIVE). A later session may flip it; this commit must not. --- .../architecture/00-capability-map.md | 3 ++ .../audits/E2D-fresh-vm-2026-07-29.md | 30 +++++++++++++++++++ documentation/backlog/OPEN-ITEMS.md | 2 +- 3 files changed, 34 insertions(+), 1 deletion(-) diff --git a/documentation/architecture/00-capability-map.md b/documentation/architecture/00-capability-map.md index 2a48ba9..792db2e 100644 --- a/documentation/architecture/00-capability-map.md +++ b/documentation/architecture/00-capability-map.md @@ -72,6 +72,7 @@ | Scenario | Components | Status | Evidence | Gap / roadmap | |---|---|---|---|---| +| Whole-guest backup lands OFF the guest's own physical device; a single-drive box is recorded **DEGRADED** rather than silently normal (installer Case A/B) | agent v0.113, host-install v1.22.0 | **PROVEN-LIVE** | `E2D-fresh-vm-2026-07-29` C1 (real 1.22.0 install, rc=0, `Day-0 provision SUCCESS`) + C2 (both DEGRADED lines verbatim, `local_backup_target=local`, install did not abort) | Case A (a second drive already present at install) has never fired naturally — only Case B has | | Nightly DB dumps (postgres/mariadb autodiscovery), atomic writes | controller v0.118 | **PROVEN-LIVE** | `CAMPAIGN-2` T-BAK-FULL (pg+mariadb autodiscovered); atomicity `CAMPAIGN-6B` P4 + `CAMPAIGN-6E` B1/B2 (SIGKILL mid-write → only `.tar.tmp` touched, last-good byte-unchanged); DB restore `CAMPAIGN-6D` P-FAB | (Cited `CAMPAIGN-3` F7 is the *finding* of non-atomic writes, and T-RST-DB was auth-hollow — corrected to the 6B/6E fix-proofs.) T-6E-1 dir-fsync asymmetry (LOW) → R-10 **DB replay route → `07-backup-architecture.md` §8 row 3** | | Tier-2 secondary-drive copy: class-driven legs, v2 relpath layout, NAS-target exclusion, safe-remove boundary | controller v0.135 | **PROVEN-LIVE** | `CAMPAIGN-6E-2026-07-15` (P-TIER2 deep-4 PASS), `CAMPAIGN-6C` | **Route + RTO → `07-backup-architecture.md` §8 rows 1, 2, 4, 5.** The matrix records that the copy's `recovery-unit/` mirror is read by no path (→ R-102) | | Offsite (restic → Hetzner Storage Box): mandatory class only, raw-data quota, enlargement gate, retention regrouping | controller v0.134, agent, hub | **PROVEN-LIVE** | `CAMPAIGN-6D-2026-07-15` (mandatory-only P-IMMICH; enlargement gate fired at real 50GiB quota P3-DELIVERY); `VALIDATION-offbox-storagebox-2026-07-09` (byte-perfect round-trip) | Raw-data quota (SP-1) + retention regrouping (SP-2) are `SPIKE-restic-snapshot-shape` **dry-run** verdicts — mechanism validated, not fired in a live product run; only the enlargement gate is live-fired. **Reinstall-continuity (controller v0.142.0, 2026-07-17):** a recreated data volume that orphaned the repo (new passphrase can't open the old keys) is now CLASSIFIED (`wrong password or no key found`) → explicit ORPHANED card + event (not nightly-spam) + a move-aside (never-delete) reset (unclaimed auto / claimed confirm), instead of a raw nightly restic error. Fake-based scenarios + red-proofs. **The live leg FIRED on its own during the 2026-07-18 rehearsal** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`, S7): after a RESET + re-enable, the first offsite run hit the previous lifecycle's ciphertext and the guard **classified it, pushed `offbox_repo_orphaned`, skipped the run and showed the card (16:58:14)** rather than nightly-spamming a raw restic error; the operator-confirmed reset then **moved the repo aside (never deleted) to `.orphaned-20260718` and re-initialised (16:59:26→16:59:32)**, and the next run produced 2 snapshots / 48.717 MiB. The guard behaved exactly as designed — **the finding is that it had to fire at all** (R-32: RESET destroys custody, so the ciphertext it leaves behind is dead by design and should be purged, while the move-aside guard stays correct for reinstall-WITHOUT-RESET). `DIAGNOSE-offbox-repo-orphaned-2026-07-17` **Route + RTO → `07-backup-architecture.md` §8 rows 4, 10, 12, 15** (incl. the R-95 delete exposure and the R-104 stale-lock defect) | @@ -90,6 +91,7 @@ | Scenario | Components | Status | Evidence | Gap / roadmap | |---|---|---|---|---| +| A second drive appearing is OFFERED as the backup target; accepting moves it; registration and the drive gate confer no role by themselves | controller v0.186.0, agent v0.113 | **PROVEN-LIVE** | `SESSION-C-2026-07-29` C4: offer rendered with `data-path`, decline path proven (target stayed `local`, no `felhom-backup` storage, agent.json unchanged), `restart_required:true`, agent did NOT self-restart, wrapper created the storage at the drive's OWN mountpoint | Accept was driven through the endpoint the button POSTs, not a browser click — no browser automation exists on DooPlex | | Drive wizard: scan/format/mount/enroll, incl. legacy-boot LVM-root hosts | controller, agent v0.87 | **PROVEN-LIVE** | `DISPOSITION-ia-finding2-systemdisks-2026-07-13` (legacy EFI+LVM host, root not offered, byte-identical); enroll/format live in `storage-lifecycle-acceptance-2026-06-15` (E10 re-enroll, data intact); agent fence self-test refuses `/dev/sda` | (Cited `CAMPAIGN-2` T-STG-ENROLL/SEC-FORMAT were auth-hollow CSRF-403.) **Fresh-USB wizard enroll+format through the customer UI PROVEN-LIVE (controller v0.141.0, 2026-07-17):** a 64 GB scratch USB driven through the real `/api/storage/init` endpoints (login+CSRF) → confirm → detached format (~27 s mkfs) → mount → register → mounted+registered at `/mnt/felhom-drives/scratch1`. **F6 (initialize-to-usable) now covered:** the wizard runs the chain as a detached, disconnect-safe, pollable job (3-step progress) with an agent format-status poll for a slow mkfs **2026-07-26 — a SILENT failure class on the channel every agent-backed capability depends on (this row, data migration, USB enrollment, guest RAM, quiesce/PBS) is now DETECTED (controller v0.173.0, R-77). No row status flips.** `controller.yaml` and `bootstrap.json` could disagree on `local_api.endpoint` indefinitely with no signal: the R-50 island migration rewrote the latter, the fleet kept dialling the former, and for 17.5 h the only alert was a generic "agent unreachable" that read as an infrastructure blip. Drift now raises its own event type (`local_api_endpoint_drift`) naming both values. It is DETECTION ONLY — the authority ruling is R-78 — so the class is now loud, not prevented. Evidence: `audits/DIAG-agent-channel-2026-07-26.md`. | | Data migration between drives (all / per-app), crash-safe | controller | **PROVEN-LIVE** | `CAMPAIGN-6C` 4P-5 (scope=app round-trip, byte-identical); `storage-lifecycle-acceptance-2026-06-15` (two migrate-all runs via dashboard UI, sha256 byte-identical) | (Cited `CAMPAIGN-2` T-STG-MIGRATE-* were auth-hollow.) "crash-safe" is design-level (copy→verify→remove) — no clean live crash-during-migration PASS | | NAS (NFS/SMB-client) verify-before-commit, uid-1000 probe, categorized Hungarian errors, DSM-validated | controller v0.113–117, agent v0.81/84/85 | **PROVEN-LIVE** | `SPIKE-nas-verify-2026-07-11`, `SPIKE-nas-dsm-2026-07-11`, `CAMPAIGN-3-2026-07-11` (boot/reassert fixes) | | @@ -119,6 +121,7 @@ | Scenario | Components | Status | Evidence | Gap / roadmap | |---|---|---|---|---| +| An ABSENT backup-target drive raises its OWN alarm, paired with a matching recovery | agent v0.114.0, controller v0.184.1+, hub v0.81.0 | **PARTIAL** | `SESSION-C-2026-07-29` C5: the drive-absent gate fires (4 s) and an alarm reaches the hub — but it is the **generic** `storage_disconnected`, while the return fires the **specific** `backup_target_restored`, so the pair cannot be matched. `backup_target_absent` never fired (count 0) | The specific alarm and its severity, Hungarian copy and hub routing are all still unexercised end-to-end → **R-116** | | Health-degradation email (edge-triggered, cooldowns, Hungarian) via hub → Resend | controller, hub | **IMPLEMENTED** | delivery pipeline live-proven for the **enlarge-block** trigger (`CAMPAIGN-6D` P3-DELIVERY, op+customer "Kedves Ügyfél!"); `NotifyHealthChange` ok→warn/fail edge-trigger implemented | The **health-degradation** trigger specifically has never fired an email live in any doc. Demoted (pipeline proven for a different event). Deliverability to HU freemail → R-4 | | Event catalog: app_start_failed, dead-app, offbox_enlarge_blocked, claim/reset codes, critical severity | controller, hub v0.31/48/50/55 | **PROVEN-LIVE** | live-delivered: `CAMPAIGN-6D` P3-DELIVERY (enlarge-block, op+customer); `DRILL-day0-vm` F-4 (claim code); `DRILL-day0-take2` F-15 (reset code) | `app_start_failed`/`dead-app` delivery is unit-only (6C inconclusive) — the pipeline + 3 event families are live, those two are not | | Prefs safety: empty-email wipe guard | controller v0.137 + hub v0.71.0 | **IMPLEMENTED** | controller leg red-proofed 07-15; hub-side no-clobber belt (`handleSavePreferences` preserves a stored non-empty address on an empty-email push) red-proofed 07-22 | Born from a live incident; controller 0.160.0 guards both its push legs, so the hub belt covers older/rogue boxes | diff --git a/documentation/audits/E2D-fresh-vm-2026-07-29.md b/documentation/audits/E2D-fresh-vm-2026-07-29.md index 37e2016..001574f 100644 --- a/documentation/audits/E2D-fresh-vm-2026-07-29.md +++ b/documentation/audits/E2D-fresh-vm-2026-07-29.md @@ -37,6 +37,36 @@ manual-installer fallback was used. **Operator STOP: not required and now retired.** `HUB_PW` is in `~/.config/credentials`; CC created the customer and performed the bind itself. The one human step that *was* needed is new — see §6. +## 1a. Phase 0 answers, preserved from the OPEN-ITEMS row + +Moved here when the E-2d register row was rewritten (2026-07-29) — the row had grown to ~820 words and +these were the facts that existed nowhere else. They are inputs to any future drill on this host, not +narrative. + +- **Storage fence.** `local-lvm` on demo-hp is a thin pool, ~144 GB allocated against ~54 GB real, + 38.8 % used, on a box running a live customer guest — a full thin pool corrupts every guest on it. + `local` has only 23.7 GB and sits on `pve-root`. **Use `/mnt/nvme-1tb` (888 GB free).** +- **The `exactMount` caveat, and the placement decision it forces.** A dir storage created at a + SUBDIRECTORY of `/mnt/nvme-1tb` fails the agent's `exactMount` check and reports `disconnected` in the + host report. Both E-2d and Session C accepted that: hub-side it is a WARN log line only — no event, + no email — and the alternative (a second storage at the live backup target's own mountpoint) risks + perturbing the drive-role resolution on a production box. The agent deliberately falls back to a + stable store id rather than borrowing the nvme's fs-UUID in this case, so there is **no durable_id + collision** with `felhom-backup`; that is what makes the subdirectory the safe choice. +- **The ISO/PAIRING → DIRECT fall-through, derived at source.** A fresh VM with no baked customer-id + lands in PAIRING mode (`scripts/iso/felhom-bootstrap.sh:537-541`), not DIRECT (`:312`), and only + DIRECT passes `--customer-id / --mode / --passphrase-file`. On a 200 from `/api/v1/appliance/poll` + the pairing loop writes the hub-delivered credentials into the 0600 env, re-sources it and calls + `run_direct` **in the same invocation** (`:495-499`), which is the single site that fetches + `$INSTALL_URL` (`:322-330`), builds the args (`:334`) and invokes + `bash "$SCRIPT_TMP" "${args[@]}"` (`:343`). So the ISO route reaches the identical installer + invocation and yields a claimable customer — which is why it is the spine and no manual 1.22.0 run + is needed as a separate scenario. +- **`drill-r50` stays blocked.** Unblocking it means the fixture stops representing anything real + (R-93). + +--- + ## 2. Timeline (VM 9300 `e2d-fresh` on demo-hp, nested PVE) | UTC | Event | diff --git a/documentation/backlog/OPEN-ITEMS.md b/documentation/backlog/OPEN-ITEMS.md index 53d2634..0618a02 100644 --- a/documentation/backlog/OPEN-ITEMS.md +++ b/documentation/backlog/OPEN-ITEMS.md @@ -11,7 +11,7 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha |---|---|---|---|---|---| | **R-88a** | ~~Failing backup re-quiesces every 5 min, no backoff~~ | **SHIPPED** (controller v0.176.0, 2026-07-27) | — | Live on both boxes; breaker 15m→4h, per-tier, never permanent | — | | **R-88b** | ~~`/backup/due` cannot say *unknown*~~ | **SHIPPED + PROVEN-LIVE** (agent v0.105.0 + controller v0.178.0, 2026-07-27) | — | `age_state=unknown` captured on real hardware during a deliberate ep0 outage; controller deferred, **zero app stacks stopped** | — | -| **E-2d** | **Prove E-2 on a fresh VM on the t740** — the only remaining route to four unproven items: a real `felhom-host-install.sh` **1.22.0** run (never done), Case B naturally (single-drive install renders the degraded banner without degrading a live box), a **claimable** customer so the three claim-gated items stop being gated, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (Session C, 2026-07-29) | — | **CLOSED by `audits/SESSION-C-2026-07-29.md`.** C1/C2 proven in E-2d; **C3 and C4 PROVEN LIVE** this session (R-114, R-112); **C5 FAILED** — the gate fires and an alarm reaches the hub, but it is the generic event, not `backup_target_absent` (→ **R-116**, the one named open leg). Per the runbook's §9, decided in advance: a failed claim closes E-2 as partially proven with a named leg rather than re-running. The arc's stated definition of done is R-106+R-109, R-108 and D5 — none of which this detour touched. **Space checked 2026-07-29 — NOT a blocker, with one constraint: the VM disk must NOT go on `local-lvm`.** That thin pool is over-subscribed (144 GB allocated against 54 GB, 38.8% used) on a box running a live customer guest, and a full thin pool corrupts every guest on it. `local` has only 23.7 GB and sits on `pve-root`. **Use `/mnt/nvme-1tb` (888 GB free).** Caveat found: a dir storage at a SUBDIRECTORY there will fail the agent's `exactMount` check and report `disconnected` in the host report — cosmetic, but decide the placement deliberately. **Do NOT unblock drill-r50** (deliberately blocked; unblocking it means the fixture stops representing anything real). **CORRECTED 2026-07-29 — the ISO is not an obstacle to the 1.22.0 proof, it is the best route to it.** `felhom-bootstrap.sh:96` fetches from `https://felhom.eu/scripts/felhom-host-install.sh`, not the hub, and that URL serves **1.22.0** (git-sync from `main`, ≤30 s — live fetch confirmed). A fresh ISO install therefore exercises 1.22.0 **automatically** — which makes the ISO leg the *stronger* proof, because it is the real customer path rather than a proxy for it, and it retires the "run 1.22.0 manually" step. **The Phase 0 question that decides the shape is ANSWERED — it was a source read, and the answer is yes.** A fresh VM with no baked customer-id lands in **PAIRING** mode (`felhom-bootstrap.sh:537-541`), not **DIRECT** (`:312`), and only DIRECT passes `--customer-id / --mode / --passphrase-file`. But on a 200 from `/api/v1/appliance/poll` the pairing loop writes the hub-delivered `FELHOM_CUSTOMER_ID` + `FELHOM_RETRIEVAL_PASSPHRASE` into the 0600 env, re-sources it and calls `run_direct` **in the same invocation** (`:495-499`) — so pairing reaches the identical installer invocation (`:322-343` — the single `$INSTALL_URL` fetch at `:322-330`, the `--customer-id/--mode/--hub-url/--passphrase-file` args array at `:334`, and the `bash "$SCRIPT_TMP" "${args[@]}"` call itself at `:343`) and the customer it yields is the one the operator bound, i.e. claimable. **So the ISO leg is the spine**; a manual 1.22.0 run is not needed as a separate scenario. **RUN ATTEMPTED 2026-07-29 — STOPPED AT PHASE 0, no VM created, nothing touched (`audits/E2D-fresh-vm-2026-07-29.md`).** The blocker is **R-111**: a fresh box installs **agent 0.96.0 + controller 0.161.0**, not `main`'s 0.113.0/0.185.1, so **C3/C4/C5 test surfaces that do not exist on it** — the degraded banner + `GET /api/storage/backup-target` landed in controller **v0.185.1** (`cdaeb36`) with the copy in **v0.185.0** (`3f7cf2a`); `backup_target_absent` in **v0.184.0** (`c1a63de`); the offer's apply needs agent **v0.113.0** (`58b598b`). **C1 (a real rc=0 1.22.0 install) and C2 (Case B natural, `configure_backup_target()` `felhom-host-install.sh:627`, warnings `:653-655`) remain ACHIEVABLE TODAY** — both are installer-side and host-install is served at 1.22.0. **C3 unblocks cheaply** by raising the hub global floor 0.156.0 → ≥0.185.0 (controller 0.185.1 IS in the registry): measured fleet impact is **nil** — both demo boxes already run 0.185.1, and Peti is DOWN 14 d and already below the current floor. **C4/C5 need R-111 first.** Phase 0 answers are all recorded in the audit, so a resumed run does not re-derive them: drive-gate cadence **30 s** (`intermediary.go:337`, registered `server.go:217`) ⇒ a 60 s two-cycle budget; hot-detach available (`virtio-scsi-single` + default hotplug, VM 300 is the working reference); ISO present (`…v1.25.0-nested-vm-generic-mkimage.iso`); `/mnt/nvme-1tb` 888 G free and the `local-lvm` fence re-measured (38.77 %, unchanged). **The §5.1a operator STOP is retired** — `HUB_PW` is in `~/.config/credentials` and hub auth was verified, so CC can bind. **One open decision carried forward:** where the VM disk's dir storage goes, since `local-lvm` is forbidden and `felhom-backup` is the live backup target — see the audit §4.1. **RUN EXECUTED 2026-07-29 after R-111 was fixed — `audits/E2D-fresh-vm-2026-07-29.md`.** Full ISO/PAIRING route on a nested VM on demo-hp; bind → running controller in **3 m 35 s**; teardown clean (`pvesm status` after == before, `local-lvm` 38.77 %, guest 9201 untouched). **C1 PROVEN** (`felhom-host-install v1.22.0`, `Day-0 provision SUCCESS`, guest 9201 running, golden = the one baked 20 min earlier). **C2 PROVEN** (both DEGRADED lines verbatim, `local_backup_target=local`, install did not abort). **C3/C4 PARTIAL — API exact and complete, but customer-invisible → R-112.** **C5 FAILED → R-113.** C4's decline path PROVEN (registration confers no role), `restart_required:true` PROVEN, agent did **not** self-restart, E-2a wrapper created the storage at the drive's own mountpoint, healthy renders nothing. **E-2d's own premise needed amending:** a fresh install does NOT yield a CC-drivable claimable customer — the claim code is bcrypt-hashed and email-only, so one operator relay was required (and the claim flow is now proven end to end). **This row stays OPEN only for the residue:** C5 re-test after R-113, and the C3/C4 UI legs after R-112 | CC | +| **E-2d** | **Prove E-2 on a fresh VM** — a real `felhom-host-install.sh` 1.22.0 run, Case B naturally, a claimable customer, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **CLOSED — PARTIALLY PROVEN** (2026-07-29) | — | **C1, C2 proven** (`audits/E2D-fresh-vm-2026-07-29.md`); **C3, C4 proven live** (`audits/SESSION-C-2026-07-29.md`); **C5 FAILED → R-116** — the gate fires and an alarm reaches the hub, but it is the generic event, so the alarm and its recovery cannot be paired. **R-116 is the single named open leg**; per the Session-C runbook §9, decided in advance, a failed claim closes the item as partially proven rather than triggering a re-run. Both audits carry the full record — the `local-lvm` fence, the ISO/PAIRING derivation, the Phase 0 answers, the per-claim observables and the teardown evidence — and are the place to read it, not this cell. **The arc's actual definition of done is R-106 + R-109, R-108 and D5**, none of which this detour touched | CC | | **R-94** | **A hand-synced version constant drifts, and the gate that would catch it is never run** — `hub/internal/web/configs.go:28` pins `hostInstallVersion = "1.19.0"` while `scripts/felhom-host-install.sh:187` is `SCRIPT_VERSION="1.22.0"` | **READY (XS)** | — | **CORRECTED 2026-07-29 — the earlier framing of this row was false and is retracted.** The constant selects no script: its only consumers are `configs.go:487` (`ScriptVersion`) and `render_test.go:219`, and it renders as a label at `customer_unified.html:494`. The install command beneath that label fetches `https://felhom.eu/scripts/felhom-host-install.sh` (`customer_unified.html:563`, `:1262`), which the website git-syncs from `main` on a 30 s period (`manifests/webpage.yaml`) — so **1.22.0 is what every install already gets** (live fetch, 2026-07-29). Every flag the generator emits is parsed by 1.22.0 (`customer_unified.html`~`:1210`–`:1238` vs `felhom-host-install.sh:1177`–`:1210`): **no functional gap, only a wrong number on the operator's screen.** Three legs, all XS: **(a)** derive the label from `SCRIPT_VERSION` rather than hand-syncing it, or delete it; **(b)** `scripts/hostinstall_gates.py` **fails today** and is invoked by no Makefile, hook or `CLAUDE.md` — wire it next to `site_gates.py` or delete it, because a gate nobody runs reads as coverage it is not providing (**this leg is one instance of → R-29**, which is the class: gates are enforced nowhere, and the enforcement decision belongs there, not here); **(c)** `render_test.go:219` compares the constant to itself and passes at any value — replace it with the cross-file assertion. **No longer blocked on E-2d** — it never gated anything. **2026-07-29: a real 1.22.0 install has now happened** (`audits/E2D-fresh-vm-2026-07-29.md`), so even the original (retracted) precaution is discharged — nothing stands in front of this row | CC | | **R-110** | **`main` is the installer's publish channel — there is no staging.** `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a 30 s period and nginx serves that working tree directly (`location /scripts/`, `root …/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`) receives. There is no tag, no pinned-version path, no staging copy and no rollback other than another push — for the artifact that runs as **root on a virgin box**, the single most privileged thing Felhom ships | **WAITING-ON-OPERATOR (S)** | operator ruling | **Two consequences worth stating:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29 — and the precaution recorded on the old R-94 row as "do not point every new box at an installer that has never run" **was never available to take**. **Open question for the operator, not a defect to fix blind:** whether `/scripts/` should serve a pinned release (tag-tracked path, or a versioned directory with the customer command naming a version) or whether `main`-tracking is the accepted shape for a one-operator product. Exposure today is zero — there are no boxes installing — which is exactly why it is cheap to decide now. **SECOND INSTANCE, found 2026-07-29 by the E-2d run and filed here rather than as a new ID:** `felhom-host-install.sh` fetches **nine** files from `raw/branch/main` (`:2072`–`:2206`) and the hub manifest vouches a sha for exactly **one** (`wrapper_sha256` → `felhom-pbs-apply`; re-checked this run, no drift). E-2a's `felhom-backup-target-apply` (`:2116`) is installed **0755 to `/usr/local/sbin` and root-fenced in sudoers**, validated only by `bash -n` — a root-executed artifact taken from `main` with no pinned integrity, which is this row's class exactly | CC | | **R-111** | ~~**The Day-0 artifact channel is 17 agent releases stale — a box installed today gets agent `0.96.0`, not `0.113.0`.**~~ `felhom-host-install.sh` does not use `main`: it reads the hub-vouched manifest (`:423-436`) and fetches Gitea generic packages (agent `:1945`, golden `:2573`). Gitea's newest are **agent 0.96.0** and **golden 0.161.0**, and the hub's manifest selects exactly those — so a fresh box lands on **agent 0.96.0 + controller 0.161.0** (global floor `v0.156.0` < the golden's 0.161.0, so no self-update) against `main`'s 0.113.0 / 0.185.1. Agent 0.113.0 reached both demo boxes by **direct deploy and was never published** | **SHIPPED 2026-07-29 — the channel now serves agent 0.113.0 + golden 0.185.1** | — | **FIXED the same day it was found.** Agent **0.113.0** built from the clean tree @ `58b598b` and published (`scripts/publish-agent.sh`), sha `5f3247f756cb658e…`, round-trip GET verified. Golden **0.185.1** baked on the nested drill VM embedding controller `0.185.1`, published, sha `dba00f3e845c415e…` — bake clean: `Result=success`, overlay2, **all 3 mounts included** (rootfs+mp0+mp1), 0 FATAL/exclusions, upload HTTP 201, token-leak grep 0; log `drill/bake-0.185.1.log`; GL-1 teardown done (guest 9100 purged, secrets shredded, disk restored to `virgin`). Hub Day-0 manifest moved **both together in one POST** so it never vouched a new agent against an old golden; `min_agent` **0.93.0 → 0.113.0**, which is what controller v0.185.0 declares (`felhom-controller/CHANGELOG.md:15`) — **zero fleet impact, verified: all three enrolled hosts already run agent 0.113.0, so no box is held.** `wrapper_sha256` preserved verbatim (re-checked against `configs/felhom-pbs-apply` — no drift). **The global controller floor was deliberately NOT raised**: the golden now bakes 0.185.1, so a fresh box needs no self-update, and raising it would have been an unnecessary fleet-wide write. Original finding follows. **Found 2026-07-29 by the E-2d Phase 0 gate, which stopped the run before a VM was created.** 17 unpublished releases (v0.97.0–v0.113.0) strand the **entire R-82 tiered-backup arc** plus **F-CRIT-2** (a failed backup looking fresh — 7 days silent) and **F-REBOOT** (a guest rebooted mid-backup never returns): a new customer's box would install without them. **Blocks E-2d's C3/C4/C5** — those test endpoints and events that do not exist in 0.96.0/0.161.0. The **controller is fine** (registry has 0.185.1, floor-driven self-update), so the gap is specific to the two Gitea-generic artifacts. **Mirror of R-110, not a duplicate:** R-110 = the installer publishes instantly with no staging; R-111 = the agent/golden publish gate exists and was never walked. Fix should decide whether publishing joins the release train rather than staying a remembered step (R-29's shape, one layer up). Evidence: `audits/E2D-fresh-vm-2026-07-29.md` **DEFERRED LEG, AND IT RECURRED → R-115.** This row's shipped half stands and is not reopened: the bump happened, was verified, and was proven end-to-end by the E-2d install. But its own closing line — *decide whether publishing joins the release train rather than staying a remembered step* — was never acted on, and agent 0.114.0 reproduced the exact condition the same afternoon. The recurrence is filed as **R-115**, not as a reopen, because the stale-channel finding is closed while the process defect that caused it is a distinct problem with a distinct fix. | CC |