E-2d: Phase 0 STOP — the Day-0 artifact channel cannot deliver the code under test

No VM created, no install run, no box touched. The run stopped at the Phase 0
gate per runbook §3, before provisioning.

felhom-host-install.sh does not install what is on main. resolve_artifacts()
(:423-436) reads the hub-vouched manifest and fetches Gitea GENERIC PACKAGES
(agent :1945, golden :2573). Gitea's newest are agent 0.96.0 and golden 0.161.0;
the hub manifest selects exactly those; the global floor v0.156.0 is below the
golden's 0.161.0 so nothing self-updates. A fresh box therefore lands on agent
0.96.0 + controller 0.161.0 against main's 0.113.0 / 0.185.1. Agent 0.113.0
reached both demo boxes by direct deploy and is not in the channel at all.

Claim impact, each pinned to its introducing commit:
- C1 (real rc=0 1.22.0 install) and C2 (Case B natural) — ACHIEVABLE, not run;
  both are installer-side and host-install is served at 1.22.0.
- C3 — BLOCKED: banner + GET /api/storage/backup-target are controller v0.185.1
  (cdaeb36), copy v0.185.0 (3f7cf2a). Unblocks cheaply by raising the hub floor
  to >=0.185.0; measured fleet impact nil (both demo boxes already 0.185.1).
- C4 — BLOCKED: needs controller v0.185.1 + agent v0.113.0 (58b598b).
- C5 — BLOCKED: needs controller v0.184.0 (c1a63de) + agent v0.112.0.

Filed R-111 (P1): 17 unpublished agent releases (v0.97.0-v0.113.0) strand the
entire R-82 tiered-backup arc plus F-CRIT-2 and F-REBOOT, so a new customer's
box installs without them. Mirror of R-110, not a duplicate.

- audits/E2D-fresh-vm-2026-07-29.md — all four Phase 0 answers recorded so a
  resumed run does not re-derive them (cadence 30s; hot-detach available; ISO
  present; local-lvm fence re-measured at 38.77%, unchanged).
- OPEN-ITEMS.md — R-111 opened; E-2d re-stated, NOT closed.
- ROADMAP.md — R-111 under P1.
- capability map NOT touched: nothing was proven live.

The §5.1a operator STOP is retired — HUB_PW is in ~/.config/credentials and hub
auth was verified, so CC can bind on a resumed run.
This commit is contained in:
2026-07-29 11:43:00 +02:00
parent 91a1dad0f3
commit f3f0d58844
4 changed files with 259 additions and 1 deletions
+58
View File
@@ -0,0 +1,58 @@
# REPORT — E-2d fresh-VM run: STOPPED at Phase 0 (2026-07-29)
`RUNBOOK-e2d-fresh-vm-2026-07-29.md`, executed by CC on DooPlex. Full evidence:
`documentation/audits/E2D-fresh-vm-2026-07-29.md`. Written as `REPORT-e2d.md` per the runbook §8.5;
root `REPORT.md` untouched.
**Outcome: no VM created, no install run, no box touched, no teardown needed.** The run stopped at the
Phase 0 gate, per §3 (*"If any gate fails, STOP and report — do not adapt around it"*).
## Why
`felhom-host-install.sh` does not install what is on `main`. It resolves the hub-vouched manifest
(`:423-436`) and fetches **Gitea generic packages** (agent `:1945`, golden `:2573`). Gitea's newest are
**agent `0.96.0`** and **golden `0.161.0`**; the hub's saved manifest selects exactly those; the global
controller floor is `v0.156.0`, below the golden's 0.161.0, so no self-update follows.
| Component | Fresh install gets | `main` / demo boxes |
|---|---|---|
| host-install | **1.22.0** | 1.22.0 |
| agent | **0.96.0** | 0.113.0 — **never published**, direct-deployed |
| controller | **0.161.0** | 0.185.1 |
## Claims
| Claim | Verdict | Pinned to |
|---|---|---|
| C1 — real rc=0 install of 1.22.0 | **ACHIEVABLE, not run** | installer-side |
| C2 — Case B fires naturally | **ACHIEVABLE, not run** | `felhom-host-install.sh:627`, `:653-655` |
| C3 — degraded banner renders | **BLOCKED** | controller v0.185.1 `cdaeb36`; copy v0.185.0 `3f7cf2a` |
| C4 — offer moves the target | **BLOCKED** | controller v0.185.1 + agent v0.113.0 `58b598b` |
| C5 — `backup_target_absent` e2e | **BLOCKED** | controller v0.184.0 `c1a63de` + agent v0.112.0 |
C3 unblocks by raising the hub floor to ≥0.185.0 (controller 0.185.1 **is** in the registry); measured
fleet impact nil — both demo boxes already run 0.185.1, Peti is DOWN 14 d and already below the current
floor. C4/C5 need agent 0.112.0/0.113.0 **published**, which runbook §0 forbids this run from doing.
## Filed
**R-111 (P1)** — the Day-0 artifact channel is 17 agent releases stale. v0.97.0v0.113.0 unpublished,
stranding the entire R-82 tiered-backup arc plus **F-CRIT-2** and **F-REBOOT**. A new customer's box
installs without them. Mirror of R-110, not a duplicate: R-110 = publishes instantly with no staging;
R-111 = the publish gate exists and was never walked.
## Record
- `OPEN-ITEMS.md`**R-111 opened** (READY (M), P1). **E-2d re-stated**, not closed: the attempt, the
blocker, the C1/C2-vs-C3/C4/C5 split, and every Phase 0 answer so a resumed run does not re-derive them.
- `ROADMAP.md` — R-111 filed under **P1 — closed-alpha blockers**.
- `architecture/00-capability-map.md`**not touched.** Nothing was proven live; no row qualifies.
## Not done
C1C5 all unproven. No VM, no ISO boot, no appliance registered/bound/discarded, no storage created or
modified on demo-hp, no drive attached or detached, no hub setting written (the floor was read only),
no code written.
The runbook's **§5.1a operator STOP is retired**: `HUB_PW` is in `~/.config/credentials` and hub auth
was verified working, so CC can perform the bind itself on a resumed run.
@@ -0,0 +1,198 @@
# E2D-fresh-vm-2026-07-29 — Phase 0 STOP: the Day-0 artifact channel cannot deliver the code under test
**Run:** `RUNBOOK-e2d-fresh-vm-2026-07-29.md`, executed by CC on DooPlex, 2026-07-29.
**Outcome:** **STOPPED at Phase 0, before any VM was created.** No VM provisioned, no install run, no
box touched, no teardown required. Per §3: *"If any gate fails, STOP and report — do not adapt around it."*
**The finding in one sentence:** a box installed today through the real customer chain receives
**agent 0.96.0** and **controller 0.161.0**, because those are the newest artifacts ever published to
the Day-0 channel — so three of the five claims (C3, C4, C5) test endpoints and events that do not
exist in the software a fresh box actually runs.
---
## 1. Baselines confirmed
| Artifact | Runbook §1 | Confirmed | Source |
|---|---|---|---|
| hub | 0.81.0 | **0.81.0** | `manifests/hub.yaml:128`; `hub/CHANGELOG.md:1`; live deploy image |
| agent | 0.113.0 | **0.113.0** on `main` @ `58b598b` | `felhom-agent` HEAD; live `felhom-agent -version` on demo-hp |
| controller | 0.185.1 | **0.185.1** on `main` @ `cdaeb36` | `felhom-controller` HEAD; live image on guest 9201 |
| host-install | 1.22.0 | **1.22.0** | `scripts/felhom-host-install.sh:187` |
| felhom.eu | `91a1dad` | **`91a1dad`**, clean, == `origin/main` | `git rev-parse` |
These are the versions **on `main` and on the demo boxes**. They are not the versions a fresh install
receives — that distinction is the whole finding.
---
## 2. THE BLOCKER — the Day-0 artifact channel is 17 agent releases stale
### 2.1 The chain, read at source
`felhom-host-install.sh` does not use `main`. It resolves a hub-vouched manifest and fetches versioned
packages from Gitea:
- `resolve_artifacts()` (`:423-436`) → `GET $HUB_URL/api/v1/artifacts/$CUSTOMER_ID``agent.version`,
`golden.version` (served by `hub/internal/api/handler.go:2120` `handleArtifactManifest`).
- agent binary ← `$GITEA_BASE/api/packages/admin/generic/felhom-agent/$ART_AGENT_VER/felhom-agent` (`:1945`)
- golden ← `$GITEA_BASE/api/packages/admin/generic/felhom-golden/$ART_GOLDEN_VER/golden.tar.zst` (`:2573`)
### 2.2 What that channel actually holds (Gitea API, authenticated, 2026-07-29)
```
felhom-agent : newest = 0.96.0
all = 0.79.0 0.80.0 0.81.0 0.84.0 0.85.0 0.86.0 0.87.0 0.88.0 0.89.0
0.90.0 0.91.0 0.91.1 0.91.2 0.92.0 0.92.1 0.93.0 0.96.0
felhom-golden: newest = 0.161.0
all = 0.136.0 0.143.0 0.146.0 0.153.0 0.161.0
```
Hub's saved Day-0 manifest (`/configuration`, selected options): **agent `0.96.0`**
(sha `af938601…`), **golden `0.161.0`** (sha `77624408…`), min-agent `0.93.0`.
Hub global controller floor: **v0.156.0** (DB override; env fallback v0.120.0).
### 2.3 Therefore a fresh box lands on
| Component | Fresh install gets | `main` / demo boxes | Gap |
|---|---|---|---|
| host-install | **1.22.0** (website git-sync from `main`) | 1.22.0 | none |
| agent | **0.96.0** | 0.113.0 | **17 releases** |
| controller | **0.161.0** (golden bake; floor 0.156.0 < 0.161.0 ⇒ no self-update) | 0.185.1 | **24 releases** |
**Agent 0.113.0 is not in the channel at all** — it reached both demo boxes by direct deploy, never
through publish. Confirmed live: demo-hp reports `felhom-agent 0.113.0` while Gitea's newest is 0.96.0.
### 2.4 What that strands — 17 unpublished agent releases (`felhom-agent/CHANGELOG.md`)
| Version | What it carries |
|---|---|
| v0.97.0v0.104.0 | **the entire R-82 per-target backup-tier arc** (local daily + PBS weekly), incl. v0.98.0's 30-minute false-failure bound, v0.101.0/0.102.0 scratch-guest + defer fixes, v0.103.0 R-84, v0.104.0 R-85 restore-test |
| v0.105.0 | R-88 Part 2 — the agent can say `unknown` |
| v0.106.0 | **F-CRIT-2** — a failed backup must not look like a fresh one (the phantom snapshot that reset the freshness clock, 7 days silent) |
| v0.107.0 | **F-REBOOT** — a guest rebooted during its backup never comes back |
| v0.108.0, v0.110.0 | F-LEAK — scratch-guest VMID band leak |
| v0.109.0 | F-OBS — the guest-power watchdog's positive observable |
| v0.111.0 | E-2c — the backup-target drive can no longer be ejected out from under the backup |
| v0.112.0 | **E-2b**`GET /disks` flags the backup-target drive |
| v0.113.0 | **E-2a** — the guarded wrapper + `POST /backup/target` |
**A new customer box installed today therefore runs an agent that predates the whole tiered-backup
model and lacks F-CRIT-2 and F-REBOOT** — two customer-impacting silent-failure fixes. That is a
larger finding than E-2d itself and is filed as **R-111**.
---
## 3. Claim-by-claim impact, each pinned to its introducing commit
| Claim | Needs | Introduced in | Fresh box has | Verdict |
|---|---|---|---|---|
| **C1** rc=0 real install, banner names 1.22.0 | host-install 1.22.0 | website ← `main` | **1.22.0** | **ACHIEVABLE** |
| **C2** Case B DEGRADED lines + `resolved local` | host-install `configure_backup_target()` `:627`, Case B `:653-655` | host-install 1.22.0 | **1.22.0** | **ACHIEVABLE** — installer-side only |
| **C3** degraded banner via `GET /api/storage/backup-target` | controller **v0.185.1** (`cdaeb36`); copy „lemezhiba ellen nem" **v0.185.0** (`3f7cf2a`) | — | controller 0.161.0 | **BLOCKED** — endpoint and copy do not exist |
| **C4** offer → `POST …/assign``restart_required` | controller **v0.185.1** (`cdaeb36`) + agent **v0.113.0** `POST /backup/target` (`58b598b`) | — | 0.161.0 / 0.96.0 | **BLOCKED** |
| **C5** `backup_target_absent` / `_restored` | controller **v0.184.0** (`c1a63de`) + agent **v0.112.0** | — | 0.161.0 / 0.96.0 | **BLOCKED** |
**C3 has a cheap unblock; C4 and C5 do not.** C3 needs only the hub's global controller floor raised
0.156.0 → ≥0.185.0, because controller **0.185.1 IS published** to the registry (60 tags, newest
0.185.1) — a fresh box would self-update on its first report. C4/C5 need agent 0.112.0/0.113.0
*published as Gitea generic packages*, which §0 of the runbook explicitly forbids this run from doing
(*"Does not touch: … any published artifact"*).
Fleet impact of raising the floor, measured rather than assumed (hub dashboard, 2026-07-29):
| Customer | Controller | Effect of floor → 0.185.1 |
|---|---|---|
| Demo Ügyfél | 0.185.1 | none — already at it |
| Demo HP | 0.185.1 | none — already at it |
| Peti Proxmox | 0.115.0, **DOWN 14d** | already below the current 0.156.0 floor, so it updates on return either way; only the target version changes |
---
## 4. Phase 0 gate answers (all four completed before the stop)
### 4.1 Storage placement — gate PASSES, fence confirmed
`pvesm status` on demo-hp, 2026-07-29 (BEFORE; there is no AFTER — no VM was created):
```
Name Type Status Total (KiB) Used (KiB) Available (KiB) %
felhom-backup dir active 983379700 2293856 931059232 0.23%
felhom-pbs pbs active 0 0 0 0.00%
local dir active 40516856 14721460 23705004 36.33%
local-lvm lvmthin active 56545280 21922605 34622674 38.77%
```
`local-lvm` is thin, pool `<53.93g`, **data 38.77 %** — matching the runbook's "~54 GB pool, 38.8 %"
to 0.03 pp, so the picture has **not** materially changed. Allocated LVs on it: `vm-9201-disk-{0,1,2}`
(32+50+20 G, the **live customer guest**), `vm-300-disk-{0,1}` (drill-r50, 32 G) and its `r50pre`
snapshots (32 G) ⇒ ~166 G allocated against 53.93 G real. `vgs` shows only 14.75 G VFree.
**The fence holds: nothing goes on `local-lvm`.**
`/mnt/nvme-1tb` = `/dev/nvme0n1`, ext4, 938 G, **888 G free**. ≥100 G requirement satisfied.
**Placement decision (recorded deliberately, per §3.1):** the only storage with `content images` is
`local-lvm` (forbidden). `felhom-backup` is a `dir` at exactly `/mnt/nvme-1tb` but is `content backup`
only **and is demo-hp's live backup target** — widening its content set would mutate the very storage
E-2c/E-2 role logic keys on, which §0/§9 forbid. So a *new* dir storage would have been required.
Both remaining shapes carry a cost and the choice was **not** forced this run because the run stopped:
a second storage at the same mountpoint risks perturbing the agent's role resolution on a live box;
a storage at a **subdirectory** fails `exactMount` and reports `disconnected` in the host report,
which is not purely cosmetic since it can reach the hub's storage monitor. **This decision is left
open and is an input to any resumed run.**
### 4.2 Drive-gate cadence — answered
- Symbol: `driveGateLoop``s.ReconcileDriveGates()`
- Definition: `controller/internal/web/intermediary.go:328` (gate itself at `:267`)
- Registration: `controller/internal/web/server.go:217``go s.driveGateLoop()`
- **Interval: `time.NewTicker(30 * time.Second)`** (`intermediary.go:337`)
⇒ Stage 7's derived budget would be **two cycles = 60 s** of controller-side reconcile, *plus* the
agent's own `/disks` refresh, since `ReconcileDriveGates` consumes `resp.Disks` (`:281`). Recorded so
a resumed run does not have to re-derive it.
### 4.3 Hot-detach — feasible, not exercised
demo-hp: AMD `svm` present, `/sys/module/kvm_amd/parameters/nested = 1`, 8 cores, 29 GB RAM (25 GB
available). The working nested-PVE reference is VM 300 (`drill-r50`): `bios: ovmf`,
`efitype=4m,pre-enrolled-keys=0` (Secure Boot OFF, as the mkimage loader requires), `machine: q35`,
`scsihw: virtio-scsi-single`, `cpu: host`. `virtio-scsi-single` + PVE's default
`hotplug: disk,network,usb` supports SCSI hot-detach, so the **live-transition** variant of Stage 7
was available. Not exercised.
### 4.4 Install route — ISO available, route selected, not used
Two generic (pairing) ISOs are present on demo-hp `local`:
`felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso` (**preferred**, the v1.25.0 train, mkimage
loader for nested VMs) and `felhom-pve-9.2-1-v1.22.0-nested-canary-generic.iso`.
The **ISO/PAIRING route was selected**. It reaches the same installer invocation — `run_pairing`
(`felhom-bootstrap.sh:441`) falls through to `run_direct` (`:495-499`) in the same invocation, which
fetches `$INSTALL_URL` (`:322-330`) and runs it (`:343`). Not used — the run stopped first.
**Operator STOP: not required.** `HUB_PW` is present in `~/.config/credentials` and hub auth was
verified working (`http://10.43.52.34:8080/` → 200 with `curl -u ":$HUB_PW"`; public ingress also 200),
so CC could have performed the bind itself. The runbook's §5.1a STOP is retired for future runs.
---
## 5. Incidental observations (filed, not acted on)
1. **A stale unclaimed appliance is sitting in the hub**: `206c8838-7755-4751-8a16-a842853d718f`,
pairing code `QWA-WJE`, MAC `bc:24:11:ea:55:d0`, "Standard PC (Q35 + ICH9, 2009)" on an
AMD V1756B with 7.7 GB — i.e. a **nested VM on demo-hp**, first and last seen **5 d ago**. It is a
leftover from the 2026-07-25 ISO work and has never been bound or discarded. A resumed run must
distinguish its own appliance from this one, or discard it first.
2. **The agent has no container image** (`/v2/admin/felhom-agent/tags/list` → 0 tags). It ships only
as a Gitea generic package, which is why the publish gap is invisible from the registry.
3. `drill-r50` (VM 300) still holds a `r50pre` snapshot pair consuming ~32 G of allocation on the
over-subscribed `local-lvm`. Untouched per §9; noted because it is part of why the pool is tight.
---
## 6. What did not happen
No VM was created. No ISO booted. No install ran. No appliance registered, bound or discarded. No
storage was created or modified on demo-hp. No drive attached or detached. No hub setting changed
(the floor was **read**, not written). No teardown was needed. **C1C5 are all UNPROVEN**, and the
E-2 / E-2d rows are unchanged except to record this blocker.
+2 -1
View File
@@ -11,9 +11,10 @@ State: `BLOCKED` · `READY` · `WAITING-ON-OPERATOR` · `WATCHING`. Every row ha
|---|---|---|---|---|---|
| **R-88a** | ~~Failing backup re-quiesces every 5 min, no backoff~~ | **SHIPPED** (controller v0.176.0, 2026-07-27) | — | Live on both boxes; breaker 15m→4h, per-tier, never permanent | — |
| **R-88b** | ~~`/backup/due` cannot say *unknown*~~ | **SHIPPED + PROVEN-LIVE** (agent v0.105.0 + controller v0.178.0, 2026-07-27) | — | `age_state=unknown` captured on real hardware during a deliberate ep0 outage; controller deferred, **zero app stacks stopped** | — |
| **E-2d** | **Prove E-2 on a fresh VM on the t740** — the only remaining route to four unproven items: a real `felhom-host-install.sh` **1.22.0** run (never done), Case B naturally (single-drive install renders the degraded banner without degrading a live box), a **claimable** customer so the three claim-gated items stop being gated, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **READY (M)** | — | **Space checked 2026-07-29 — NOT a blocker, with one constraint: the VM disk must NOT go on `local-lvm`.** That thin pool is over-subscribed (144 GB allocated against 54 GB, 38.8% used) on a box running a live customer guest, and a full thin pool corrupts every guest on it. `local` has only 23.7 GB and sits on `pve-root`. **Use `/mnt/nvme-1tb` (888 GB free).** Caveat found: a dir storage at a SUBDIRECTORY there will fail the agent's `exactMount` check and report `disconnected` in the host report — cosmetic, but decide the placement deliberately. **Do NOT unblock drill-r50** (deliberately blocked; unblocking it means the fixture stops representing anything real). **CORRECTED 2026-07-29 — the ISO is not an obstacle to the 1.22.0 proof, it is the best route to it.** `felhom-bootstrap.sh:96` fetches from `https://felhom.eu/scripts/felhom-host-install.sh`, not the hub, and that URL serves **1.22.0** (git-sync from `main`, ≤30 s — live fetch confirmed). A fresh ISO install therefore exercises 1.22.0 **automatically** — which makes the ISO leg the *stronger* proof, because it is the real customer path rather than a proxy for it, and it retires the "run 1.22.0 manually" step. **The Phase 0 question that decides the shape is ANSWERED — it was a source read, and the answer is yes.** A fresh VM with no baked customer-id lands in **PAIRING** mode (`felhom-bootstrap.sh:537-541`), not **DIRECT** (`:312`), and only DIRECT passes `--customer-id / --mode / --passphrase-file`. But on a 200 from `/api/v1/appliance/poll` the pairing loop writes the hub-delivered `FELHOM_CUSTOMER_ID` + `FELHOM_RETRIEVAL_PASSPHRASE` into the 0600 env, re-sources it and calls `run_direct` **in the same invocation** (`:495-499`) — so pairing reaches the identical installer invocation (`:322-343` — the single `$INSTALL_URL` fetch at `:322-330`, the `--customer-id/--mode/--hub-url/--passphrase-file` args array at `:334`, and the `bash "$SCRIPT_TMP" "${args[@]}"` call itself at `:343`) and the customer it yields is the one the operator bound, i.e. claimable. **So the ISO leg is the spine**; a manual 1.22.0 run is not needed as a separate scenario | CC |
| **E-2d** | **Prove E-2 on a fresh VM on the t740** — the only remaining route to four unproven items: a real `felhom-host-install.sh` **1.22.0** run (never done), Case B naturally (single-drive install renders the degraded banner without degrading a live box), a **claimable** customer so the three claim-gated items stop being gated, then add a drive (the offer) and unplug it (`backup_target_absent` end-to-end) | **READY (M)** | — | **Space checked 2026-07-29 — NOT a blocker, with one constraint: the VM disk must NOT go on `local-lvm`.** That thin pool is over-subscribed (144 GB allocated against 54 GB, 38.8% used) on a box running a live customer guest, and a full thin pool corrupts every guest on it. `local` has only 23.7 GB and sits on `pve-root`. **Use `/mnt/nvme-1tb` (888 GB free).** Caveat found: a dir storage at a SUBDIRECTORY there will fail the agent's `exactMount` check and report `disconnected` in the host report — cosmetic, but decide the placement deliberately. **Do NOT unblock drill-r50** (deliberately blocked; unblocking it means the fixture stops representing anything real). **CORRECTED 2026-07-29 — the ISO is not an obstacle to the 1.22.0 proof, it is the best route to it.** `felhom-bootstrap.sh:96` fetches from `https://felhom.eu/scripts/felhom-host-install.sh`, not the hub, and that URL serves **1.22.0** (git-sync from `main`, ≤30 s — live fetch confirmed). A fresh ISO install therefore exercises 1.22.0 **automatically** — which makes the ISO leg the *stronger* proof, because it is the real customer path rather than a proxy for it, and it retires the "run 1.22.0 manually" step. **The Phase 0 question that decides the shape is ANSWERED — it was a source read, and the answer is yes.** A fresh VM with no baked customer-id lands in **PAIRING** mode (`felhom-bootstrap.sh:537-541`), not **DIRECT** (`:312`), and only DIRECT passes `--customer-id / --mode / --passphrase-file`. But on a 200 from `/api/v1/appliance/poll` the pairing loop writes the hub-delivered `FELHOM_CUSTOMER_ID` + `FELHOM_RETRIEVAL_PASSPHRASE` into the 0600 env, re-sources it and calls `run_direct` **in the same invocation** (`:495-499`) — so pairing reaches the identical installer invocation (`:322-343` — the single `$INSTALL_URL` fetch at `:322-330`, the `--customer-id/--mode/--hub-url/--passphrase-file` args array at `:334`, and the `bash "$SCRIPT_TMP" "${args[@]}"` call itself at `:343`) and the customer it yields is the one the operator bound, i.e. claimable. **So the ISO leg is the spine**; a manual 1.22.0 run is not needed as a separate scenario. **RUN ATTEMPTED 2026-07-29 — STOPPED AT PHASE 0, no VM created, nothing touched (`audits/E2D-fresh-vm-2026-07-29.md`).** The blocker is **R-111**: a fresh box installs **agent 0.96.0 + controller 0.161.0**, not `main`'s 0.113.0/0.185.1, so **C3/C4/C5 test surfaces that do not exist on it** — the degraded banner + `GET /api/storage/backup-target` landed in controller **v0.185.1** (`cdaeb36`) with the copy in **v0.185.0** (`3f7cf2a`); `backup_target_absent` in **v0.184.0** (`c1a63de`); the offer's apply needs agent **v0.113.0** (`58b598b`). **C1 (a real rc=0 1.22.0 install) and C2 (Case B natural, `configure_backup_target()` `felhom-host-install.sh:627`, warnings `:653-655`) remain ACHIEVABLE TODAY** — both are installer-side and host-install is served at 1.22.0. **C3 unblocks cheaply** by raising the hub global floor 0.156.0 → ≥0.185.0 (controller 0.185.1 IS in the registry): measured fleet impact is **nil** — both demo boxes already run 0.185.1, and Peti is DOWN 14 d and already below the current floor. **C4/C5 need R-111 first.** Phase 0 answers are all recorded in the audit, so a resumed run does not re-derive them: drive-gate cadence **30 s** (`intermediary.go:337`, registered `server.go:217`) ⇒ a 60 s two-cycle budget; hot-detach available (`virtio-scsi-single` + default hotplug, VM 300 is the working reference); ISO present (`…v1.25.0-nested-vm-generic-mkimage.iso`); `/mnt/nvme-1tb` 888 G free and the `local-lvm` fence re-measured (38.77 %, unchanged). **The §5.1a operator STOP is retired**`HUB_PW` is in `~/.config/credentials` and hub auth was verified, so CC can bind. **One open decision carried forward:** where the VM disk's dir storage goes, since `local-lvm` is forbidden and `felhom-backup` is the live backup target — see the audit §4.1 | CC |
| **R-94** | **A hand-synced version constant drifts, and the gate that would catch it is never run**`hub/internal/web/configs.go:28` pins `hostInstallVersion = "1.19.0"` while `scripts/felhom-host-install.sh:187` is `SCRIPT_VERSION="1.22.0"` | **READY (XS)** | — | **CORRECTED 2026-07-29 — the earlier framing of this row was false and is retracted.** The constant selects no script: its only consumers are `configs.go:487` (`ScriptVersion`) and `render_test.go:219`, and it renders as a label at `customer_unified.html:494`. The install command beneath that label fetches `https://felhom.eu/scripts/felhom-host-install.sh` (`customer_unified.html:563`, `:1262`), which the website git-syncs from `main` on a 30 s period (`manifests/webpage.yaml`) — so **1.22.0 is what every install already gets** (live fetch, 2026-07-29). Every flag the generator emits is parsed by 1.22.0 (`customer_unified.html`~`:1210``:1238` vs `felhom-host-install.sh:1177``:1210`): **no functional gap, only a wrong number on the operator's screen.** Three legs, all XS: **(a)** derive the label from `SCRIPT_VERSION` rather than hand-syncing it, or delete it; **(b)** `scripts/hostinstall_gates.py` **fails today** and is invoked by no Makefile, hook or `CLAUDE.md` — wire it next to `site_gates.py` or delete it, because a gate nobody runs reads as coverage it is not providing (**this leg is one instance of → R-29**, which is the class: gates are enforced nowhere, and the enforcement decision belongs there, not here); **(c)** `render_test.go:219` compares the constant to itself and passes at any value — replace it with the cross-file assertion. **No longer blocked on E-2d** — it never gated anything | CC |
| **R-110** | **`main` is the installer's publish channel — there is no staging.** `manifests/webpage.yaml` git-syncs `/scripts/` from `--branch=main` on a 30 s period and nginx serves that working tree directly (`location /scripts/`, `root …/current`). So pushing `scripts/felhom-host-install.sh` **is** publishing it: within thirty seconds it is what every subsequent `felhom-bootstrap.sh` fetch (`scripts/iso/felhom-bootstrap.sh:96`) and every operator-run day-0 command (`customer_unified.html:563`) receives. There is no tag, no pinned-version path, no staging copy and no rollback other than another push — for the artifact that runs as **root on a virgin box**, the single most privileged thing Felhom ships | **WAITING-ON-OPERATOR (S)** | operator ruling | **Two consequences worth stating:** E-2d is not a gate *before* exposure — 1.22.0 has been the live installer since it hit `main` on 2026-07-29 — and the precaution recorded on the old R-94 row as "do not point every new box at an installer that has never run" **was never available to take**. **Open question for the operator, not a defect to fix blind:** whether `/scripts/` should serve a pinned release (tag-tracked path, or a versioned directory with the customer command naming a version) or whether `main`-tracking is the accepted shape for a one-operator product. Exposure today is zero — there are no boxes installing — which is exactly why it is cheap to decide now | CC |
| **R-111** | **The Day-0 artifact channel is 17 agent releases stale — a box installed today gets agent `0.96.0`, not `0.113.0`.** `felhom-host-install.sh` does not use `main`: it reads the hub-vouched manifest (`:423-436`) and fetches Gitea generic packages (agent `:1945`, golden `:2573`). Gitea's newest are **agent 0.96.0** and **golden 0.161.0**, and the hub's manifest selects exactly those — so a fresh box lands on **agent 0.96.0 + controller 0.161.0** (global floor `v0.156.0` < the golden's 0.161.0, so no self-update) against `main`'s 0.113.0 / 0.185.1. Agent 0.113.0 reached both demo boxes by **direct deploy and was never published** | **READY (M) — P1, gates the first remote tester** | — | **Found 2026-07-29 by the E-2d Phase 0 gate, which stopped the run before a VM was created.** 17 unpublished releases (v0.97.0v0.113.0) strand the **entire R-82 tiered-backup arc** plus **F-CRIT-2** (a failed backup looking fresh — 7 days silent) and **F-REBOOT** (a guest rebooted mid-backup never returns): a new customer's box would install without them. **Blocks E-2d's C3/C4/C5** — those test endpoints and events that do not exist in 0.96.0/0.161.0. The **controller is fine** (registry has 0.185.1, floor-driven self-update), so the gap is specific to the two Gitea-generic artifacts. **Mirror of R-110, not a duplicate:** R-110 = the installer publishes instantly with no staging; R-111 = the agent/golden publish gate exists and was never walked. Fix should decide whether publishing joins the release train rather than staying a remembered step (R-29's shape, one layer up). Evidence: `audits/E2D-fresh-vm-2026-07-29.md` | CC |
| **R-29** | **The green gates are not enforced anywhere — one was RED for 16 releases before anyone ran it.** This is the **class**, not an instance: a gate that exists, asserts something true, is red, and is invoked by nothing reads as coverage it is not providing. `controller/scripts/docker_run_volume_path_gate.py` failed continuously from **2026-07-14 (v0.129.0)** until R-7b's close-out ran it by hand at v0.145.0 — sixteen releases in which every REPORT said "green" | **READY (S for (a) / M for (b))** | — | **This item has existed at `ROADMAP.md:158` since before the register was rebuilt (2026-07-27) and was never carried across — that omission is itself part of the finding**, because it is an open item *about work not getting done* that then went missing from the page that decides what gets done. Two separable parts, per R-29's own analysis: **(a)** the `docker_run_volume_path_gate` finding is benign and the fix is a 3-line ALLOWLIST addition with its why — **not** a rewrite of the flagged call — and it gets its own reviewed diff, never bundled into a feature commit; **(b)** the systemic half, the real item: decide where gates run (pre-push hook, `build.sh` step, or CI) and make a red gate block the train the way the Go green gate does. **Two further orphans confirmed 2026-07-29** by repo-wide grep across all file types + sibling repos + `~/.claude` settings/skills/hooks + `.git/hooks` (none non-sample) + Makefile/justfile/Taskfile find (only `hub/Makefile`, zero `gate` occurrences) + CI-directory find (**this repo has no CI at all**) — every one of the 19 hits is a docstring, a code comment or prose, and **not one is an invocation**: `scripts/hostinstall_gates.py`**RED today** (`hub Setup-tab hostInstallVersion=1.19.0 != SCRIPT_VERSION=1.22.0`, exit 1), the same finding as **R-94 leg (b)** — and `scripts/hub_confirm_gate.py`. Of the four gates in `scripts/`, only `site_gates.py` is mandated anywhere (`CLAUDE.md:153`) and `manifest_bearer_gate.py` is named in `runbooks/secrets.md:76`. **In R-29's own words, carried forward deliberately: do not mint a new ID for a new instance** — the 2026-07-18 rehearsal independently re-raised this item and no second ID was minted then either | CC |
| **R-95** | restic offsite credential **can delete** (`readonly=False`, `forget --prune` runs from the box); SFTP cannot express append-only | **READY** | — | Root exposure still open. Mitigation now ARMED — split prune off-box or move to REST `--append-only` | CC |
| **R-86** | Restore-tests are interval-scheduled, not backup-aligned | **READY** | R-90 (ep0 headroom) informs cadence | Trigger a tier ~24 h after **its own** newest archive | CC |
+1
View File
@@ -20,6 +20,7 @@
| ID | Item | Size | Status | Notes / map rows flipped |
|----|------|------|--------|--------------------------|
| R-111 | **The Day-0 artifact channel is 17 agent releases stale — a box installed today gets agent `0.96.0`, not `0.113.0`** | M | idea — found 2026-07-29 by the E-2d Phase 0 gate | **The fleet's live versions are not the fleet's INSTALLABLE versions, and only the first were ever checked.** `felhom-host-install.sh` does not use `main`: `resolve_artifacts()` (`:423-436`) reads the hub-vouched manifest (`GET /api/v1/artifacts/<customer>`, `hub/internal/api/handler.go:2120`) and fetches versioned **Gitea generic packages** — agent from `:1945`, golden from `:2573`. Gitea holds **`felhom-agent` newest `0.96.0`** and **`felhom-golden` newest `0.161.0`**; the hub's saved manifest selects exactly those. So a fresh box lands on **agent 0.96.0 + controller 0.161.0** (golden bake; the global floor is `v0.156.0` < 0.161.0, so it does not self-update) against `main`'s 0.113.0 / 0.185.1. Agent 0.113.0 reached both demo boxes by **direct deploy and is not in the channel at all** — demo-hp reports `felhom-agent 0.113.0` while Gitea's newest is 0.96.0. **17 unpublished releases (`felhom-agent/CHANGELOG.md` v0.97.0v0.113.0)**, including the ENTIRE R-82 per-target backup-tier arc (v0.97.0v0.104.0), **F-CRIT-2** (v0.106.0 — a failed backup looking fresh, 7 days silent), **F-REBOOT** (v0.107.0 — a guest rebooted mid-backup never returns), F-LEAK (v0.108.0/0.110.0), F-OBS (v0.109.0), E-2c (v0.111.0), E-2b (v0.112.0), E-2a (v0.113.0). **P1 because it gates the first remote tester:** their box would install an agent predating the tiered-backup model and both silent-failure fixes. **Mirror of R-110, not a duplicate:** R-110 is *the installer publishes instantly with no staging*; this is *the agent and golden have a deliberate publish+vouch gate and it was never walked* — opposite failure modes of one subject, different fixes. Contrast worth keeping: the **controller** is fine (registry has 0.185.1; it self-updates from the floor), so the gap is specific to the two Gitea-generic artifacts. **Decide as part of the fix:** whether publishing becomes part of the release train rather than a separate remembered step — this is R-29's shape (a gate that exists and is never walked) one layer up. Evidence: `audits/E2D-fresh-vm-2026-07-29.md` §2 |
| R-1 | **Peti convergence***the appliance half is DONE; this item is now Peti-only.* **Rehearsal EXECUTED 2026-07-18** (`tests/VALIDATION-n100-rehearsal-2026-07-18.md`): the full final-product flow ran on real metal in one pass (RESET → generic ISO → **customer self-bind** → day-0 → floor lift → escrow ceremony → offsite snapshots), which retires the "supervised rehearsal" dependency that R-13/R-21/R-23/R-24/R-27/R-28 were all parked behind. **Surviving half: Peti's clean-slate proxmox2 reinstall + the parked publish trains on a REAL REMOTE customer** — the one thing a demo box on the operator's own LAN can never prove. | L | **rehearsal DONE; Peti half open** | Flips: publish train PARTIAL→PROVEN-LIVE; appliance/BYO/day-0 "real customer" notes; escrow ceremony. The single biggest unproven surface — an alpha where fixes can't ship remotely is dead. **Reinstall arc SHIPPED hub v0.57.0 (2026-07-16):** the clean-slate reinstall-of-existing-customer path is now first-class — claim re-issue (F2), offsite re-issue (F3), escrow-honesty-on-re-issue (2.3) all auto-fire on re-enrollment. Peti's proxmox2 clean-slate now walks a supported path |
| R-2 | ~~Resolve ~215 lines of foreign WIP in felhom.eu clone (`hub/internal/notify/`, `store.go`, `hub/internal/claim/`)~~ | S | **killed** (2026-07-16) | Not a real issue: the "foreign WIP" was in-flight code from a concurrent CC session on the customer-claim arc, snapshotted before it committed. All of it landed cleanly — `notify/`+`claim/engine.go` in `6b40eb8` (v0.50.0), `store.go` in `a1d0450` (v0.54.0), plus follow-up `e205a2d`; v0.55.0 shipped. Working tree is clean, no stashes. Lesson already codified: never run two writing sessions on one felhom.eu clone (CLAUDE.md §git add -A) |
| R-3 | Friend-alpha onboarding runbook (generalized from `pilot/RUNBOOK-peti-return-2026-07-13`): hardware prep → golden → install → claim → ceremony → "first restore by the customer" scripted step | M | idea | Flips: "customer performs a restore" MISSING row; produces the tester-agreement sibling of `PETI-tester-agreement.md`. **Next from-scratch rehearsal to include customer DELETE + re-create** — the ESCROW cascade is now DEFINED (hub v0.60.1): host delete DEMOTES escrow to retained custody (never destroys), customer Danger-zone delete PURGES it (the one true purge point). **S6b (manual stale-host delete before re-enroll) is OBSOLETE** — re-enrollment upserts the existing host row cleanly (`store.UpsertHost` ON CONFLICT DO UPDATE; `handleAdminCreateHost` no duplicate refusal) + the v0.57.0 arc auto-fires the re-issues; the rehearsal live-confirms it. **NON-escrow offboarding NOW ANSWERED by the middle-tier Customer RESET (hub v0.61.0, LIVE):** one operator action deprovisions the Hetzner sub-account/box (repo data destroyed), destroys the PBS namespace + backup groups + token, clears the DR recipe / one-time secret / claim state / retained escrow custody (separate ack) — identity + basic config survive. WG peer release rides host delete (peers are host-scoped, gone before RESET runs — RESET refuses while any host row exists). **Remaining consistency gap:** the customer Danger-zone DELETE still leaves host rows and does NOT run the offsite/PBS teardown (RESET is the teardown path; DELETE is escrow-purge + config-drop). Decide whether DELETE should require a prior RESET (or subsume it) — new item R-25b |