# E2D-fresh-vm-2026-07-29 — Phase 0 STOP: the Day-0 artifact channel cannot deliver the code under test **Run:** `RUNBOOK-e2d-fresh-vm-2026-07-29.md`, executed by CC on DooPlex, 2026-07-29. **Outcome:** **STOPPED at Phase 0, before any VM was created.** No VM provisioned, no install run, no box touched, no teardown required. Per §3: *"If any gate fails, STOP and report — do not adapt around it."* **The finding in one sentence:** a box installed today through the real customer chain receives **agent 0.96.0** and **controller 0.161.0**, because those are the newest artifacts ever published to the Day-0 channel — so three of the five claims (C3, C4, C5) test endpoints and events that do not exist in the software a fresh box actually runs. --- ## 1. Baselines confirmed | Artifact | Runbook §1 | Confirmed | Source | |---|---|---|---| | hub | 0.81.0 | **0.81.0** | `manifests/hub.yaml:128`; `hub/CHANGELOG.md:1`; live deploy image | | agent | 0.113.0 | **0.113.0** on `main` @ `58b598b` | `felhom-agent` HEAD; live `felhom-agent -version` on demo-hp | | controller | 0.185.1 | **0.185.1** on `main` @ `cdaeb36` | `felhom-controller` HEAD; live image on guest 9201 | | host-install | 1.22.0 | **1.22.0** | `scripts/felhom-host-install.sh:187` | | felhom.eu | `91a1dad` | **`91a1dad`**, clean, == `origin/main` | `git rev-parse` | These are the versions **on `main` and on the demo boxes**. They are not the versions a fresh install receives — that distinction is the whole finding. --- ## 2. THE BLOCKER — the Day-0 artifact channel is 17 agent releases stale ### 2.1 The chain, read at source `felhom-host-install.sh` does not use `main`. It resolves a hub-vouched manifest and fetches versioned packages from Gitea: - `resolve_artifacts()` (`:423-436`) → `GET $HUB_URL/api/v1/artifacts/$CUSTOMER_ID` → `agent.version`, `golden.version` (served by `hub/internal/api/handler.go:2120` `handleArtifactManifest`). - agent binary ← `$GITEA_BASE/api/packages/admin/generic/felhom-agent/$ART_AGENT_VER/felhom-agent` (`:1945`) - golden ← `$GITEA_BASE/api/packages/admin/generic/felhom-golden/$ART_GOLDEN_VER/golden.tar.zst` (`:2573`) ### 2.2 What that channel actually holds (Gitea API, authenticated, 2026-07-29) ``` felhom-agent : newest = 0.96.0 all = 0.79.0 0.80.0 0.81.0 0.84.0 0.85.0 0.86.0 0.87.0 0.88.0 0.89.0 0.90.0 0.91.0 0.91.1 0.91.2 0.92.0 0.92.1 0.93.0 0.96.0 felhom-golden: newest = 0.161.0 all = 0.136.0 0.143.0 0.146.0 0.153.0 0.161.0 ``` Hub's saved Day-0 manifest (`/configuration`, selected options): **agent `0.96.0`** (sha `af938601…`), **golden `0.161.0`** (sha `77624408…`), min-agent `0.93.0`. Hub global controller floor: **v0.156.0** (DB override; env fallback v0.120.0). ### 2.3 Therefore a fresh box lands on | Component | Fresh install gets | `main` / demo boxes | Gap | |---|---|---|---| | host-install | **1.22.0** (website git-sync from `main`) | 1.22.0 | none | | agent | **0.96.0** | 0.113.0 | **17 releases** | | controller | **0.161.0** (golden bake; floor 0.156.0 < 0.161.0 ⇒ no self-update) | 0.185.1 | **24 releases** | **Agent 0.113.0 is not in the channel at all** — it reached both demo boxes by direct deploy, never through publish. Confirmed live: demo-hp reports `felhom-agent 0.113.0` while Gitea's newest is 0.96.0. ### 2.4 What that strands — 17 unpublished agent releases (`felhom-agent/CHANGELOG.md`) | Version | What it carries | |---|---| | v0.97.0–v0.104.0 | **the entire R-82 per-target backup-tier arc** (local daily + PBS weekly), incl. v0.98.0's 30-minute false-failure bound, v0.101.0/0.102.0 scratch-guest + defer fixes, v0.103.0 R-84, v0.104.0 R-85 restore-test | | v0.105.0 | R-88 Part 2 — the agent can say `unknown` | | v0.106.0 | **F-CRIT-2** — a failed backup must not look like a fresh one (the phantom snapshot that reset the freshness clock, 7 days silent) | | v0.107.0 | **F-REBOOT** — a guest rebooted during its backup never comes back | | v0.108.0, v0.110.0 | F-LEAK — scratch-guest VMID band leak | | v0.109.0 | F-OBS — the guest-power watchdog's positive observable | | v0.111.0 | E-2c — the backup-target drive can no longer be ejected out from under the backup | | v0.112.0 | **E-2b** — `GET /disks` flags the backup-target drive | | v0.113.0 | **E-2a** — the guarded wrapper + `POST /backup/target` | **A new customer box installed today therefore runs an agent that predates the whole tiered-backup model and lacks F-CRIT-2 and F-REBOOT** — two customer-impacting silent-failure fixes. That is a larger finding than E-2d itself and is filed as **R-111**. --- ## 3. Claim-by-claim impact, each pinned to its introducing commit | Claim | Needs | Introduced in | Fresh box has | Verdict | |---|---|---|---|---| | **C1** rc=0 real install, banner names 1.22.0 | host-install 1.22.0 | website ← `main` | **1.22.0** | **ACHIEVABLE** | | **C2** Case B DEGRADED lines + `resolved local` | host-install `configure_backup_target()` `:627`, Case B `:653-655` | host-install 1.22.0 | **1.22.0** | **ACHIEVABLE** — installer-side only | | **C3** degraded banner via `GET /api/storage/backup-target` | controller **v0.185.1** (`cdaeb36`); copy „lemezhiba ellen nem" **v0.185.0** (`3f7cf2a`) | — | controller 0.161.0 | **BLOCKED** — endpoint and copy do not exist | | **C4** offer → `POST …/assign` → `restart_required` | controller **v0.185.1** (`cdaeb36`) + agent **v0.113.0** `POST /backup/target` (`58b598b`) | — | 0.161.0 / 0.96.0 | **BLOCKED** | | **C5** `backup_target_absent` / `_restored` | controller **v0.184.0** (`c1a63de`) + agent **v0.112.0** | — | 0.161.0 / 0.96.0 | **BLOCKED** | **C3 has a cheap unblock; C4 and C5 do not.** C3 needs only the hub's global controller floor raised 0.156.0 → ≥0.185.0, because controller **0.185.1 IS published** to the registry (60 tags, newest 0.185.1) — a fresh box would self-update on its first report. C4/C5 need agent 0.112.0/0.113.0 *published as Gitea generic packages*, which §0 of the runbook explicitly forbids this run from doing (*"Does not touch: … any published artifact"*). Fleet impact of raising the floor, measured rather than assumed (hub dashboard, 2026-07-29): | Customer | Controller | Effect of floor → 0.185.1 | |---|---|---| | Demo Ügyfél | 0.185.1 | none — already at it | | Demo HP | 0.185.1 | none — already at it | | Peti Proxmox | 0.115.0, **DOWN 14d** | already below the current 0.156.0 floor, so it updates on return either way; only the target version changes | --- ## 4. Phase 0 gate answers (all four completed before the stop) ### 4.1 Storage placement — gate PASSES, fence confirmed `pvesm status` on demo-hp, 2026-07-29 (BEFORE; there is no AFTER — no VM was created): ``` Name Type Status Total (KiB) Used (KiB) Available (KiB) % felhom-backup dir active 983379700 2293856 931059232 0.23% felhom-pbs pbs active 0 0 0 0.00% local dir active 40516856 14721460 23705004 36.33% local-lvm lvmthin active 56545280 21922605 34622674 38.77% ``` `local-lvm` is thin, pool `<53.93g`, **data 38.77 %** — matching the runbook's "~54 GB pool, 38.8 %" to 0.03 pp, so the picture has **not** materially changed. Allocated LVs on it: `vm-9201-disk-{0,1,2}` (32+50+20 G, the **live customer guest**), `vm-300-disk-{0,1}` (drill-r50, 32 G) and its `r50pre` snapshots (32 G) ⇒ ~166 G allocated against 53.93 G real. `vgs` shows only 14.75 G VFree. **The fence holds: nothing goes on `local-lvm`.** `/mnt/nvme-1tb` = `/dev/nvme0n1`, ext4, 938 G, **888 G free**. ≥100 G requirement satisfied. **Placement decision (recorded deliberately, per §3.1):** the only storage with `content images` is `local-lvm` (forbidden). `felhom-backup` is a `dir` at exactly `/mnt/nvme-1tb` but is `content backup` only **and is demo-hp's live backup target** — widening its content set would mutate the very storage E-2c/E-2 role logic keys on, which §0/§9 forbid. So a *new* dir storage would have been required. Both remaining shapes carry a cost and the choice was **not** forced this run because the run stopped: a second storage at the same mountpoint risks perturbing the agent's role resolution on a live box; a storage at a **subdirectory** fails `exactMount` and reports `disconnected` in the host report, which is not purely cosmetic since it can reach the hub's storage monitor. **This decision is left open and is an input to any resumed run.** ### 4.2 Drive-gate cadence — answered - Symbol: `driveGateLoop` → `s.ReconcileDriveGates()` - Definition: `controller/internal/web/intermediary.go:328` (gate itself at `:267`) - Registration: `controller/internal/web/server.go:217` — `go s.driveGateLoop()` - **Interval: `time.NewTicker(30 * time.Second)`** (`intermediary.go:337`) ⇒ Stage 7's derived budget would be **two cycles = 60 s** of controller-side reconcile, *plus* the agent's own `/disks` refresh, since `ReconcileDriveGates` consumes `resp.Disks` (`:281`). Recorded so a resumed run does not have to re-derive it. ### 4.3 Hot-detach — feasible, not exercised demo-hp: AMD `svm` present, `/sys/module/kvm_amd/parameters/nested = 1`, 8 cores, 29 GB RAM (25 GB available). The working nested-PVE reference is VM 300 (`drill-r50`): `bios: ovmf`, `efitype=4m,pre-enrolled-keys=0` (Secure Boot OFF, as the mkimage loader requires), `machine: q35`, `scsihw: virtio-scsi-single`, `cpu: host`. `virtio-scsi-single` + PVE's default `hotplug: disk,network,usb` supports SCSI hot-detach, so the **live-transition** variant of Stage 7 was available. Not exercised. ### 4.4 Install route — ISO available, route selected, not used Two generic (pairing) ISOs are present on demo-hp `local`: `felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso` (**preferred**, the v1.25.0 train, mkimage loader for nested VMs) and `felhom-pve-9.2-1-v1.22.0-nested-canary-generic.iso`. The **ISO/PAIRING route was selected**. It reaches the same installer invocation — `run_pairing` (`felhom-bootstrap.sh:441`) falls through to `run_direct` (`:495-499`) in the same invocation, which fetches `$INSTALL_URL` (`:322-330`) and runs it (`:343`). Not used — the run stopped first. **Operator STOP: not required.** `HUB_PW` is present in `~/.config/credentials` and hub auth was verified working (`http://10.43.52.34:8080/` → 200 with `curl -u ":$HUB_PW"`; public ingress also 200), so CC could have performed the bind itself. The runbook's §5.1a STOP is retired for future runs. --- ## 5. Incidental observations (filed, not acted on) 1. **A stale unclaimed appliance is sitting in the hub**: `206c8838-7755-4751-8a16-a842853d718f`, pairing code `QWA-WJE`, MAC `bc:24:11:ea:55:d0`, "Standard PC (Q35 + ICH9, 2009)" on an AMD V1756B with 7.7 GB — i.e. a **nested VM on demo-hp**, first and last seen **5 d ago**. It is a leftover from the 2026-07-25 ISO work and has never been bound or discarded. A resumed run must distinguish its own appliance from this one, or discard it first. 2. **The agent has no container image** (`/v2/admin/felhom-agent/tags/list` → 0 tags). It ships only as a Gitea generic package, which is why the publish gap is invisible from the registry. 3. `drill-r50` (VM 300) still holds a `r50pre` snapshot pair consuming ~32 G of allocation on the over-subscribed `local-lvm`. Untouched per §9; noted because it is part of why the pool is tight. --- ## 6. What did not happen No VM was created. No ISO booted. No install ran. No appliance registered, bound or discarded. No storage was created or modified on demo-hp. No drive attached or detached. No hub setting changed (the floor was **read**, not written). No teardown was needed. **C1–C5 are all UNPROVEN**, and the E-2 / E-2d rows are unchanged except to record this blocker.