Files
felhom.eu/documentation/audits/E2D-fresh-vm-2026-07-29.md
T
admin f3975cf5bc E-2d executed on a fresh box: C1/C2 proven, C3/C4 partial, C5 FAILS — R-112/113/114
Full ISO/PAIRING route on a nested PVE VM on demo-hp, after R-111 was fixed
earlier in the session. Bind -> running controller in 3m35s. The install fetched
the artifacts published an hour before and restored the golden baked 20 minutes
before, so the publish train is proven end to end on a real install.

C1 PROVEN: "felhom-host-install v1.22.0", "Day-0 provision SUCCESS", guest 9201
running, bootstrap unit wrote its done-flag and self-disabled. This retires E-2's
"installer-logic-tested, not install-tested".

C2 PROVEN: both DEGRADED warning lines verbatim, backup.local_backup_target=local,
no felhom-backup storage created, and the install did not abort.

C3/C4 PARTIAL and C5 FAILED — three findings, none fixed:

R-112 (P1): E-2's degraded banner and offer have NO UI CONSUMER. The endpoint
returns byte-exact copy; grep 'backup-target' across every html/js/css is 0 hits
and no page handler injects the state. Templates fetch 18 distinct /api/storage/*
endpoints; these two are the only ones with zero references. v0.185.1 fixed the
router mount and stopped one layer short of the render. Fifth instance of
seam-built-but-never-wired.

R-113 (P1): the drive-absent gate CANNOT FIRE on device loss. planDriveGates
reads presence from BoundUnderParent = "is this path in the guest's mountinfo".
The raw mount is a device-bound systemd unit and dies with the device; the
agent's own bind is not device-bound and outlives it, so the gate sees "present"
forever. Live: agent reported the drive absent every 20s for 4.5 minutes, the
controller logged 0 [gate] lines, the hub received zero events -- neither
backup_target_absent nor the generic storage_disconnected. Sixth instance of the
class: E-2b wired the seam to a condition that cannot occur.

R-114: on target-drive loss the message claims the backup is on the system disk
(false) and offers the drive that just vanished. Invisible only because of R-112,
so it must be fixed BEFORE R-112 is wired.

Also filed as a second instance under R-110 rather than a new ID: host-install
fetches nine files from raw/branch/main and the hub vouches a sha for one;
E-2a's wrapper is installed 0755 to /usr/local/sbin, root-fenced in sudoers,
validated only by bash -n.

C4 is fully proven at API level: decline path (registration confers no role),
restart_required:true, agent did NOT self-restart (in-flight check performed and
recorded first), E-2a wrapper created the storage at the drive's own mountpoint,
and healthy renders nothing.

Teardown: VM destroyed, scratch storage removed, pvesm status after == before
(local-lvm 38.77%), guest 9201 and drill-r50 untouched. Hub records for e2d-fresh
remain -- delete correctly refused at four gates, finally "host is ONLINE";
deletable once it ages to DOWN. Command recorded in OPEN-ITEMS.md.

capability-map NOT touched: the customer-facing legs are broken rather than
proven, and the map has no E-2 rows at all.
2026-07-29 13:13:56 +02:00

15 KiB
Raw Blame History

E2D-fresh-vm-2026-07-29 — E-2 proven on a fresh box: C1/C2 pass, C3/C4 partial, C5 FAILS

Run: RUNBOOK-e2d-fresh-vm-2026-07-29.md, executed by CC on DooPlex, 2026-07-29. Preceded by: a Phase 0 STOP earlier the same day (R-111 — the Day-0 channel was 17 agent releases stale). R-111 was fixed first; this run then proceeded on the real customer path.

Headline: the installer and its Case B are proven on a real install. The two customer-facing halves of E-2 are not reachable by a customer at all, and the drive-absent alarm cannot fire on device loss. Both were invisible to a green unit suite and to an API-level check; only the live run found them.

Claim Verdict
C1 host-install 1.22.0 completes a real install, rc=0 PROVEN
C2 Case B fires naturally on a single-drive box PROVEN
C3 the degraded banner renders to a customer ⚠️ PARTIAL — API exact, NO UI CONSUMER (R-112)
C4 the offer appears and moves the target when accepted ⚠️ PARTIAL — full API flow proven; offer equally invisible (R-112)
C5 backup_target_absent fires end to end FAILED — no event on any channel (R-113)

1. Baselines as actually confirmed

Artifact Confirmed Source
hub 0.81.0 manifests/hub.yaml:128; live deploy image
agent 0.113.0 main @ 58b598b; published this run, sha 5f3247f7…
golden 0.185.1 baked this run, sha dba00f3e…, embeds controller 0.185.1
controller 0.185.1 main @ cdaeb36; live in the guest
host-install 1.22.0 scripts/felhom-host-install.sh:187; fetched from the website at run time
felhom.eu 3dff357 main HEAD at run time

Route: ISO/PAIRING (the real customer chain). ISO felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso. Every C1C5 result is on the real route; no manual-installer fallback was used.

Operator STOP: not required and now retired. HUB_PW is in ~/.config/credentials; CC created the customer and performed the bind itself. The one human step that was needed is new — see §6.

2. Timeline (VM 9300 e2d-fresh on demo-hp, nested PVE)

UTC Event
10:29 VM created — q35/OVMF SB-off, 4c/8G, one 160 G disk, hotplug disk, outside the felhom pool
10:34:52 PVE auto-install done, first boot, DHCP 192.168.0.125 (no R-59 gate trip)
10:34:53 registered as unclaimed appliance, pairing code SB4-7ZK, console banner rendered
10:36:52 bound to customer e2d-fresh by CC; credentials delivered 2 s later
10:37:31 host enrolled e2d-fresh-ac9f09; break-glass root credential vaulted
10:37:33 artifact manifest served: agent=0.113.0 golden=0.185.1
10:40:27 controller_started (0.185.1)bind → running controller in 3 m 35 s
10:56 second 100 G disk hot-attached; wizard init→mount→register
10:57:08 offer accepted → restart_required:true; agent restarted at 10:57:36
10:58:37 target drive hot-detached (volume survives as unused0)
10:58:3711:03 no event on any channel for 4½ minutes (budget was 60 s)
11:03:31 reattached; state returns healthy; still no event

3. C1 — PROVEN

From the felhom-bootstrap.service journal on the box:

felhom-bootstrap: fetching host-install: https://felhom.eu/scripts/felhom-host-install.sh
[INFO] felhom-host-install v1.22.0 — mode=appliance customer=e2d-fresh vmid=9201
...
[OK]   controller: Up 21 seconds (healthy)
[INFO]   controller image: gitea.dooplex.hu/admin/felhom-controller:0.185.1
[OK] Day-0 provision SUCCESS — vmid=9201 host_id=e2d-fresh-ac9f09 customer=e2d-fresh
     golden=local:backup/vzdump-lxc-9100-2026_07_29-12_37_56.tar.zst
felhom-bootstrap: host-install SUCCESS — writing done-flag, disabling unit, scrubbing secrets

rc=0 is corroborated structurally: the unit wrote its done-flag, self-disabled, and Deactivated successfully. pct list showed guest 9201 e2d-fresh running.

The publish train is proven end to end: the golden restored is vzdump-lxc-9100-2026_07_29-12_37_56 — the golden baked ~20 minutes earlier in the same session. This retires E-2's "installer-logic-tested, not install-tested".

4. C2 — PROVEN

Both required warning lines, verbatim, ANSI-stripped:

[WARN]   backup target: DEGRADED — no eligible second drive, so the whole-system backup stays on the SYSTEM drive.
[WARN]   It protects against file corruption but NOT against a disk failure. Attach a second drive and assign it in the dashboard.
  • agent.jsonbackup.local_backup_target = 'local' (the "resolved local" observable)
  • no felhom-backup storage created at install
  • the install did not abort — a single-drive appliance is a valid product

5. THE FINDINGS

5.1 R-112 — the degraded banner and the offer have NO UI consumer (customer-invisible)

The endpoint is correct. Authenticated GET /api/storage/backup-target returned, exactly:

{"degraded":true,"known":true,"target":"local",
 "message":"A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok
            ellen véd, lemezhiba ellen nem. Csatlakoztass egy második meghajtót a teljes védelemhez."}

byte-identical to the runbook's required copy. known is a field separate from degraded, so UNKNOWN genuinely cannot render as degraded (R-88 Part 2's lesson held).

And nothing in the product ever asks for it. Negative claims with their search scope:

Search Result
grep -rn 'backup-target' --include='*.html' --include='*.js' --include='*.css' controller/ 0 hits
grep -rn 'kijelölheted|ugyanazon a lemezen|OfferPath|OfferLabel|Degraded' controller/internal/web/templates/ 0 hits
consumers of backupTargetDegradedText / backupTargetOfferText only degradedMessageFor (:136) and the JSON handler (:153) — both inside backup_target_offer.go
consumers of resolveBackupTargetState / degradedMessageFor across all Go only the API handler. No page handler injects the state.

The decisive contrast: the templates fetch **18 distinct /api/storage/* endpoints**. backup-targetandbackup-target/assign` are the only two referenced by zero templates.

The handler's own doc comment reads "serves GET /api/storage/backup-target — the dashboard's source for the degraded banner and the offer" — an invariant comment asserting a consumer that does not exist (CLAUDE.md's "a comment asserting an invariant needs a test pinning it, or it is a wish", instance #7). And v0.185.1's own test, TestBackupTargetRoutesLiveUnderTheStorageAPIMount, pins that the router dispatches the path — not that anything renders it. v0.185.1 shipped as "the offer endpoints were mounted where nothing routed to them": it fixed the mount and stopped one layer short.

Fifth instance of the seam-built-but-never-wired class. CLAUDE.md's seam rule names exactly this: "a feature is not shipped until its entry point is reachable… handler tests that POST directly prove nothing about reachability."

5.2 R-113 — the drive-absent gate CANNOT fire on device loss (E-2b's alarm is unreachable)

Detached the assigned target drive at 10:58:37Z under a running agent. Over the next 4½ minutes:

  • agent, every 20 s: storage: enrolled drive absent by UUID — not re-asserting / reconcile: enrolled drive not present (durable-id absent) — skippingthe agent knows
  • controller: docker logs | grep -c '\[gate\]'0. The gate never acted, ABSENT or RETURNED
  • hub: zero events for the customer across the whole window — no backup_target_absent, and no generic storage_disconnected either

Root cause, established at source and confirmed live. planDriveGates (intermediary.go:216-262) computes presence as present[GuestPath] = present[GuestPath] || d.BoundUnderParent, and the agent derives BoundUnderParent from GuestSeesMount()"does the guest's /proc/<pid>/mountinfo list this path as a mount target" (localapi/disks.go:210, localapi/intermediary.go GuestSeesMount).

Measured on the box with the device removed:

-- raw mount /mnt/mentes2 --                NOT mounted          <- systemd device-bound unit, died with the device
-- stable bind /mnt/felhom-drives/mentes2 -- /dev/sdb[/felhom-data] ext4   <- the agent's MANUAL bind, SURVIVES

The raw mount is a device-bound systemd unit and dies correctly; the agent's own bind under the shared parent is not device-bound, so its mountinfo entry outlives the device. The gate reads that surviving entry as "present" ⇒ !present[...] is never true ⇒ no Stop action ⇒ notifyDriveAbsent is never called.

This is not a virtualisation artefact: the asymmetry is between a device-bound mount and a manual bind, which is identical on physical hardware. Caveat kept honest: proven on a SCSI hot-detach; a physical USB unplug was not staged.

Consequence. E-2b's celebrated fix — "THE SEAM THAT WAS NEVER WIRED… a drive that is ONLY a backup target has no apps to stop, so it was silent twice over" (intermediary.go:288-296) — wired the notify to a branch that cannot execute on device loss. The seam is wired; the condition is unreachable. Sixth instance of the class, one layer deeper than the fifth.

Mirror scenario: not separately staged, and it does not need to be — both the specific and the generic event are emitted from the same a.Stop branch, which never executed. The generic storage_disconnected is equally unreachable by this path. Recorded as reasoned, not observed.

5.3 R-114 — on target-drive loss the customer is told the wrong story and offered the missing drive

While the target drive was absent, the endpoint returned:

{"degraded":true,"target":"felhom-backup",
 "message":"A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer …",
 "offer_path":"/mnt/felhom-drives/mentes2","offer_label":"Mentés meghajtó"}

Two defects in one payload: the message claims the backup is on the system disk, which is false — the target is felhom-backup on a drive that has vanished; and the remedy offered is the drive that just disappeared. resolveBackupTargetState falls through to the generic degraded branch whenever no disk satisfies d.BackupTarget && d.MountPath != "", without distinguishing never configured from configured and now missing.

Interaction worth stating: R-114 is currently invisible only because of R-112. Fixing R-112 alone — wiring the banner — would immediately start showing customers this wrong message. They must be fixed together, R-114 first.

Also observed: after reattach the drive returned as /dev/sdc, while the stable bind still recorded /dev/sdb[/felhom-data]. The state read healthy (degraded:false) with the guest-visible bind still naming the dead device node. Not chased further; recorded as part of R-113's shape.

5.4 Smaller findings (recorded, not filed as their own IDs)

  1. A "hard min" that only warns. [WARN] local-lvm free ~83 GiB < hard min 120 GiB — the installer names a hard minimum and proceeds. Either it is not hard, or the wording is wrong.
  2. felhom-backup-target-apply is fetched unvouched. host-install pulls nine files from raw/branch/main (:2072:2206); the hub manifest vouches a sha for exactly one (wrapper_sha256felhom-pbs-apply, verified this run: no drift). E-2a's wrapper is installed 0755 to /usr/local/sbin and root-fenced in sudoers, validated only by bash -n. Filed as a second instance under R-110, not a new ID — same class (a root-executed artifact taken from main with no pinned integrity).

6. C4 — what IS proven (API level)

Everything except customer reachability:

  • The offer appeared with all three fields: offer_path=/mnt/felhom-drives/mentes2, offer_label=Mentés meghajtó, offer_message=Ezt a meghajtót kijelölheted a rendszermentés helyéül…
  • The decline path (§6.4) — PROVEN. After registering the drive and not accepting: target still local, no felhom-backup storage, agent.json unchanged. Registration does not confer a role — E-2 §3's central invariant, live.
  • Accept{"assigned":"/mnt/felhom-drives/mentes2","restart_required":true}
  • The agent did NOT self-restartActiveEnterTimestamp unchanged at 12:37:53 CEST twenty minutes later. In-flight check performed and recorded before restarting: 0 running PVE tasks, no vzdump process, 0 backup lines in the agent journal.
  • The E-2a root-fenced wrapper worked on a fresh box: created dir: felhom-backup, path /mnt/mentes2, is_mountpoint 1 — the drive's own mountpoint (law F-2).
  • Healthy renders nothing: after the restart the payload is {"degraded":false,"known":true,"label":"Mentés meghajtó","target":"felhom-backup"}no message field at all. No "backup protected" reassurance (E-2 Scenario E).

7. E-2d's own premise needs amending

The runbook assumed a fresh install yields a claimable customer CC can then drive. It does not: the claim code is bcrypt-hashed in the hub and only ever emailed (claim/engine.go:58-95), and the gate covers everything except /claim, /claim/request-new-code, /api/health, /static/ (claim.go:222-229). C3/C4/C5 all sit behind it. This run cleared it by registering an operator email, resending, and having the operator relay the code — the one genuine human step, and it also proved the claim flow end to end (code → password → 401 "dashboard not yet claimed" becoming 401 "authentication required", a positive discriminator).

8. Teardown

  • VM 9300 destroyed --purge; e2d-images storage removed; scratch dir removed.
  • pvesm status after == before: local-lvm 38.77 %, felhom-backup 931059224 KiB available — byte-identical to the pre-run measurement. Freed space returned.
  • Guest 9201 (live customer) untouched and running; drill-r50 VM 300 untouched.
  • Drill VM (golden bake) torn down per GL-1 earlier: guest 9100 purged, secrets shredded, disk restored to virgin, token-leak grep 0.

Hub records NOT yet removed — stated, not silent. e2d-fresh + host e2d-fresh-ac9f09 remain. The delete was attempted and correctly refused at four successive gates: acknowledgements → typed confirm_idexpect_hosts stale-preview → finally host e2d-fresh-ac9f09 is ONLINE. The host is online only because its last report is recent; the VM is destroyed, so it ages OK → WARN (30 m) → DOWN (>1 h) and is then deletable. Cleanup command, once it reads DOWN:

POST /configs/e2d-fresh/delete  ack_hosts=1 ack_reset=1 ack_purge=1 confirm_id=e2d-fresh expect_hosts=1

Tracked in OPEN-ITEMS.md. Deleting rather than keeping is deliberate — R-93 records what a half-real fixture costs.

Also still present: the stale unclaimed appliance from 2026-07-25 (206c8838…, code QWA-WJE). Not mine; not discarded. A future run must distinguish its own appliance from it.