Files
felhom.eu/documentation/audits/E2D-fresh-vm-2026-07-29.md
admin 952ebf4862 Record work, banked first: shrink the E-2d row, create the missing capability-map rows
Unconditional and three sessions overdue, so it commits before any code is
touched — E-2d itself stopped at Phase 0 and banked nothing.

E-2d row: 822 words -> 121, and the contradiction resolved. Its State read
CLOSED — PARTIALLY PROVEN while the cell's final sentence read "This row stays
OPEN only for the residue"; a reader could not tell which. It is CLOSED, with
R-116 the single named open leg.

Nothing unique was binned. Three facts existed ONLY in that cell and are moved
into audits/E2D-fresh-vm-2026-07-29.md as a new §1a: the local-lvm fence figures
with the 888 GB nvme alternative, the exactMount subdirectory caveat and why the
subdirectory is nonetheless the safe placement (no durable_id collision), and
the ISO/PAIRING -> DIRECT fall-through derived at source with its line
citations. drill-r50's blocked status was already in both audits.

Capability map: it had ZERO rows for the backup-target work — grep gives 0 hits
for backup_target and one for "E-2" that is a campaign date string. Three
scenario rows added, at today's honest status, not the value hoped for later:

  C. Protection & recovery — installer Case A/B, DEGRADED recorded not hidden
     PROVEN-LIVE, cites E2D-fresh-vm C1+C2
  D. Storage & devices — the offer, and that registration confers no role
     PROVEN-LIVE, cites SESSION-C C4 + the decline path
  F. Notifications & monitoring — the absent-target alarm and its pairing
     PARTIAL, cites SESSION-C C5, leg named, -> R-116

Row F is PARTIAL today per the doc's own strict enum (a leg not exercised live
is PARTIAL with the leg named, never PROVEN-LIVE). A later session may flip it;
this commit must not.
2026-07-29 23:34:06 +02:00

18 KiB
Raw Permalink Blame History

E2D-fresh-vm-2026-07-29 — E-2 proven on a fresh box: C1/C2 pass, C3/C4 partial, C5 FAILS

Run: RUNBOOK-e2d-fresh-vm-2026-07-29.md, executed by CC on DooPlex, 2026-07-29. Preceded by: a Phase 0 STOP earlier the same day (R-111 — the Day-0 channel was 17 agent releases stale). R-111 was fixed first; this run then proceeded on the real customer path.

Headline: the installer and its Case B are proven on a real install. The two customer-facing halves of E-2 are not reachable by a customer at all, and the drive-absent alarm cannot fire on device loss. Both were invisible to a green unit suite and to an API-level check; only the live run found them.

Claim Verdict
C1 host-install 1.22.0 completes a real install, rc=0 PROVEN
C2 Case B fires naturally on a single-drive box PROVEN
C3 the degraded banner renders to a customer ⚠️ PARTIAL — API exact, NO UI CONSUMER (R-112)
C4 the offer appears and moves the target when accepted ⚠️ PARTIAL — full API flow proven; offer equally invisible (R-112)
C5 backup_target_absent fires end to end FAILED — no event on any channel (R-113)

1. Baselines as actually confirmed

Artifact Confirmed Source
hub 0.81.0 manifests/hub.yaml:128; live deploy image
agent 0.113.0 main @ 58b598b; published this run, sha 5f3247f7…
golden 0.185.1 baked this run, sha dba00f3e…, embeds controller 0.185.1
controller 0.185.1 main @ cdaeb36; live in the guest
host-install 1.22.0 scripts/felhom-host-install.sh:187; fetched from the website at run time
felhom.eu 3dff357 main HEAD at run time

Route: ISO/PAIRING (the real customer chain). ISO felhom-pve-9.2-1-v1.25.0-nested-vm-generic-mkimage.iso. Every C1C5 result is on the real route; no manual-installer fallback was used.

Operator STOP: not required and now retired. HUB_PW is in ~/.config/credentials; CC created the customer and performed the bind itself. The one human step that was needed is new — see §6.

1a. Phase 0 answers, preserved from the OPEN-ITEMS row

Moved here when the E-2d register row was rewritten (2026-07-29) — the row had grown to ~820 words and these were the facts that existed nowhere else. They are inputs to any future drill on this host, not narrative.

  • Storage fence. local-lvm on demo-hp is a thin pool, ~144 GB allocated against ~54 GB real, 38.8 % used, on a box running a live customer guest — a full thin pool corrupts every guest on it. local has only 23.7 GB and sits on pve-root. Use /mnt/nvme-1tb (888 GB free).
  • The exactMount caveat, and the placement decision it forces. A dir storage created at a SUBDIRECTORY of /mnt/nvme-1tb fails the agent's exactMount check and reports disconnected in the host report. Both E-2d and Session C accepted that: hub-side it is a WARN log line only — no event, no email — and the alternative (a second storage at the live backup target's own mountpoint) risks perturbing the drive-role resolution on a production box. The agent deliberately falls back to a stable store id rather than borrowing the nvme's fs-UUID in this case, so there is no durable_id collision with felhom-backup; that is what makes the subdirectory the safe choice.
  • The ISO/PAIRING → DIRECT fall-through, derived at source. A fresh VM with no baked customer-id lands in PAIRING mode (scripts/iso/felhom-bootstrap.sh:537-541), not DIRECT (:312), and only DIRECT passes --customer-id / --mode / --passphrase-file. On a 200 from /api/v1/appliance/poll the pairing loop writes the hub-delivered credentials into the 0600 env, re-sources it and calls run_direct in the same invocation (:495-499), which is the single site that fetches $INSTALL_URL (:322-330), builds the args (:334) and invokes bash "$SCRIPT_TMP" "${args[@]}" (:343). So the ISO route reaches the identical installer invocation and yields a claimable customer — which is why it is the spine and no manual 1.22.0 run is needed as a separate scenario.
  • drill-r50 stays blocked. Unblocking it means the fixture stops representing anything real (R-93).

2. Timeline (VM 9300 e2d-fresh on demo-hp, nested PVE)

UTC Event
10:29 VM created — q35/OVMF SB-off, 4c/8G, one 160 G disk, hotplug disk, outside the felhom pool
10:34:52 PVE auto-install done, first boot, DHCP 192.168.0.125 (no R-59 gate trip)
10:34:53 registered as unclaimed appliance, pairing code SB4-7ZK, console banner rendered
10:36:52 bound to customer e2d-fresh by CC; credentials delivered 2 s later
10:37:31 host enrolled e2d-fresh-ac9f09; break-glass root credential vaulted
10:37:33 artifact manifest served: agent=0.113.0 golden=0.185.1
10:40:27 controller_started (0.185.1)bind → running controller in 3 m 35 s
10:56 second 100 G disk hot-attached; wizard init→mount→register
10:57:08 offer accepted → restart_required:true; agent restarted at 10:57:36
10:58:37 target drive hot-detached (volume survives as unused0)
10:58:3711:03 no event on any channel for 4½ minutes (budget was 60 s)
11:03:31 reattached; state returns healthy; still no event

3. C1 — PROVEN

From the felhom-bootstrap.service journal on the box:

felhom-bootstrap: fetching host-install: https://felhom.eu/scripts/felhom-host-install.sh
[INFO] felhom-host-install v1.22.0 — mode=appliance customer=e2d-fresh vmid=9201
...
[OK]   controller: Up 21 seconds (healthy)
[INFO]   controller image: gitea.dooplex.hu/admin/felhom-controller:0.185.1
[OK] Day-0 provision SUCCESS — vmid=9201 host_id=e2d-fresh-ac9f09 customer=e2d-fresh
     golden=local:backup/vzdump-lxc-9100-2026_07_29-12_37_56.tar.zst
felhom-bootstrap: host-install SUCCESS — writing done-flag, disabling unit, scrubbing secrets

rc=0 is corroborated structurally: the unit wrote its done-flag, self-disabled, and Deactivated successfully. pct list showed guest 9201 e2d-fresh running.

The publish train is proven end to end: the golden restored is vzdump-lxc-9100-2026_07_29-12_37_56 — the golden baked ~20 minutes earlier in the same session. This retires E-2's "installer-logic-tested, not install-tested".

4. C2 — PROVEN

Both required warning lines, verbatim, ANSI-stripped:

[WARN]   backup target: DEGRADED — no eligible second drive, so the whole-system backup stays on the SYSTEM drive.
[WARN]   It protects against file corruption but NOT against a disk failure. Attach a second drive and assign it in the dashboard.
  • agent.jsonbackup.local_backup_target = 'local' (the "resolved local" observable)
  • no felhom-backup storage created at install
  • the install did not abort — a single-drive appliance is a valid product

5. THE FINDINGS

5.1 R-112 — the degraded banner and the offer have NO UI consumer (customer-invisible)

The endpoint is correct. Authenticated GET /api/storage/backup-target returned, exactly:

{"degraded":true,"known":true,"target":"local",
 "message":"A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer — így hibás fájlok
            ellen véd, lemezhiba ellen nem. Csatlakoztass egy második meghajtót a teljes védelemhez."}

byte-identical to the runbook's required copy. known is a field separate from degraded, so UNKNOWN genuinely cannot render as degraded (R-88 Part 2's lesson held).

And nothing in the product ever asks for it. Negative claims with their search scope:

Search Result
grep -rn 'backup-target' --include='*.html' --include='*.js' --include='*.css' controller/ 0 hits
grep -rn 'kijelölheted|ugyanazon a lemezen|OfferPath|OfferLabel|Degraded' controller/internal/web/templates/ 0 hits
consumers of backupTargetDegradedText / backupTargetOfferText only degradedMessageFor (:136) and the JSON handler (:153) — both inside backup_target_offer.go
consumers of resolveBackupTargetState / degradedMessageFor across all Go only the API handler. No page handler injects the state.

The decisive contrast: the templates fetch **18 distinct /api/storage/* endpoints**. backup-targetandbackup-target/assign` are the only two referenced by zero templates.

The handler's own doc comment reads "serves GET /api/storage/backup-target — the dashboard's source for the degraded banner and the offer" — an invariant comment asserting a consumer that does not exist (CLAUDE.md's "a comment asserting an invariant needs a test pinning it, or it is a wish", instance #7). And v0.185.1's own test, TestBackupTargetRoutesLiveUnderTheStorageAPIMount, pins that the router dispatches the path — not that anything renders it. v0.185.1 shipped as "the offer endpoints were mounted where nothing routed to them": it fixed the mount and stopped one layer short.

Fifth instance of the seam-built-but-never-wired class. CLAUDE.md's seam rule names exactly this: "a feature is not shipped until its entry point is reachable… handler tests that POST directly prove nothing about reachability."

5.2 R-113 — the drive-absent gate CANNOT fire on device loss (E-2b's alarm is unreachable)

Detached the assigned target drive at 10:58:37Z under a running agent. Over the next 4½ minutes:

  • agent, every 20 s: storage: enrolled drive absent by UUID — not re-asserting / reconcile: enrolled drive not present (durable-id absent) — skippingthe agent knows
  • controller: docker logs | grep -c '\[gate\]'0. The gate never acted, ABSENT or RETURNED
  • hub: zero events for the customer across the whole window — no backup_target_absent, and no generic storage_disconnected either

Root cause, established at source and confirmed live. planDriveGates (intermediary.go:216-262) computes presence as present[GuestPath] = present[GuestPath] || d.BoundUnderParent, and the agent derives BoundUnderParent from GuestSeesMount()"does the guest's /proc/<pid>/mountinfo list this path as a mount target" (localapi/disks.go:210, localapi/intermediary.go GuestSeesMount).

Measured on the box with the device removed:

-- raw mount /mnt/mentes2 --                NOT mounted          <- systemd device-bound unit, died with the device
-- stable bind /mnt/felhom-drives/mentes2 -- /dev/sdb[/felhom-data] ext4   <- the agent's MANUAL bind, SURVIVES

The raw mount is a device-bound systemd unit and dies correctly; the agent's own bind under the shared parent is not device-bound, so its mountinfo entry outlives the device. The gate reads that surviving entry as "present" ⇒ !present[...] is never true ⇒ no Stop action ⇒ notifyDriveAbsent is never called.

This is not a virtualisation artefact: the asymmetry is between a device-bound mount and a manual bind, which is identical on physical hardware. Caveat kept honest: proven on a SCSI hot-detach; a physical USB unplug was not staged.

Consequence. E-2b's celebrated fix — "THE SEAM THAT WAS NEVER WIRED… a drive that is ONLY a backup target has no apps to stop, so it was silent twice over" (intermediary.go:288-296) — wired the notify to a branch that cannot execute on device loss. The seam is wired; the condition is unreachable. Sixth instance of the class, one layer deeper than the fifth.

Mirror scenario: not separately staged, and it does not need to be — both the specific and the generic event are emitted from the same a.Stop branch, which never executed. The generic storage_disconnected is equally unreachable by this path. Recorded as reasoned, not observed.

5.3 R-114 — on target-drive loss the customer is told the wrong story and offered the missing drive

While the target drive was absent, the endpoint returned:

{"degraded":true,"target":"felhom-backup",
 "message":"A rendszermentés jelenleg ugyanazon a lemezen van, mint a rendszer …",
 "offer_path":"/mnt/felhom-drives/mentes2","offer_label":"Mentés meghajtó"}

Two defects in one payload: the message claims the backup is on the system disk, which is false — the target is felhom-backup on a drive that has vanished; and the remedy offered is the drive that just disappeared. resolveBackupTargetState falls through to the generic degraded branch whenever no disk satisfies d.BackupTarget && d.MountPath != "", without distinguishing never configured from configured and now missing.

Interaction worth stating: R-114 is currently invisible only because of R-112. Fixing R-112 alone — wiring the banner — would immediately start showing customers this wrong message. They must be fixed together, R-114 first.

Also observed: after reattach the drive returned as /dev/sdc, while the stable bind still recorded /dev/sdb[/felhom-data]. The state read healthy (degraded:false) with the guest-visible bind still naming the dead device node. Not chased further; recorded as part of R-113's shape.

5.4 Smaller findings (recorded, not filed as their own IDs)

  1. A "hard min" that only warns. [WARN] local-lvm free ~83 GiB < hard min 120 GiB — the installer names a hard minimum and proceeds. Either it is not hard, or the wording is wrong.
  2. felhom-backup-target-apply is fetched unvouched. host-install pulls nine files from raw/branch/main (:2072:2206); the hub manifest vouches a sha for exactly one (wrapper_sha256felhom-pbs-apply, verified this run: no drift). E-2a's wrapper is installed 0755 to /usr/local/sbin and root-fenced in sudoers, validated only by bash -n. Filed as a second instance under R-110, not a new ID — same class (a root-executed artifact taken from main with no pinned integrity).

6. C4 — what IS proven (API level)

Everything except customer reachability:

  • The offer appeared with all three fields: offer_path=/mnt/felhom-drives/mentes2, offer_label=Mentés meghajtó, offer_message=Ezt a meghajtót kijelölheted a rendszermentés helyéül…
  • The decline path (§6.4) — PROVEN. After registering the drive and not accepting: target still local, no felhom-backup storage, agent.json unchanged. Registration does not confer a role — E-2 §3's central invariant, live.
  • Accept{"assigned":"/mnt/felhom-drives/mentes2","restart_required":true}
  • The agent did NOT self-restartActiveEnterTimestamp unchanged at 12:37:53 CEST twenty minutes later. In-flight check performed and recorded before restarting: 0 running PVE tasks, no vzdump process, 0 backup lines in the agent journal.
  • The E-2a root-fenced wrapper worked on a fresh box: created dir: felhom-backup, path /mnt/mentes2, is_mountpoint 1 — the drive's own mountpoint (law F-2).
  • Healthy renders nothing: after the restart the payload is {"degraded":false,"known":true,"label":"Mentés meghajtó","target":"felhom-backup"}no message field at all. No "backup protected" reassurance (E-2 Scenario E).

7. E-2d's own premise needs amending

The runbook assumed a fresh install yields a claimable customer CC can then drive. It does not: the claim code is bcrypt-hashed in the hub and only ever emailed (claim/engine.go:58-95), and the gate covers everything except /claim, /claim/request-new-code, /api/health, /static/ (claim.go:222-229). C3/C4/C5 all sit behind it. This run cleared it by registering an operator email, resending, and having the operator relay the code — the one genuine human step, and it also proved the claim flow end to end (code → password → 401 "dashboard not yet claimed" becoming 401 "authentication required", a positive discriminator).

8. Teardown

  • VM 9300 destroyed --purge; e2d-images storage removed; scratch dir removed.
  • pvesm status after == before: local-lvm 38.77 %, felhom-backup 931059224 KiB available — byte-identical to the pre-run measurement. Freed space returned.
  • Guest 9201 (live customer) untouched and running; drill-r50 VM 300 untouched.
  • Drill VM (golden bake) torn down per GL-1 earlier: guest 9100 purged, secrets shredded, disk restored to virgin, token-leak grep 0.

Hub records removed — teardown COMPLETE. The delete was first attempted while the host still read ONLINE and was correctly refused at four successive gates: acknowledgements → typed confirm_idexpect_hosts stale-preview → host e2d-fresh-ac9f09 is ONLINE. Rather than force it, the run waited for the destroyed host to age OK → WARN → DOWN, at which point delete-impact reported {"deletable":true,"status":"down"} and the documented cascade ran:

POST /configs/e2d-fresh/delete  ack_hosts=1 ack_reset=1 ack_purge=1 confirm_id=e2d-fresh expect_hosts=1
→ customer DELETE cascade started (journal #4, 1 host)
→ host e2d-fresh-ac9f09 deleted (escrow DEMOTED to retained custody)
→ tenantsync: deprovision ok (ns=e2d-fresh, existed=false); PBS tenancy deprovisioned
→ claim reset to unclaimed; residue purged (reports=5 app_telemetry=5 notif_prefs=1
  appliance_registrations=1)
→ customer DELETE cascade COMPLETE — full teardown

Verified after: 0 occurrences of e2d anywhere on the hosts page; demo-felhom + demo-hp ONLINE on agent 0.113.0; drill-r50 and peti-felhom unchanged; demo-hp carries only guest 9201 and VM 300. Deleting rather than keeping was deliberate — R-93 records what a half-real fixture costs.

Scoping note, checked rather than assumed: the cascade purged appliance_registrations=1 — that one was this run's own appliance (810d10c5…, bound to e2d-fresh). The unrelated stale appliance from 2026-07-25 (206c8838…, code QWA-WJE) was not touched by the cascade; the operator removed it separately. A future run must still distinguish its own appliance from any pre-existing one.