Files
felhom.eu/REPORT-e2d.md
T
admin d839ddcb60 E-2d teardown complete: drill customer + host removed from the hub
The delete was correctly refused at four successive gates while the host still
read ONLINE (acknowledgements -> typed confirm_id -> expect_hosts stale-preview
-> "host is ONLINE"). Rather than force it, the run waited for the destroyed
host to age to DOWN; delete-impact then reported deletable:true and the
documented cascade ran:

  host deleted (escrow demoted to retained custody), tenantsync deprovisioned,
  PBS tenancy deprovisioned, claim reset to unclaimed, residue purged
  (reports=5 app_telemetry=5 notif_prefs=1 appliance_registrations=1)

Verified after: 0 occurrences of "e2d" anywhere on the hosts page; demo-felhom
and demo-hp ONLINE on agent 0.113.0; drill-r50 and peti-felhom unchanged;
demo-hp carries only guest 9201 and VM 300.

Scoping checked rather than assumed: the single purged appliance_registration
was this run's own appliance (810d10c5, bound to e2d-fresh). The unrelated stale
2026-07-25 appliance (206c8838 / QWA-WJE) was NOT touched by the cascade — the
operator removed it separately.

- OPEN-ITEMS.md: the drill-cleanup WATCHING row is removed (done, not open).
- audits/E2D-fresh-vm-2026-07-29.md §8 + REPORT-e2d.md: teardown recorded as
  complete, with the cascade output and the appliance-scoping note.
2026-07-29 15:56:49 +02:00

5.8 KiB

REPORT — R-111 fixed, then E-2 proven on a fresh box (2026-07-29)

Two phases in one session. Full evidence: documentation/audits/E2D-fresh-vm-2026-07-29.md. Root REPORT.md untouched.

Phase 1 — R-111: the Day-0 channel now serves the current software

A Phase 0 gate earlier the same day stopped the E-2d run before any VM existed: a fresh box would have installed agent 0.96.0 + controller 0.161.0, ~17 and ~24 releases behind main.

Before Now
agent (Gitea generic) 0.96.0 0.113.0, sha 5f3247f7…, round-trip verified
golden (Gitea generic) 0.161.0 0.185.1, sha dba00f3e…, embeds controller 0.185.1
hub min_agent 0.93.0 0.113.0 (what controller v0.185.0 declares)

Bake clean on every marker: Result=success, overlay2, all three mounts in the archive, 0 FATAL/exclusions, HTTP 201, token-leak grep 0. GL-1 teardown: guest 9100 purged, secrets shredded, drill disk restored to virgin. Agent + golden moved in one manifest POST so it never vouched a new agent against an old golden. min_agent verified zero-impact first (all three enrolled hosts already at 0.113.0). Global floor deliberately not raised — the golden now bakes 0.185.1.

Commit 3dff357.

Phase 2 — the E-2d run, full ISO/PAIRING route

Nested PVE VM on demo-hp, one disk, outside the felhom pool. Bind → running controller in 3 m 35 s. The install fetched exactly the artifacts published an hour earlier and restored vzdump-lxc-9100-2026_07_29-12_37_56 — the golden baked 20 minutes before. The publish train is proven end to end on a real install.

Claim Verdict
C1 host-install 1.22.0 completes a real install, rc=0 PROVEN
C2 Case B fires naturally PROVEN — both DEGRADED lines verbatim, local_backup_target=local, install did not abort
C3 degraded banner renders to a customer ⚠️ PARTIAL — API byte-exact; no UI consumer existsR-112
C4 offer appears and moves the target ⚠️ PARTIAL — decline path, restart_required:true, no self-restart, E-2a wrapper, healthy-renders-nothing all PROVEN at API level; offer equally invisible → R-112
C5 backup_target_absent end to end FAILED — zero events on any channel → R-113

The three findings

R-112 (P1) — E-2's banner and offer have no UI consumer. The endpoint returns byte-exact copy; grep 'backup-target' across every *.html/*.js/*.css0 hits, and no page handler injects the state. Decisive contrast: templates fetch 18 distinct /api/storage/* endpoints; backup-target and backup-target/assign are the only two with zero references. v0.185.1 fixed the router mount and stopped one layer short of the render. Fifth instance of seam-built-but-never-wired.

R-113 (P1) — the drive-absent gate cannot fire on device loss. planDriveGates reads presence from BoundUnderParent = "is this path in the guest's mountinfo". The raw mount is a device-bound systemd unit and dies; the agent's own bind is not device-bound and outlives the device, so the gate sees "present" forever. Live: agent said enrolled drive absent by UUID every 20 s for 4½ minutes, controller logged 0 [gate] lines, hub got zero events — neither the specific nor the generic one. Sixth instance of the class, one layer deeper: E-2b wired the seam to a condition that cannot occur.

R-114 — on target-drive loss the message says the backup is "on the same disk as the system" (false) and offers the drive that just vanished. Invisible today only because of R-112 — so R-114 must be fixed before R-112 is wired.

Also filed as a second instance under R-110 (not a new ID): host-install fetches nine files from raw/branch/main and the hub vouches a sha for one; E-2a's wrapper is installed 0755 to /usr/local/sbin, root-fenced in sudoers, validated only by bash -n.

Record

  • OPEN-ITEMS.mdR-112/R-113/R-114 opened; E-2d re-stated with results and left open for the residue; E-2's "NOT yet live-proven" list resolved into proven / known-broken; R-94 fully unblocked; R-110 extended. The drill-cleanup row was opened and then closed the same session once the teardown completed, so it is not carried in the register.
  • ROADMAP.md — R-112/R-113/R-114 under P1; R-111 marked SHIPPED.
  • architecture/00-capability-map.md not touched — for two reasons: the customer-facing legs are broken rather than proven, and the map has no E-2 / backup-target rows at all (worth noting against the ROADMAP's coupling rule).

Teardown

VM destroyed, scratch storage removed, pvesm status after == before (local-lvm 38.77 %, byte-identical), guest 9201 and drill-r50 untouched. Hub records removed — teardown complete. The delete was correctly refused at four gates while the host still read ONLINE; once the destroyed host aged to DOWN (delete-impactdeletable:true) the documented cascade ran and completed: host deleted, PBS tenancy deprovisioned, claim reset, residue purged. Verified after: 0 e2d occurrences on the hosts page, fleet unchanged. The one purged appliance_registrations=1 was this run's own appliance; the unrelated stale 2026-07-25 appliance (206c8838…) was not touched by the cascade — the operator removed it separately.

One human step, and a premise correction

The runbook's §5.1a operator STOP (the bind) is retired — CC did it. But E-2d's premise that a fresh install yields a CC-drivable claimable customer is wrong: the claim code is bcrypt-hashed and email-only, and the gate covers everything except /claim, /api/health, /static/. One operator relay of the emailed code was required — which also proved the claim flow end to end.