Files
felhom.eu/REPORT-e2d.md
T
admin d839ddcb60 E-2d teardown complete: drill customer + host removed from the hub
The delete was correctly refused at four successive gates while the host still
read ONLINE (acknowledgements -> typed confirm_id -> expect_hosts stale-preview
-> "host is ONLINE"). Rather than force it, the run waited for the destroyed
host to age to DOWN; delete-impact then reported deletable:true and the
documented cascade ran:

  host deleted (escrow demoted to retained custody), tenantsync deprovisioned,
  PBS tenancy deprovisioned, claim reset to unclaimed, residue purged
  (reports=5 app_telemetry=5 notif_prefs=1 appliance_registrations=1)

Verified after: 0 occurrences of "e2d" anywhere on the hosts page; demo-felhom
and demo-hp ONLINE on agent 0.113.0; drill-r50 and peti-felhom unchanged;
demo-hp carries only guest 9201 and VM 300.

Scoping checked rather than assumed: the single purged appliance_registration
was this run's own appliance (810d10c5, bound to e2d-fresh). The unrelated stale
2026-07-25 appliance (206c8838 / QWA-WJE) was NOT touched by the cascade — the
operator removed it separately.

- OPEN-ITEMS.md: the drill-cleanup WATCHING row is removed (done, not open).
- audits/E2D-fresh-vm-2026-07-29.md §8 + REPORT-e2d.md: teardown recorded as
  complete, with the cascade output and the appliance-scoping note.
2026-07-29 15:56:49 +02:00

86 lines
5.8 KiB
Markdown

# REPORT — R-111 fixed, then E-2 proven on a fresh box (2026-07-29)
Two phases in one session. Full evidence: `documentation/audits/E2D-fresh-vm-2026-07-29.md`.
Root `REPORT.md` untouched.
## Phase 1 — R-111: the Day-0 channel now serves the current software
A Phase 0 gate earlier the same day stopped the E-2d run before any VM existed: a fresh box would
have installed **agent 0.96.0 + controller 0.161.0**, ~17 and ~24 releases behind `main`.
| | Before | Now |
|---|---|---|
| agent (Gitea generic) | 0.96.0 | **0.113.0**, sha `5f3247f7…`, round-trip verified |
| golden (Gitea generic) | 0.161.0 | **0.185.1**, sha `dba00f3e…`, embeds controller 0.185.1 |
| hub `min_agent` | 0.93.0 | **0.113.0** (what controller v0.185.0 declares) |
Bake clean on every marker: `Result=success`, overlay2, **all three mounts in the archive**, 0
FATAL/exclusions, HTTP 201, token-leak grep 0. GL-1 teardown: guest 9100 purged, secrets shredded,
drill disk restored to `virgin`. Agent + golden moved in **one** manifest POST so it never vouched a
new agent against an old golden. `min_agent` verified zero-impact first (all three enrolled hosts
already at 0.113.0). Global floor deliberately **not** raised — the golden now bakes 0.185.1.
Commit `3dff357`.
## Phase 2 — the E-2d run, full ISO/PAIRING route
Nested PVE VM on demo-hp, one disk, outside the `felhom` pool. Bind → running controller in
**3 m 35 s**. The install fetched exactly the artifacts published an hour earlier and restored
`vzdump-lxc-9100-2026_07_29-12_37_56` — the golden baked 20 minutes before. The publish train is
proven end to end on a real install.
| Claim | Verdict |
|---|---|
| **C1** host-install 1.22.0 completes a real install, rc=0 | ✅ **PROVEN** |
| **C2** Case B fires naturally | ✅ **PROVEN** — both DEGRADED lines verbatim, `local_backup_target=local`, install did not abort |
| **C3** degraded banner renders **to a customer** | ⚠️ **PARTIAL** — API byte-exact; **no UI consumer exists****R-112** |
| **C4** offer appears and moves the target | ⚠️ **PARTIAL** — decline path, `restart_required:true`, no self-restart, E-2a wrapper, healthy-renders-nothing all PROVEN at API level; offer equally invisible → **R-112** |
| **C5** `backup_target_absent` end to end | ❌ **FAILED** — zero events on any channel → **R-113** |
## The three findings
**R-112 (P1)** — E-2's banner and offer have **no UI consumer**. The endpoint returns byte-exact copy;
`grep 'backup-target'` across every `*.html`/`*.js`/`*.css`**0 hits**, and no page handler injects
the state. Decisive contrast: templates fetch **18** distinct `/api/storage/*` endpoints;
`backup-target` and `backup-target/assign` are the only two with zero references. v0.185.1 fixed the
router mount and stopped one layer short of the render. Fifth instance of seam-built-but-never-wired.
**R-113 (P1)** — the drive-absent gate **cannot fire on device loss**. `planDriveGates` reads presence
from `BoundUnderParent` = "is this path in the guest's mountinfo". The raw mount is a device-bound
systemd unit and dies; **the agent's own bind is not device-bound and outlives the device**, so the
gate sees "present" forever. Live: agent said `enrolled drive absent by UUID` every 20 s for 4½
minutes, controller logged **0** `[gate]` lines, hub got **zero** events — neither the specific nor the
generic one. Sixth instance of the class, one layer deeper: E-2b wired the seam to a condition that
cannot occur.
**R-114** — on target-drive loss the message says the backup is *"on the same disk as the system"*
(false) and offers **the drive that just vanished**. Invisible today only because of R-112 — so
**R-114 must be fixed before R-112 is wired.**
Also filed as a **second instance under R-110** (not a new ID): host-install fetches **nine** files
from `raw/branch/main` and the hub vouches a sha for **one**; E-2a's wrapper is installed 0755 to
`/usr/local/sbin`, root-fenced in sudoers, validated only by `bash -n`.
## Record
- `OPEN-ITEMS.md`**R-112/R-113/R-114 opened**; E-2d re-stated with results and left open for the
residue; E-2's "NOT yet live-proven" list resolved into proven / known-broken; R-94 fully unblocked;
R-110 extended. The drill-cleanup row was opened and then **closed the same session** once the
teardown completed, so it is not carried in the register.
- `ROADMAP.md` — R-112/R-113/R-114 under P1; R-111 marked SHIPPED.
- **`architecture/00-capability-map.md` not touched** — for two reasons: the customer-facing legs are
broken rather than proven, and the map has **no E-2 / backup-target rows at all** (worth noting
against the ROADMAP's coupling rule).
## Teardown
VM destroyed, scratch storage removed, **`pvesm status` after == before** (`local-lvm` 38.77 %,
byte-identical), guest 9201 and drill-r50 untouched. **Hub records removed — teardown complete.** The delete was correctly refused at four gates while the host still read ONLINE; once the destroyed host aged to DOWN (`delete-impact``deletable:true`) the documented cascade ran and completed: host deleted, PBS tenancy deprovisioned, claim reset, residue purged. Verified after: **0** `e2d` occurrences on the hosts page, fleet unchanged. The one purged `appliance_registrations=1` was this run's own appliance; the unrelated stale 2026-07-25 appliance (`206c8838…`) was not touched by the cascade — the operator removed it separately.
## One human step, and a premise correction
The runbook's §5.1a operator STOP (the bind) is **retired** — CC did it. But E-2d's premise that a
fresh install yields a CC-drivable claimable customer is **wrong**: the claim code is bcrypt-hashed and
email-only, and the gate covers everything except `/claim`, `/api/health`, `/static/`. One operator
relay of the emailed code was required — which also proved the claim flow end to end.