d839ddcb60
The delete was correctly refused at four successive gates while the host still read ONLINE (acknowledgements -> typed confirm_id -> expect_hosts stale-preview -> "host is ONLINE"). Rather than force it, the run waited for the destroyed host to age to DOWN; delete-impact then reported deletable:true and the documented cascade ran: host deleted (escrow demoted to retained custody), tenantsync deprovisioned, PBS tenancy deprovisioned, claim reset to unclaimed, residue purged (reports=5 app_telemetry=5 notif_prefs=1 appliance_registrations=1) Verified after: 0 occurrences of "e2d" anywhere on the hosts page; demo-felhom and demo-hp ONLINE on agent 0.113.0; drill-r50 and peti-felhom unchanged; demo-hp carries only guest 9201 and VM 300. Scoping checked rather than assumed: the single purged appliance_registration was this run's own appliance (810d10c5, bound to e2d-fresh). The unrelated stale 2026-07-25 appliance (206c8838 / QWA-WJE) was NOT touched by the cascade — the operator removed it separately. - OPEN-ITEMS.md: the drill-cleanup WATCHING row is removed (done, not open). - audits/E2D-fresh-vm-2026-07-29.md §8 + REPORT-e2d.md: teardown recorded as complete, with the cascade output and the appliance-scoping note.
86 lines
5.8 KiB
Markdown
86 lines
5.8 KiB
Markdown
# REPORT — R-111 fixed, then E-2 proven on a fresh box (2026-07-29)
|
|
|
|
Two phases in one session. Full evidence: `documentation/audits/E2D-fresh-vm-2026-07-29.md`.
|
|
Root `REPORT.md` untouched.
|
|
|
|
## Phase 1 — R-111: the Day-0 channel now serves the current software
|
|
|
|
A Phase 0 gate earlier the same day stopped the E-2d run before any VM existed: a fresh box would
|
|
have installed **agent 0.96.0 + controller 0.161.0**, ~17 and ~24 releases behind `main`.
|
|
|
|
| | Before | Now |
|
|
|---|---|---|
|
|
| agent (Gitea generic) | 0.96.0 | **0.113.0**, sha `5f3247f7…`, round-trip verified |
|
|
| golden (Gitea generic) | 0.161.0 | **0.185.1**, sha `dba00f3e…`, embeds controller 0.185.1 |
|
|
| hub `min_agent` | 0.93.0 | **0.113.0** (what controller v0.185.0 declares) |
|
|
|
|
Bake clean on every marker: `Result=success`, overlay2, **all three mounts in the archive**, 0
|
|
FATAL/exclusions, HTTP 201, token-leak grep 0. GL-1 teardown: guest 9100 purged, secrets shredded,
|
|
drill disk restored to `virgin`. Agent + golden moved in **one** manifest POST so it never vouched a
|
|
new agent against an old golden. `min_agent` verified zero-impact first (all three enrolled hosts
|
|
already at 0.113.0). Global floor deliberately **not** raised — the golden now bakes 0.185.1.
|
|
|
|
Commit `3dff357`.
|
|
|
|
## Phase 2 — the E-2d run, full ISO/PAIRING route
|
|
|
|
Nested PVE VM on demo-hp, one disk, outside the `felhom` pool. Bind → running controller in
|
|
**3 m 35 s**. The install fetched exactly the artifacts published an hour earlier and restored
|
|
`vzdump-lxc-9100-2026_07_29-12_37_56` — the golden baked 20 minutes before. The publish train is
|
|
proven end to end on a real install.
|
|
|
|
| Claim | Verdict |
|
|
|---|---|
|
|
| **C1** host-install 1.22.0 completes a real install, rc=0 | ✅ **PROVEN** |
|
|
| **C2** Case B fires naturally | ✅ **PROVEN** — both DEGRADED lines verbatim, `local_backup_target=local`, install did not abort |
|
|
| **C3** degraded banner renders **to a customer** | ⚠️ **PARTIAL** — API byte-exact; **no UI consumer exists** → **R-112** |
|
|
| **C4** offer appears and moves the target | ⚠️ **PARTIAL** — decline path, `restart_required:true`, no self-restart, E-2a wrapper, healthy-renders-nothing all PROVEN at API level; offer equally invisible → **R-112** |
|
|
| **C5** `backup_target_absent` end to end | ❌ **FAILED** — zero events on any channel → **R-113** |
|
|
|
|
## The three findings
|
|
|
|
**R-112 (P1)** — E-2's banner and offer have **no UI consumer**. The endpoint returns byte-exact copy;
|
|
`grep 'backup-target'` across every `*.html`/`*.js`/`*.css` → **0 hits**, and no page handler injects
|
|
the state. Decisive contrast: templates fetch **18** distinct `/api/storage/*` endpoints;
|
|
`backup-target` and `backup-target/assign` are the only two with zero references. v0.185.1 fixed the
|
|
router mount and stopped one layer short of the render. Fifth instance of seam-built-but-never-wired.
|
|
|
|
**R-113 (P1)** — the drive-absent gate **cannot fire on device loss**. `planDriveGates` reads presence
|
|
from `BoundUnderParent` = "is this path in the guest's mountinfo". The raw mount is a device-bound
|
|
systemd unit and dies; **the agent's own bind is not device-bound and outlives the device**, so the
|
|
gate sees "present" forever. Live: agent said `enrolled drive absent by UUID` every 20 s for 4½
|
|
minutes, controller logged **0** `[gate]` lines, hub got **zero** events — neither the specific nor the
|
|
generic one. Sixth instance of the class, one layer deeper: E-2b wired the seam to a condition that
|
|
cannot occur.
|
|
|
|
**R-114** — on target-drive loss the message says the backup is *"on the same disk as the system"*
|
|
(false) and offers **the drive that just vanished**. Invisible today only because of R-112 — so
|
|
**R-114 must be fixed before R-112 is wired.**
|
|
|
|
Also filed as a **second instance under R-110** (not a new ID): host-install fetches **nine** files
|
|
from `raw/branch/main` and the hub vouches a sha for **one**; E-2a's wrapper is installed 0755 to
|
|
`/usr/local/sbin`, root-fenced in sudoers, validated only by `bash -n`.
|
|
|
|
## Record
|
|
|
|
- `OPEN-ITEMS.md` — **R-112/R-113/R-114 opened**; E-2d re-stated with results and left open for the
|
|
residue; E-2's "NOT yet live-proven" list resolved into proven / known-broken; R-94 fully unblocked;
|
|
R-110 extended. The drill-cleanup row was opened and then **closed the same session** once the
|
|
teardown completed, so it is not carried in the register.
|
|
- `ROADMAP.md` — R-112/R-113/R-114 under P1; R-111 marked SHIPPED.
|
|
- **`architecture/00-capability-map.md` not touched** — for two reasons: the customer-facing legs are
|
|
broken rather than proven, and the map has **no E-2 / backup-target rows at all** (worth noting
|
|
against the ROADMAP's coupling rule).
|
|
|
|
## Teardown
|
|
|
|
VM destroyed, scratch storage removed, **`pvesm status` after == before** (`local-lvm` 38.77 %,
|
|
byte-identical), guest 9201 and drill-r50 untouched. **Hub records removed — teardown complete.** The delete was correctly refused at four gates while the host still read ONLINE; once the destroyed host aged to DOWN (`delete-impact` → `deletable:true`) the documented cascade ran and completed: host deleted, PBS tenancy deprovisioned, claim reset, residue purged. Verified after: **0** `e2d` occurrences on the hosts page, fleet unchanged. The one purged `appliance_registrations=1` was this run's own appliance; the unrelated stale 2026-07-25 appliance (`206c8838…`) was not touched by the cascade — the operator removed it separately.
|
|
|
|
## One human step, and a premise correction
|
|
|
|
The runbook's §5.1a operator STOP (the bind) is **retired** — CC did it. But E-2d's premise that a
|
|
fresh install yields a CC-drivable claimable customer is **wrong**: the claim code is bcrypt-hashed and
|
|
email-only, and the gate covers everything except `/claim`, `/api/health`, `/static/`. One operator
|
|
relay of the emailed code was required — which also proved the claim flow end to end.
|