8.3 KiB
DRILL — Day-0 "take two" on the Demo-VM (DR-tier-by-default validation), 2026-07-12 evening
The §13 live validation of the DR-tier-by-default batch (installer v1.15.0 + agent v0.86.0
- hub v0.51.0). Same box as DRILL-day0-vm-2026-07-12 (qm 300 on felhom-pve, rolled back to
pre-day0-clean), same actors (Viktor: GO decisions, claim, ceremony R-moment, backup click; CC: everything else). Outcome: the ENTIRE first-drill §5 sequence completed with ZERO fix-and-continue stops on the Day-0/DR chain — every drill F-fix proven live — in 1h 22m (vs 1h 55m), while covering MORE (claim gate, floor-driven self-update first proof, reset flow).
1. Reset (not counted in the headline metric)
- 21:13 qm 300 →
pre-day0-clean+ start (fresh box 192.168.0.152). - ep0 tenancy cleared manually (token
felhom@pbs!demo-vm-felhomdeleted + empty ns dir removed) — REQUIRED, see finding F-14 below. - 21:44 hub stale-host delete of
demo-vm-felhom-2482b0via the v0.47.0 danger-zone flow (impact probe + escrow-ack + typed confirm; one-tx cascade incl. WG peer, staged PBS secret, break-glass credential, escrow + DR bundle) — first live exercise of the escrow-ack delete. - Customer
demo-vm-felhomkept; its dr_tier flag was ON via the v0.51.0 migration backfill (initialize-from-reality proven live: the pre-delete host carried an enabled descriptor).
2. The run (wall-clock 21:48 → 23:10, CEST)
| Time | Stage | Evidence |
|---|---|---|
| 21:48 | installer fetched; -h = v1.15.0 (F-1: single source) |
passphrase staged by Viktor (0600 file, shredded after) |
| 21:49 | dry-run reviewed | F-2 honest anonymous-fetch lines; felhom-pbs "expected absent — grant pre-positioned" INFO; manifest agent 0.86.0 d77355c5… |
| 21:50–21:52 | real install, exit 0, zero stops | age installed (F-10), felhom-pbs-apply shipped (F-7), wg_tunnel.enabled: true rendered (F-9), full ACL incl. /storage/felhom-pbs; F-8 rotation pointer in 4b + final summary; host demo-vm-felhom-2f4b00; guest 9201 controller 0.120.0 healthy |
| 21:50:51 | first capability self-check: ok=59 total=62 degraded=0 inactive=3 |
the drill's "3 pbsdr-* born DEGRADED" is GONE — inactive is the honest disabled-pending state (agent v0.86.0) |
| 21:50:52 | WG registered hands-free (10.77.0.3/32) | |
| 21:50:53 | hub auto-provisioned the PBS-DR descriptor — 1 s after WG registration, zero operator steps (scenario A) | hub log: "pbsdr auto-provisioned for demo-vm-felhom on WG registration (hands-free cascade)" |
| 21:52 | F-3 proven: guests{,/9201} felhom-agent-owned, bootstrap subtree guest-root |
root-run provision parent chown (agent v0.86.0) |
| ~21:54 | floor-driven self-update 0.120.0 → 0.122.0, zero steps | FIRST LIVE PROOF of the update path (drill-1 §4.3 limitation closed); real edge / → 302 claim page |
| 22:05:56 | agent bridge applied the descriptor on its first desired-fetch after gen 2 | consume-once → entry felhom-pbs active → K files → escrow.pbs_storage_id seeded → state=applied |
| ~22:13 | Viktor claimed the box (after the F-15 reset-code lag — see findings) | dashboard password customer-set; gate closed |
| 22:11–22:22 | offsite re-attach: hub "Re-issue offsite credentials" (the documented F4 recovery for the fresh guest's consumed-password dead-end) → ConfigVersion bump → controller applied, key installed, restic password staged to the agent | drill-1 repo remnant wiped via the controller's key first (cryptographically dead scratch — its password died with the deleted escrow) |
| 22:46:11 | ceremony (Viktor, R on paper, nothing in CC's transcript) | 383-byte zero-knowledge blob in hub custody; staged secret wiped |
| 22:52:50 | auto-confirm, hands-off: pending → escrowed in 6m 39s, zero clicks | beats the drill's 7.5 min; "offsite runs enabled" |
| 23:0x | ActualBudget installed via the dashboard (D.4 smoke, "Telepítés sikeres") + toggled for NAS | CC drove Viktor's session (he authorized) |
| 23:08 | offsite backup: 1 app, 1 snapshot, 10 s | 4.6 KB / 1 pillanatkép on the box |
| 23:10:39 | verification restore: "restored actualbudget → …/offbox-restore/actualbudget", 48K | THE ROUND-TRIP IS PROVEN (same shape as drill-1) |
| 23:12 | post-apply capability check: ok=62 total=62 degraded=0 inactive=0 |
inactive→ok flip after tier apply; agent-restart probe also proves DRConfigured's marker fallback live |
Cascade UI (scenario D) verified at both states: pre-run all-4-done on the old host; waiting→done progression on the new host (screenshots in the session record). Hub host page renders the NEW capability chip table (v0.51.0) live.
3. Headline metrics
| Metric | take two | drill 1 |
|---|---|---|
| Wall-clock, install-start → offsite round-trip proven | 1h 22m | 1h 55m |
| Fix-and-continue stops on the Day-0/DR chain | 0 | 3 (+ mid-drill fork) |
| Operator steps for the DR tier | 0 (flag was already ON; cascade hands-free) | manual wrapper+age+wg+ACL retrofit + hub enable |
| Ceremony → escrowed (auto-confirm) | 6m 39s | ~7.5 min |
| pbsdr capabilities at first boot | inactive (neutral) | DEGRADED (red) |
| Floor-driven self-update | PROVEN (0.120→0.122, 0 steps) | not stressable (baked == floor) |
4. New findings
| # | Sev | Finding | Direction |
|---|---|---|---|
| F-14 | MEDIUM | Host delete + re-enroll while the ep0 tenancy survives = DR re-attach dead-end: auto-provision AND config save hard-error token_exists; "Re-issue PBS credentials" 400s (requires the descriptor the deleted host took with it). Recovery today = manual ep0 root token-delete (done during this reset) |
SHIPPED 2026-07-13 (hub v0.53.0, operator ruling): auto-Reissue permitted ONLY when the hub's own deletion record (host_deletions, written in the DeleteHost tx) shows the owning host was removed through the escrow-ack flow — acknowledged destruction, not silent re-keying; no record (incl. THIS drill's pre-record reset) / un-acked → the refusal + manual path, byte-unchanged. Scenario A/B tests + red-proofs; live validation = the next real host-reset cycle (fixtures carry it until then) |
| F-15 | MEDIUM/UX — operator-flagged must-fix | Claim/reset code NOT immediately usable: the hub rotates the hash at reset-request, but the box learns it only on its next report ACK (~15 min). Viktor hit it live ("Hibás vagy lejárt kód" with a fresh code) | SHIPPED 2026-07-13 (hub v0.52.0 + controller v0.123.0): the reset-request response carries the rotated hash, applied via the ACK's generation-guarded consumer. Live re-run of the exact failure path: applied 1 s after the request, code accepted first try |
| F-16 | LOW | Native confirm() on hub buttons (offsite/PBS re-issue, freeze) froze CC's browser automation — the hub-side siblings of drill F-11 |
SHIPPED 2026-07-13 (hub v0.52.0 + controller v0.123.0): inline "Igen/Mégse" two-step everywhere, both repos; hub_confirm_gate.py/native_confirm_gate.py enforce zero native confirms. Live: the offsite re-issue completed under automation without freezing |
| obs. | — | A re-provisioned guest over an EXISTING offsite repo needs the old repo password (gone with the deleted escrow) or a repo wipe — this IS the S5 DR-restore scenario, already queued as its own drill; the wipe was the correct reset action here, not a product gap | S5 drill |
| obs. | — | The controller's very first offbox run with zero toggled apps reports "backup OK, 0 snapshots" — arguably should hint "no apps toggled" in the UI | SHIPPED 2026-07-13 (controller v0.123.0): toggle-list hint + "Sikeres — nincs mentésre jelölt alkalmazás" run copy, live-proven |
5. Disposition
F-1/F-2/F-3/F-7/F-8/F-9/F-10: CLOSED, live-proven on a fresh box (this run). F-6: closed by policy (decision 4; the coupling guard + always-capable plumbing held live). F-4/F-5: closed by the claim arc (re-proven here incl. a reset-flow first exercise → F-15). Drill-1 §4.3 floor-update limitation: closed (first live firing). Blast radius honored: guest demo 9201, Peti's hub entry, the demo offbox, and the OTHER ep0 namespaces were never touched.
Box state kept at end: qm 300 running the full take-two result (fresh snapshots can be taken for
re-drills); drill-1's snapshots (post-install, post-drill, post-claim-arc) remain.