Files
felhom.eu/documentation/audits/DRILL-day0-take2-2026-07-12.md
T

8.3 KiB
Raw Blame History

DRILL — Day-0 "take two" on the Demo-VM (DR-tier-by-default validation), 2026-07-12 evening

The §13 live validation of the DR-tier-by-default batch (installer v1.15.0 + agent v0.86.0

  • hub v0.51.0). Same box as DRILL-day0-vm-2026-07-12 (qm 300 on felhom-pve, rolled back to pre-day0-clean), same actors (Viktor: GO decisions, claim, ceremony R-moment, backup click; CC: everything else). Outcome: the ENTIRE first-drill §5 sequence completed with ZERO fix-and-continue stops on the Day-0/DR chain — every drill F-fix proven live — in 1h 22m (vs 1h 55m), while covering MORE (claim gate, floor-driven self-update first proof, reset flow).

1. Reset (not counted in the headline metric)

  • 21:13 qm 300 → pre-day0-clean + start (fresh box 192.168.0.152).
  • ep0 tenancy cleared manually (token felhom@pbs!demo-vm-felhom deleted + empty ns dir removed) — REQUIRED, see finding F-14 below.
  • 21:44 hub stale-host delete of demo-vm-felhom-2482b0 via the v0.47.0 danger-zone flow (impact probe + escrow-ack + typed confirm; one-tx cascade incl. WG peer, staged PBS secret, break-glass credential, escrow + DR bundle) — first live exercise of the escrow-ack delete.
  • Customer demo-vm-felhom kept; its dr_tier flag was ON via the v0.51.0 migration backfill (initialize-from-reality proven live: the pre-delete host carried an enabled descriptor).

2. The run (wall-clock 21:48 → 23:10, CEST)

Time Stage Evidence
21:48 installer fetched; -h = v1.15.0 (F-1: single source) passphrase staged by Viktor (0600 file, shredded after)
21:49 dry-run reviewed F-2 honest anonymous-fetch lines; felhom-pbs "expected absent — grant pre-positioned" INFO; manifest agent 0.86.0 d77355c5…
21:5021:52 real install, exit 0, zero stops age installed (F-10), felhom-pbs-apply shipped (F-7), wg_tunnel.enabled: true rendered (F-9), full ACL incl. /storage/felhom-pbs; F-8 rotation pointer in 4b + final summary; host demo-vm-felhom-2f4b00; guest 9201 controller 0.120.0 healthy
21:50:51 first capability self-check: ok=59 total=62 degraded=0 inactive=3 the drill's "3 pbsdr-* born DEGRADED" is GONE — inactive is the honest disabled-pending state (agent v0.86.0)
21:50:52 WG registered hands-free (10.77.0.3/32)
21:50:53 hub auto-provisioned the PBS-DR descriptor — 1 s after WG registration, zero operator steps (scenario A) hub log: "pbsdr auto-provisioned for demo-vm-felhom on WG registration (hands-free cascade)"
21:52 F-3 proven: guests{,/9201} felhom-agent-owned, bootstrap subtree guest-root root-run provision parent chown (agent v0.86.0)
~21:54 floor-driven self-update 0.120.0 → 0.122.0, zero steps FIRST LIVE PROOF of the update path (drill-1 §4.3 limitation closed); real edge / → 302 claim page
22:05:56 agent bridge applied the descriptor on its first desired-fetch after gen 2 consume-once → entry felhom-pbs active → K files → escrow.pbs_storage_id seeded → state=applied
~22:13 Viktor claimed the box (after the F-15 reset-code lag — see findings) dashboard password customer-set; gate closed
22:1122:22 offsite re-attach: hub "Re-issue offsite credentials" (the documented F4 recovery for the fresh guest's consumed-password dead-end) → ConfigVersion bump → controller applied, key installed, restic password staged to the agent drill-1 repo remnant wiped via the controller's key first (cryptographically dead scratch — its password died with the deleted escrow)
22:46:11 ceremony (Viktor, R on paper, nothing in CC's transcript) 383-byte zero-knowledge blob in hub custody; staged secret wiped
22:52:50 auto-confirm, hands-off: pending → escrowed in 6m 39s, zero clicks beats the drill's 7.5 min; "offsite runs enabled"
23:0x ActualBudget installed via the dashboard (D.4 smoke, "Telepítés sikeres") + toggled for NAS CC drove Viktor's session (he authorized)
23:08 offsite backup: 1 app, 1 snapshot, 10 s 4.6 KB / 1 pillanatkép on the box
23:10:39 verification restore: "restored actualbudget → …/offbox-restore/actualbudget", 48K THE ROUND-TRIP IS PROVEN (same shape as drill-1)
23:12 post-apply capability check: ok=62 total=62 degraded=0 inactive=0 inactive→ok flip after tier apply; agent-restart probe also proves DRConfigured's marker fallback live

Cascade UI (scenario D) verified at both states: pre-run all-4-done on the old host; waiting→done progression on the new host (screenshots in the session record). Hub host page renders the NEW capability chip table (v0.51.0) live.

3. Headline metrics

Metric take two drill 1
Wall-clock, install-start → offsite round-trip proven 1h 22m 1h 55m
Fix-and-continue stops on the Day-0/DR chain 0 3 (+ mid-drill fork)
Operator steps for the DR tier 0 (flag was already ON; cascade hands-free) manual wrapper+age+wg+ACL retrofit + hub enable
Ceremony → escrowed (auto-confirm) 6m 39s ~7.5 min
pbsdr capabilities at first boot inactive (neutral) DEGRADED (red)
Floor-driven self-update PROVEN (0.120→0.122, 0 steps) not stressable (baked == floor)

4. New findings

# Sev Finding Direction
F-14 MEDIUM Host delete + re-enroll while the ep0 tenancy survives = DR re-attach dead-end: auto-provision AND config save hard-error token_exists; "Re-issue PBS credentials" 400s (requires the descriptor the deleted host took with it). Recovery today = manual ep0 root token-delete (done during this reset) SHIPPED 2026-07-13 (hub v0.53.0, operator ruling): auto-Reissue permitted ONLY when the hub's own deletion record (host_deletions, written in the DeleteHost tx) shows the owning host was removed through the escrow-ack flow — acknowledged destruction, not silent re-keying; no record (incl. THIS drill's pre-record reset) / un-acked → the refusal + manual path, byte-unchanged. Scenario A/B tests + red-proofs; live validation = the next real host-reset cycle (fixtures carry it until then)
F-15 MEDIUM/UX — operator-flagged must-fix Claim/reset code NOT immediately usable: the hub rotates the hash at reset-request, but the box learns it only on its next report ACK (~15 min). Viktor hit it live ("Hibás vagy lejárt kód" with a fresh code) SHIPPED 2026-07-13 (hub v0.52.0 + controller v0.123.0): the reset-request response carries the rotated hash, applied via the ACK's generation-guarded consumer. Live re-run of the exact failure path: applied 1 s after the request, code accepted first try
F-16 LOW Native confirm() on hub buttons (offsite/PBS re-issue, freeze) froze CC's browser automation — the hub-side siblings of drill F-11 SHIPPED 2026-07-13 (hub v0.52.0 + controller v0.123.0): inline "Igen/Mégse" two-step everywhere, both repos; hub_confirm_gate.py/native_confirm_gate.py enforce zero native confirms. Live: the offsite re-issue completed under automation without freezing
obs. A re-provisioned guest over an EXISTING offsite repo needs the old repo password (gone with the deleted escrow) or a repo wipe — this IS the S5 DR-restore scenario, already queued as its own drill; the wipe was the correct reset action here, not a product gap S5 drill
obs. The controller's very first offbox run with zero toggled apps reports "backup OK, 0 snapshots" — arguably should hint "no apps toggled" in the UI SHIPPED 2026-07-13 (controller v0.123.0): toggle-list hint + "Sikeres — nincs mentésre jelölt alkalmazás" run copy, live-proven

5. Disposition

F-1/F-2/F-3/F-7/F-8/F-9/F-10: CLOSED, live-proven on a fresh box (this run). F-6: closed by policy (decision 4; the coupling guard + always-capable plumbing held live). F-4/F-5: closed by the claim arc (re-proven here incl. a reset-flow first exercise → F-15). Drill-1 §4.3 floor-update limitation: closed (first live firing). Blast radius honored: guest demo 9201, Peti's hub entry, the demo offbox, and the OTHER ep0 namespaces were never touched.

Box state kept at end: qm 300 running the full take-two result (fresh snapshots can be taken for re-drills); drill-1's snapshots (post-install, post-drill, post-claim-arc) remain.