Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
19 KiB
DRILL GL-6 — supervised go-live drill (demo box, two full cycles) — 2026-07-08
Go-live item G7 + the live-validation landing zone for GL-1…GL-5b deferrals. Destructive- supervised: guest 9201 destroyed + rebuilt twice; Viktor present, gating every phase. This record is filled per phase as evidence lands; timings + deviations verbatim. No secrets.
Inputs verified at start: felhom.eu bd97bc34 (host-install v1.11.1, operator pubkeys
pinned) · felhom-agent 7f073936 (v0.76.0) · hub Day-0 manifest vouches agent
0.76.0/9828c5f7…f50b + golden 0.103.0/8481e8a1…6026e (fetched via the box-facing API).
Phase 0 — preflight + key validation (nothing destroyed)
| Step | Result |
|---|---|
| 0a baselines | PASS — repo heads + SCRIPT_VERSION="1.11.1" + manifest all as stated above |
| 0b key validation leg 1 | PASS — all four fingerprints EQUAL (below) |
| 0c key validation leg 2 | [Viktor] — pending at Gate 0 |
| 0d abort-path backup + restore-test | PASS (local + PBS backups fresh; PBS restore-test parity ok, 5m35s) |
| 0e busy-drive prop | STAGED — gl6-busyprop unit cwd-pinned into /mnt/felhom-drives/felhom-usb (4h) |
| 0f mechanism re-read | DONE — Phase-3 delivery pinned (below) |
0b — the four fingerprints, side by side
repo op (felhom-op-1): SHA256:7YqN4rXO08yixTeOO+UtQ8jHyIGycICuctQgRYVGnWw
box op (felhom-op-1): SHA256:7YqN4rXO08yixTeOO+UtQ8jHyIGycICuctQgRYVGnWw
repo rec (felhom-rec-1): SHA256:/ixgTesZqykAGJpFUUd4kLAiHFgKOkYFLNC3AQXWP+k
box rec (felhom-rec-1): SHA256:/ixgTesZqykAGJpFUUd4kLAiHFgKOkYFLNC3AQXWP+k
Box source: felhom-pve agent config (systemctl cat → -config path) authz.signers[];
repo source: the v1.11.1 OPERATOR_KEY_*_LINE constants. Both match the runbook's stated values.
0d — abort-path backups (the session's safety net)
- Local (agent pipeline):
--selftest=backup -vmid 9201→ snapshot-mode, crash-consistent,local:backup/vzdump-lxc-9201-2026_07_08-11_43_42.tar.zst(6.82 GB, 5m52s). Expecteduncovered_volumes:/etc/felhom-bootstrap+/mnt/felhom-drives(bind mounts — regenerated / re-enrolled, per doctrine). - PBS (offsite tier, PVE-native vzdump per the S4-proven path):
felhom-pbs:backup/ct/9201/2026-07-08T09:51:00Z— 3m04s (dedup) — NOTE: the live config deliberately pinslocal_backup_target: local; the PBS leg was driven withvzdump 9201 --storage felhom-pbs --mode snapshot(tunnel healthy, storage active). - Restore-test vs the fresh PBS archive: PASS —
mount_parity: ok, inventorymp0=/var/lib/docker (200G),mp1=/mnt/sys_drive (50G), mp8/mp9 throwaways; scratch 990000 restored+booted+verified+torn-down in 5m35s (first G12 PBS-tier runtime datum). Observation (cosmetic): the record'ssource_tiersays "local" — the tier label derives from the CONFIG backup target, not the explicit-archive's storage; display-only.
0e — busy prop + data-intact witness
Prop: transient unit gl6-busyprop, cwd inside /mnt/felhom-drives/felhom-usb.
Witness file (for the Phase-1/5 drive-data-intact checks):
userdata/actualbudget/canary_hdd_actualbudget.txt on /dev/sdd1
(UUID da9e7089-cf8e-4617-adcb-a377743fae00), sha256 prefix 7b00bc3550b10e3952e0e0965f5a2628.
0f — Phase-3 signed-update mechanism (confirmed at source, NOT improvised)
- [Viktor] sign:
felhom-opsign -op agent_update -host <host_id> -key-id felhom-op-1 -agent-version <v> -sha256 <hex> -key <operational private key> [-ttl 30m]→ signed-op blob. - Deliver:
POST /api/v1/admin/hosts/{host_id}/jobs{job_id, blob_b64}— hub GLOBAL key required (Viktor), or the v0.70.0-precedent directsigned_jobssqlite insert on the hub pod (CC via kubectl on 180). The agent's heartbeat seeshas_signed_ops, fetches, verifies against the PINNED operational key + durable nonce store, downloads the binary from Gitea, re-verifies the SIGNED sha (the only integrity root), thenfelhom-selfupdate-guarded applyA/B-swaps with crash-rollback armed; dwell →commit.
Gate 0 — [Viktor]
- Input 1 (byo caps): --cores 4 --memory 8192
- Input 2 (Phase-6 repopulation): fresh + redeploy apps
- 0c fingerprint confirmation vs password manager: CONFIRMED — match
- Verdict: GO — with one pre-Phase-1 STOP + ruling (below)
Gate-0 deviation — the demo customer's empty git credentials (Viktor ruling → v1.11.2)
Pre-destroy check found demo-felhom's git.username/git.token still EMPTY (the GL-2 finding);
the v1.11.1 installer would die at step 5/8. Viktor ruled: the artifacts are world-readable by
design — anonymous fetch is the supported path. All 12 installer fetch targets (2 packages + 10
raw configs) validated HTTP 200 anonymously, then host-install v1.11.2 shipped mid-drill
(felhom.eu 0640aa06): empty creds → loud WARN + anonymous fetch (conditional curl auth; the
hub-vouched sha256 stays the integrity root; configured creds still used when present). Harness
26/26 (new GL6-ANON case; GL4-C2 assertion updated for the pinned constants).
Phase 1 — uninstall #1 (full mode, busy drive staged) — PASS with 2 findings
Real --uninstall --vmid 9201 (typed confirm, v1.11.2 fetched from the served URL):
- Busy-drive behavior (the staged prop on felhom-usb): WARN + "Eject the drive via the
dashboard, or stop the apps using it and re-run" — NO
-l/-f, teardown CONTINUED; the other two drives unmounted ("data stays on the drive"); the root bind correctly left mounted with its own guard warn. The KEPT-vs-WIPED statement listed felhom-usb as "still mounted (busy — stop the apps and retry)" — i.e. the code takes the statement-lists-it-as-retry path; after releasing the prop a plain manualumountcompleted the job (no re-run needed). - KEPT-vs-WIPED statement: printed verbatim as designed (guest+volumes / agent / self-update artifacts / watchdog+OOB / pveum+pool / state = WIPED; drives+data / PBS backups / hub record / escrow = KEPT). The recovery-credential KEPT line correctly keyed off the state file — which on THIS box predates the 4b state marker, so the line was absent although a G1-vaulted credential exists (observation: state-keyed conditionals can under-report on pre-v1.10.0 installs).
- 1b audit: agent binary+baks, selfupdate wrapper+slots, rollback unit+drop-in, watchdog set, mkfs wrapper, snippets, dnsmasq drop-ins, sudoers, state dirs, pveum user/roles/ACL/pool — ALL GONE. PVE printed a benign "ignore invalid acl role 'FelhomAgentStore'" notice during ACL removal (ordering noise, recorded).
- Drive data intact: witness file re-hashed on the still-mounted primary path —
7b00bc3550b10e3952e0e0965f5a2628== the Phase-0 value.
Findings
- GL6-F1 (MED, security-adjacent):
/etc/felhom-agent/SURVIVES the uninstall holdingagent.json.bak-*files (incl. the operator's same-day ceremony backup) — old configs with secret material (pve token now dead; the per-host hub api_key is still live because the hub record survives). The uninstall removesagent.jsonbut not.bak-*siblings. Fix candidate: purge"$agent_cfg"*/ the dir in 4b (v1.11.3). Files left in place pending Viktor's call (one is his own backup). - GL6-F2 (MED): the agent-created per-drive systemd mount units
(
mnt-felhom-usb.mount+ flash + teszt_enroll) and their ACTIVE primary mounts at/mnt/<name>survive — the statement's "physically removable now" overclaims (the binds are gone but the partition is still mounted). Fix candidate: uninstall stops+removes the agent'smnt-*units for enrolled namespaces (data untouched). Left as-is for Phase 2 (the re-enroll behavior against pre-existing mounts is itself drill-relevant).
Gate 1 — [Viktor]: GO (delete the .bak secrets now)
GL6-F1 residue resolved at the gate: shred -u /etc/felhom-agent/agent.json.bak-* + rmdir
(incl. Viktor's ceremony backup, confirmed disposable). F1/F2 kept as v1.11.3 fix candidates.
Phase 2 — byo-mode day-0 (Peti's exact path, on real hardware) — PASS (happy path), 3 findings
Command (Viktor typed the demo-felhom hostname ack — interactive consent leg):
./felhom-host-install.sh --customer-id demo-felhom --mode byo --vmid 9201 --cores 4 --memory 8192 --force-gitea-golden --acl-storages "local local-lvm" --passphrase-file <0600> (+ --allow-new-leaf --resume after F3/F4, see below).
End state — all green (verify step, verbatim highlights):
- disclosure printed; Viktor's typed-hostname ack accepted (
byo install acknowledged for host demo-felhom); - NO 4b / chpasswd / recovery-credential anywhere (grep of the full transcript = 0);
- anonymous artifact fetch worked (v1.11.2): agent v0.76.0 sha
9828c5f7…verified, golden 0.103.0 sha8481e8a1…verified, both against the hub manifest; - installed agent
--version== 0.76.0; daemon active; authz signers: 2 (operator-signed self-update armed)from first boot — key validation leg 3 ✓ (config carries operational:felhom-op-1 + recovery:felhom-rec-1);- pool member + all 4 scoped-ACL grants (user+token) present; caps on the guest: cores 4 / memory 8192;
- byo config asserts pass (lan_resolver/wg_tunnel/oob all off);
- controller Up (healthy) 0.103.0 + traefik + cloudflared up; dashboard
HTTP 200(Host-header probe); hub config-retrieve OK.
Findings — the byo install's resume path is fragile (happy single-shot path is clean)
- GL6-F3 (MED, GL-7-relevant): on a BYO host carrying the OWNER's (non-Felhom) guests, a FRESH
install trips the populated-host leaf guard (
pct list | grep -q .counts ANY guest) and dies demanding--preserve-state-from/--allow-new-leaf— even though no Felhom controller pin exists to protect on a fresh install. Peti's box has his own guests → his day-0 command MUST include--allow-new-leaf(GL-7). Better fix candidate: narrow the guard tofelhom_guestsso a fresh byo install on a populated non-Felhom host doesn't trip. - GL6-F4 (HIGH, resume-correctness):
--resumefully skips completed steps viashould_skip X || step_X, but later steps consume the skipped steps' IN-MEMORY outputs, which are then empty → broken results: (a)enrollskipped →HOST_ID/HOST_API_KEYempty →step_agent_configwrites a config missinghub.host_id(+ emptyproxmox.tokenwhentokenis skipped) → daemon crash-loops (hub.host_id is required); the mint-once token secret is unrecoverable, so recovery = re-runtoken+enroll. (b)growsskipped →ROOTFS_GROW/… empty → provision passes-rootfs-grow ""→ flag parse error. Notestep_tokeneven has an internal resume-guard ([[ -n "$PVE_TOKEN" ]]) that the|| step_tokendispatch DEFEATS — strong signal the dispatch is the bug (steps whose outputs feed later steps must re-run, or re-derive from the written config, on resume). Drill recovery: dropped the affected completed marks fromstate.jsonto force re-run — the install then completed cleanly. The single-shot happy path (no interruption) is unaffected, which is Peti's path. - GL6-F2 (from Phase 1, MED): the agent's per-drive
mnt-*.mountunits + primary/mnt/<name>mounts survived uninstall #1; on this reinstall the parent bind/mnt/felhom-drivesis back but the drives are NOT yet bound under it — re-enrollment is the Phase-2b UI leg.
F4 fixed mid-drill (Viktor: HOLD → fix now) — host-install v1.11.3 (98bf5a43)
Root cause confirmed: should_skip X || step_X fully skips a completed step on --resume, but
token/enroll/grows are PRODUCERS whose in-memory outputs later steps consume; step_token even had
an internal resume-guard the || step_token dispatch DEFEATED. Fix: token/enroll/grows run every
pass (all idempotent — reuse-or-rotate / mint-once-reuse POST / pure recompute); golden's
GOLDEN_VOLID re-derived from the local archive in the resume block. Harness +GL6-F4 invariant,
27/27, shellcheck clean.
Live proof (v1.11.3, crafted --resume --dry-run on felhom-pve, throwaway FELHOM_INSTALL_STATE_DIR,
zero mutations): marks = the exact pre-crash set [preflight,token,grows,enroll] (agent_config
NOT done). Transcript: steps 2/8 token, 3/8 grows, 4/8 enroll all RE-RUN; the agent_config write
shows host_id=<dry-run-host-id> (non-empty) and provision shows -rootfs-grow 0 -datavol-grow 184 -sysdata-grow 42 — both the exact values that were empty (→ crash-loop / flag parse error)
pre-fix. Real 9201 + its state.json untouched (verified: daemon active, controller Up healthy).
GL6-F4 CLOSED + live-validated. (GL6-F3 stays a GL-7 command requirement --allow-new-leaf;
guard-narrowing to felhom_guests recorded as a candidate. GL6-F1/F2 remain v1.11.x candidates.)
Phase 2b — bring-to-life: BLOCKED on public DNS (GL6-F5), browser bridge confirmed working
- Browser bridge: WORKS — created a tab group, navigated to the dashboard; the failure below is the SITE, not the bridge.
- GL6-F5 (onboarding, Viktor/Cloudflare): after the byo reinstall,
felhom.demo-felhom.eudoes NOT resolve publicly —curlfrom the box (outbound internet) returnsCould not resolve host. The cloudflared tunnel is UP (QUIC + HTTP/2 to argotunnel, precheck all pass, same tunnel token), and the dashboard is healthy IN-GUEST (HTTP 200, "Beállítás" setup page) — so the controller + tunnel are fine; the missing piece is the public DNS CNAMEfelhom.demo-felhom.eu → <tunnel>.cfargotunnel.com. The controller does not auto-(re)create the dashboard's own DNS record on reinstall. This blocks the 2b real-pipeline rehearsal (password set + drive re-enroll need the dashboard reachable in a browser). NOT an install-correctness defect — the install is green; it is a Cloudflare-zone / onboarding gap that also matters for Peti (GL-7 / G9).
Phase 2b — bring-to-life (browser, CC-driven) — drive re-enroll + data-visible PASS
DNS unblock: Viktor added a Pi-hole split-horizon A record felhom.demo-felhom.eu → 192.168.0.105
(LAN path to the guest); dashboard then loaded in-browser. (The public Cloudflare path is ALSO
correct — forced-DoH resolution returned HTTP 200 with the real page; only the client-side resolver
via the box's broken dnsmasq (F6) was failing.)
- Dashboard live: "Demo Ügyfél", 3 running / 56 total apps, mem 1.1/8 GB, backup "Még nem futott".
- Drive re-enrollment (CC-driven UI): Storage → Meghajtók showed all 3 drives detected-but-
UNREGISTERED (the agent's registration state was wiped in uninstall #1; the drives + data survived
per F2, tagged "Adatot tartalmaz"). Clicked Regisztrálás on
felhom-usb(uuidda9e7089…, /dev/sdd1) — it moved to Adattárolók → Aktív, ext4 /dev/sdd1[/felhom-data] HD710 PRO, additive (no format). Data-visible PASS: the witness file is now readable in the guest at/mnt/felhom-drives/felhom-usb/userdata/actualbudget/canary_hdd_actualbudget.txt, hash7b00bc35…== the Phase-0 value. Additive doctrine proven end-to-end through the real UI. - GL6-F6 (MED, byo coexistence): the box runs a leftover broken dnsmasq (active, bound to
0.0.0.0:53, but queries time out — empty/etc/dnsmasq.d/, no upstream) on a byo host wherelan_resolveris OFF. It breaks DNS for LAN clients that use the box as a resolver (it broke the drill's browser resolution). Fix pile with F1/F2: byo install/uninstall should stop+mask the leftover dnsmasq. Immediate drill unblock was the Pi-hole entry. - G9/G10 — password is OPERATOR-set, currently UNSET → dashboard OPEN. Settings → Biztonság: "A jelszavas védelem nincs beállítva. Kérd az üzemeltetőt a beállításhoz." The dashboard password comes through the real config pipeline (hub customer-config → controller.yaml bcrypt — the anti-F9 path), NOT a customer-set field here. Until the operator sets it the public-reachable dashboard is UNAUTHENTICATED (the known G10 exposure, now confirmed live). Geo-restriction also OFF. Operator leg pending: Viktor sets the dashboard password via the hub, then CC verifies login is enforced.
Gate 2 — [Viktor]: GO — Peti's byo path live-proven end to end
Green: anonymous fetch, sha-verified 0.76.0 + 0.103.0, signers ARMED from first boot, caps applied, no root@pam touched, controller healthy, drive re-enrolled additively with pre-drill data intact. F4 found + fixed (v1.11.3) + live-proven. Deferred to GL-7 with Peti: the dashboard password (operator-set via hub — G9) + geo-restriction; G10 (open dashboard until the operator sets a password) is now a GL-7 onboarding prerequisite. Findings F1/F2/F3/F6 → post-drill installer pile.
Phase 3 — armed key end-to-end (the only leg that needed real keys)
PASS — the armed-key chain proven both directions with the REAL operational key.
- 3a: drill agent 0.76.1 published (sha
ff18972f…aed16c), version-only bump, throwaway. - Delivery mechanism (deviation from the [CC] tag, recorded): the hub pod is a minimal image
(no sqlite3/python — only the felhom-hub binary) and the admin-enqueue API is GLOBAL-key gated, so
CC cannot inject the signed job.
felhom-opsignhas a built-in-upload+-hub-key, so the sign + enqueue is ONE operator command (Viktor's secrets stay in it). The-hub-keyis the hub'sreport_api_key(live ConfigMap == committedmanifests/hub.yaml:77; the operator/global key, NOT the web-UI login) — a point of confusion worth a GL-7 note. - 3b/3c FORWARD (0.76.0→0.76.1):
felhom-opsign -op agent_update -host demo-felhom-01 -key-id felhom-op-1 -agent-version 0.76.1 -sha256 ff18972f… -key <op-priv> -upload https://hub.felhom.eu -hub-key <report_api_key>. Agent (restart-forced poll) → auditgate decision class=agent_update disposition=destructive allowed=true reason=signed key_id=felhom-op-1 nonce=4176e8e0…→ download (anon Gitea) → sha re-verify →felhom-selfupdate-guarded applyA/B flip →new version running 0.76.1 prev=0.76.0 dwell=1m0s→update committed version=0.76.1..prev=0.76.0 retained; crash-guard silent. - RETURN (0.76.1→0.76.0): same with the 0.76.0 published sha
9828c5f7…+ a FRESH nonce → committed back to 0.76.0 (the manifest-vouched version; sha confirmed). Double round-trip = sign→verify→ nonce→swap proven twice with the real key. - 3d: 0.76.1 package DELETED (HTTP 204 → 404 anon). Box ends at 0.76.0, daemon active.
Key validation leg 4 (the whole point of GL-1's key-pin): the pinned operational key actually authorizes a real signed op end-to-end. Combined with legs 1–3 (box==repo==manager fingerprints, armed-from-first-boot), the operator-signed self-update path is fully live-validated.
Gate 3 — [Viktor]:
(Phases 3–6 appended as they run.)