Files
felhom.eu/documentation/pilot/DRILL-GL6-2026-07-08.md
T

11 KiB
Raw Blame History

DRILL GL-6 — supervised go-live drill (demo box, two full cycles) — 2026-07-08

Go-live item G7 + the live-validation landing zone for GL-1…GL-5b deferrals. Destructive- supervised: guest 9201 destroyed + rebuilt twice; Viktor present, gating every phase. This record is filled per phase as evidence lands; timings + deviations verbatim. No secrets.

Inputs verified at start: felhom.eu bd97bc34 (host-install v1.11.1, operator pubkeys pinned) · felhom-agent 7f073936 (v0.76.0) · hub Day-0 manifest vouches agent 0.76.0/9828c5f7…f50b + golden 0.103.0/8481e8a1…6026e (fetched via the box-facing API).

Phase 0 — preflight + key validation (nothing destroyed)

Step Result
0a baselines PASS — repo heads + SCRIPT_VERSION="1.11.1" + manifest all as stated above
0b key validation leg 1 PASS — all four fingerprints EQUAL (below)
0c key validation leg 2 [Viktor] — pending at Gate 0
0d abort-path backup + restore-test PASS (local + PBS backups fresh; PBS restore-test parity ok, 5m35s)
0e busy-drive prop STAGEDgl6-busyprop unit cwd-pinned into /mnt/felhom-drives/felhom-usb (4h)
0f mechanism re-read DONE — Phase-3 delivery pinned (below)

0b — the four fingerprints, side by side

repo op  (felhom-op-1):  SHA256:7YqN4rXO08yixTeOO+UtQ8jHyIGycICuctQgRYVGnWw
box  op  (felhom-op-1):  SHA256:7YqN4rXO08yixTeOO+UtQ8jHyIGycICuctQgRYVGnWw
repo rec (felhom-rec-1): SHA256:/ixgTesZqykAGJpFUUd4kLAiHFgKOkYFLNC3AQXWP+k
box  rec (felhom-rec-1): SHA256:/ixgTesZqykAGJpFUUd4kLAiHFgKOkYFLNC3AQXWP+k

Box source: felhom-pve agent config (systemctl cat-config path) authz.signers[]; repo source: the v1.11.1 OPERATOR_KEY_*_LINE constants. Both match the runbook's stated values.

0d — abort-path backups (the session's safety net)

  • Local (agent pipeline): --selftest=backup -vmid 9201 → snapshot-mode, crash-consistent, local:backup/vzdump-lxc-9201-2026_07_08-11_43_42.tar.zst (6.82 GB, 5m52s). Expected uncovered_volumes: /etc/felhom-bootstrap + /mnt/felhom-drives (bind mounts — regenerated / re-enrolled, per doctrine).
  • PBS (offsite tier, PVE-native vzdump per the S4-proven path): felhom-pbs:backup/ct/9201/2026-07-08T09:51:00Z — 3m04s (dedup) — NOTE: the live config deliberately pins local_backup_target: local; the PBS leg was driven with vzdump 9201 --storage felhom-pbs --mode snapshot (tunnel healthy, storage active).
  • Restore-test vs the fresh PBS archive: PASSmount_parity: ok, inventory mp0=/var/lib/docker (200G), mp1=/mnt/sys_drive (50G), mp8/mp9 throwaways; scratch 990000 restored+booted+verified+torn-down in 5m35s (first G12 PBS-tier runtime datum). Observation (cosmetic): the record's source_tier says "local" — the tier label derives from the CONFIG backup target, not the explicit -archive's storage; display-only.

0e — busy prop + data-intact witness

Prop: transient unit gl6-busyprop, cwd inside /mnt/felhom-drives/felhom-usb. Witness file (for the Phase-1/5 drive-data-intact checks): userdata/actualbudget/canary_hdd_actualbudget.txt on /dev/sdd1 (UUID da9e7089-cf8e-4617-adcb-a377743fae00), sha256 prefix 7b00bc3550b10e3952e0e0965f5a2628.

0f — Phase-3 signed-update mechanism (confirmed at source, NOT improvised)

  1. [Viktor] sign: felhom-opsign -op agent_update -host <host_id> -key-id felhom-op-1 -agent-version <v> -sha256 <hex> -key <operational private key> [-ttl 30m] → signed-op blob.
  2. Deliver: POST /api/v1/admin/hosts/{host_id}/jobs {job_id, blob_b64}hub GLOBAL key required (Viktor), or the v0.70.0-precedent direct signed_jobs sqlite insert on the hub pod (CC via kubectl on 180). The agent's heartbeat sees has_signed_ops, fetches, verifies against the PINNED operational key + durable nonce store, downloads the binary from Gitea, re-verifies the SIGNED sha (the only integrity root), then felhom-selfupdate-guarded apply A/B-swaps with crash-rollback armed; dwell → commit.

Gate 0 — [Viktor]

  • Input 1 (byo caps): --cores 4 --memory 8192
  • Input 2 (Phase-6 repopulation): fresh + redeploy apps
  • 0c fingerprint confirmation vs password manager: CONFIRMED — match
  • Verdict: GO — with one pre-Phase-1 STOP + ruling (below)

Gate-0 deviation — the demo customer's empty git credentials (Viktor ruling → v1.11.2)

Pre-destroy check found demo-felhom's git.username/git.token still EMPTY (the GL-2 finding); the v1.11.1 installer would die at step 5/8. Viktor ruled: the artifacts are world-readable by design — anonymous fetch is the supported path. All 12 installer fetch targets (2 packages + 10 raw configs) validated HTTP 200 anonymously, then host-install v1.11.2 shipped mid-drill (felhom.eu 0640aa06): empty creds → loud WARN + anonymous fetch (conditional curl auth; the hub-vouched sha256 stays the integrity root; configured creds still used when present). Harness 26/26 (new GL6-ANON case; GL4-C2 assertion updated for the pinned constants).

Phase 1 — uninstall #1 (full mode, busy drive staged) — PASS with 2 findings

Real --uninstall --vmid 9201 (typed confirm, v1.11.2 fetched from the served URL):

  • Busy-drive behavior (the staged prop on felhom-usb): WARN + "Eject the drive via the dashboard, or stop the apps using it and re-run" — NO -l/-f, teardown CONTINUED; the other two drives unmounted ("data stays on the drive"); the root bind correctly left mounted with its own guard warn. The KEPT-vs-WIPED statement listed felhom-usb as "still mounted (busy — stop the apps and retry)" — i.e. the code takes the statement-lists-it-as-retry path; after releasing the prop a plain manual umount completed the job (no re-run needed).
  • KEPT-vs-WIPED statement: printed verbatim as designed (guest+volumes / agent / self-update artifacts / watchdog+OOB / pveum+pool / state = WIPED; drives+data / PBS backups / hub record / escrow = KEPT). The recovery-credential KEPT line correctly keyed off the state file — which on THIS box predates the 4b state marker, so the line was absent although a G1-vaulted credential exists (observation: state-keyed conditionals can under-report on pre-v1.10.0 installs).
  • 1b audit: agent binary+baks, selfupdate wrapper+slots, rollback unit+drop-in, watchdog set, mkfs wrapper, snippets, dnsmasq drop-ins, sudoers, state dirs, pveum user/roles/ACL/pool — ALL GONE. PVE printed a benign "ignore invalid acl role 'FelhomAgentStore'" notice during ACL removal (ordering noise, recorded).
  • Drive data intact: witness file re-hashed on the still-mounted primary path — 7b00bc3550b10e3952e0e0965f5a2628 == the Phase-0 value.

Findings

  • GL6-F1 (MED, security-adjacent): /etc/felhom-agent/ SURVIVES the uninstall holding agent.json.bak-* files (incl. the operator's same-day ceremony backup) — old configs with secret material (pve token now dead; the per-host hub api_key is still live because the hub record survives). The uninstall removes agent.json but not .bak-* siblings. Fix candidate: purge "$agent_cfg"* / the dir in 4b (v1.11.3). Files left in place pending Viktor's call (one is his own backup).
  • GL6-F2 (MED): the agent-created per-drive systemd mount units (mnt-felhom-usb.mount + flash + teszt_enroll) and their ACTIVE primary mounts at /mnt/<name> survive — the statement's "physically removable now" overclaims (the binds are gone but the partition is still mounted). Fix candidate: uninstall stops+removes the agent's mnt-* units for enrolled namespaces (data untouched). Left as-is for Phase 2 (the re-enroll behavior against pre-existing mounts is itself drill-relevant).

Gate 1 — [Viktor]: GO (delete the .bak secrets now)

GL6-F1 residue resolved at the gate: shred -u /etc/felhom-agent/agent.json.bak-* + rmdir (incl. Viktor's ceremony backup, confirmed disposable). F1/F2 kept as v1.11.3 fix candidates.

Phase 2 — byo-mode day-0 (Peti's exact path, on real hardware) — PASS (happy path), 3 findings

Command (Viktor typed the demo-felhom hostname ack — interactive consent leg): ./felhom-host-install.sh --customer-id demo-felhom --mode byo --vmid 9201 --cores 4 --memory 8192 --force-gitea-golden --acl-storages "local local-lvm" --passphrase-file <0600> (+ --allow-new-leaf --resume after F3/F4, see below).

End state — all green (verify step, verbatim highlights):

  • disclosure printed; Viktor's typed-hostname ack accepted (byo install acknowledged for host demo-felhom);
  • NO 4b / chpasswd / recovery-credential anywhere (grep of the full transcript = 0);
  • anonymous artifact fetch worked (v1.11.2): agent v0.76.0 sha 9828c5f7… verified, golden 0.103.0 sha 8481e8a1… verified, both against the hub manifest;
  • installed agent --version == 0.76.0; daemon active;
  • authz signers: 2 (operator-signed self-update armed) from first boot — key validation leg 3 ✓ (config carries operational:felhom-op-1 + recovery:felhom-rec-1);
  • pool member + all 4 scoped-ACL grants (user+token) present; caps on the guest: cores 4 / memory 8192;
  • byo config asserts pass (lan_resolver/wg_tunnel/oob all off);
  • controller Up (healthy) 0.103.0 + traefik + cloudflared up; dashboard HTTP 200 (Host-header probe); hub config-retrieve OK.

Findings — the byo install's resume path is fragile (happy single-shot path is clean)

  • GL6-F3 (MED, GL-7-relevant): on a BYO host carrying the OWNER's (non-Felhom) guests, a FRESH install trips the populated-host leaf guard (pct list | grep -q . counts ANY guest) and dies demanding --preserve-state-from/--allow-new-leaf — even though no Felhom controller pin exists to protect on a fresh install. Peti's box has his own guests → his day-0 command MUST include --allow-new-leaf (GL-7). Better fix candidate: narrow the guard to felhom_guests so a fresh byo install on a populated non-Felhom host doesn't trip.
  • GL6-F4 (HIGH, resume-correctness): --resume fully skips completed steps via should_skip X || step_X, but later steps consume the skipped steps' IN-MEMORY outputs, which are then empty → broken results: (a) enroll skipped → HOST_ID/HOST_API_KEY empty → step_agent_config writes a config missing hub.host_id (+ empty proxmox.token when token is skipped) → daemon crash-loops (hub.host_id is required); the mint-once token secret is unrecoverable, so recovery = re-run token+enroll. (b) grows skipped → ROOTFS_GROW/… empty → provision passes -rootfs-grow ""flag parse error. Note step_token even has an internal resume-guard ([[ -n "$PVE_TOKEN" ]]) that the || step_token dispatch DEFEATS — strong signal the dispatch is the bug (steps whose outputs feed later steps must re-run, or re-derive from the written config, on resume). Drill recovery: dropped the affected completed marks from state.json to force re-run — the install then completed cleanly. The single-shot happy path (no interruption) is unaffected, which is Peti's path.
  • GL6-F2 (from Phase 1, MED): the agent's per-drive mnt-*.mount units + primary /mnt/<name> mounts survived uninstall #1; on this reinstall the parent bind /mnt/felhom-drives is back but the drives are NOT yet bound under it — re-enrollment is the Phase-2b UI leg.

Gate 2 — [Viktor]: <pending — interactive bring-to-life legs below>

(Phases 36 appended as they run.)