Files
felhom.eu/scripts/CHANGELOG.md
T

27 KiB
Raw Blame History

Felhom scripts — Changelog

docs — golden 0.98.3 live: D.1b RETIRED, drill ledger B1/B5 → FIXED, backlog note resolved (2026-07-03)

Companion to felhom-agent's build-golden.sh v2.0.0 (@ ceca355): the golden now bakes the CURRENT controller (0.98.3, mandatory-tag convention — B5) and a felhom-controller-bootstrap.path unit (controller deploys on the bootstrap-mount hot-plug, no reboot — B1). Clean-room validated end-to-end BEFORE publish (bake → isolated hot-plug proof → local-golden Day-0 install → publish → vouch → --force-gitea-golden install); evidence: documentation/audits/DRILL-golden-098-2026-07-03.md.

  • documentation/runbooks/day0-install.md: D.1b reduced to a one-line version check (fresh boxes land current + self-manage); old manual-update procedure moved to Part F troubleshooting keyed on "golden older than 0.86.0 was vouched"; header validated-versions line → script v1.9.1 / agent v0.63.0 / golden v0.98.3; A.3 records the drilled-known-good pair + "vouch ≥ 0.98.3"; A.4 floor text rewritten (fresh boxes now floor-honoring; recommend raising the global floor after a golden rebuild — operator, 1 min).
  • documentation/audits/DRILL-day0-cleanroom-2026-07-03.md ledger: B1, B5 → FIXED (pointers); R6 note: the installer's post-provision reboot is retained as a belt, removal recorded as a candidate cleanup (not done).
  • documentation/backlog/FOLLOWUP-golden-default-controller-tag.md + backlog/README.md: RESOLVED per the M18/M19 convention (file kept + annotated; README entry marked FIXED).
  • New evidence doc: documentation/audits/DRILL-golden-098-2026-07-03.md (AD transcripts, unit states, resolution-order + fetch/sha proofs, cleanup, observations — incl. the SECURITY observation that the customer git.token has package-WRITE rights → scope-down + rotate follow-up).

docs — day0-install D.1b narrowed + drill ledger B2/B3 → FIXED agent v0.63.0 (2026-07-03)

Agent v0.63.0 fixed both drill findings (B3 token reload-on-miss + B2 guesthook snippets-dir mkdir; red-proofed, deployed on felhom-pve, Gitea-published sha256 b4a89c81…). Guide follow-through: the D.1b "restart the agent first" step is now CONDITIONAL (only for an installed agent < v0.63.0 — the Day-0 manifest still vouches 0.62.0, so today's fresh installs still hit it); the 401 troubleshooting row records the fix version; the drill ledger + go/no-go item 8 marked FIXED. Operator follow-up unchanged: vouch agent 0.63.0 in the Day-0 manifest UI, then the step is dead.

felhom-host-install.sh v1.9.1 — clean-room drill fixes: residue-free uninstall + post-provision reboot (2026-07-03)

Companion to the Day-0 go-live package (documentation/runbooks/day0-install.md + documentation/audits/DRILL-day0-cleanroom-2026-07-03.md). Every fix was found by the clean-room drill (virgin nested PVE 9.2.2) and re-verified there (v1.9.1 uninstall → zero-Felhom-residue diff vs the pre-install baseline; v1.9.1 install → controller up with no manual intervention).

  • Header/version sync (the header said v1.8.0 while SCRIPT_VERSION said 1.9.0); keep-in-sync note on SCRIPT_VERSION; usage sed range follows the header (2,95).
  • Uninstall now removes the drill-found residue (R1R5): the agent config (resolved from the unit's -config BEFORE the unit is removed — it holds the per-host hub api_key), the felhom-shared-parent unit + wants links + /usr/local/sbin/felhom-shared-parent.sh
    • the /mnt/felhom-drives self-bind/dir, /usr/local/sbin/felhom-mkfs-guarded, /var/lib/vz/snippets/felhom-guest-hook.sh, and /etc/dnsmasq.d/felhom-*.conf (+ dnsmasq restart when touched). All tolerate-absent; summary lines updated (sudo AND dnsmasq packages are the documented package remnants).
  • Post-provision guest reboot (R6): the golden's felhom-controller-bootstrap.service evaluates ConditionPathExists=/etc/felhom-bootstrap/bootstrap.json at BOOT, but the agent back-half hot-plugs the mount into the running guest — on slower hardware the first boot loses that race deterministically and the controller never deploys. step_provision now reboots the guest once (the agent's own output says "next: reboot the guest"); step_verify waits bounded (180 s) for the controller container instead of a momentary look.

felhom-host-install.sh v1.9.0 — Pool.Audit for the stale-lock reaper (A1) (2026-07-03)

Companion to felhom-agent v0.62.0 (audit A1: pool-membership ownership check). PVE_PRIVS_GUEST gains Pool.Audit (12 → 13 privs, granted at /pool/felhom via the existing FelhomAgentGuest role) so the agent can read GET /pools/felhom — its stale-lock reaper's ownership registry. Pool.Allocate does NOT satisfy the read (spike SPIKE-a1-pool-membership-read-2026-07-03 T2). No structural change: _ensure_role already role modifys to the exact priv set, so re-running --rescope-acl (or a fresh install) upgrades an existing box idempotently; remove_scoped_acl deletes by role name and needs nothing. Deploy order on a live box: rescope FIRST, then deploy agent v0.62.0 — the added read priv is harmless to an older agent, while the new agent on an old ACL fail-safes its reaper (skips) and reports pve:pool-read degraded until the rescope lands.

docs — SPIKE: A1 pool-membership read for the stale-lock reaper (2026-07-03)

Findings doc documentation/audits/SPIKE-a1-pool-membership-read-2026-07-03.md. Live-probed on felhom-pve under the PRODUCTION scoped token vs root: LXC enumeration IS already pool-filtered (token sees only 9201 of 4 guests); GET /pools/felhom 403s naming Pool.Audit; a throwaway token with ONLY Pool.Audit@/pool/felhom reads members (minimal delta proven, fully torn down); /cluster/resources withholds the pool field without Pool.Audit; local ownership records are all partial. Recommendation for the A1 impl spec: add Pool.Audit to PVE_PRIVS_GUEST in felhom-host-install.sh (L183) + a GET /pools/felhom cross-check in the agent's staleLockController.Guests(), fail-safe skip on read failure. No script/agent change in this commit — docs only. Appendix: committed-secrets (felhom.secret.yaml) rotation micro-runbook, operator follow-up.

install_skills.py — new: Claude Code skills installer (2026-07-03)

Installs skills/*/SKILL.md (felhom-build-deploy, felhom-ui-design, felhom-testing) into ~/.claude/skills/ as Windows junctions (mklink /J) so repo edits are live immediately; falls back to a full copy if junction creation fails or isn't followed (copy mode prints a re-run reminder). Idempotent — re-runs detect a correct junction and leave it. Verified: junctions ARE followed by Claude Code skill discovery (fresh-session probe found all three).

reuse_refs_check.py — new gate: REUSE.md citation checker (2026-07-03)

Staleness defense for the new per-repo REUSE.md reuse maps. Takes repo roots as argv, extracts every cited *.go/*.py/*.html/*.css/*.yml/*.yaml/*.sh path (slash-containing tokens only — bare filenames are conventions, not citations), verifies each exists; prints offenders, non-zero exit on any missing path. Symbols are spot-verified by the reviewer, not this script. Usage: python scripts/reuse_refs_check.py <repo-root> [...].

felhom-host-install.sh v1.8.0 — install the guarded-mkfs wrapper (Impl-1 Part B) (2026-07-01)

Companion to felhom-agent v0.54.0 (format-safety foundation). During agent install, fetch + install the guarded-mkfs wrapper so the agent's format path is safe on any box.

  • New step in step_agent_install: fetch configs/felhom-mkfs-guarded.sh from Gitea, bash -n validate, install -m0755 -o root -g root/usr/local/sbin/felhom-mkfs-guarded. Installed BEFORE the sudoers (which now allowlists ONLY the wrapper, not raw mkfs.*), so the ordering is gap-free.
  • The agent v0.54.0 sudoers (fetched by the same step) drops the raw mkfs.ext4 -F /dev/* / mkfs.xfs -f /dev/* allowlist and permits only felhom-mkfs-guarded /dev/* *, plus read-only pvs/zpool for the agent's unclaimed-disk guard. No other host-install change.
  • bash -n + shellcheck clean (0 new warnings). Live-validated on felhom-pve (agent v0.54.0 deploy): wrapper refuses the OS disk + an LVM-PV partition, raw mkfs is sudo-denied, an unclaimed throwaway disk formats; the agent guard's sudo reads (pvs/lsblk/zpool) all work as the felhom-agent user.

felhom-host-install.sh v1.7.0 — 3b-fix: Datastore.Audit box-wide (restore drive visibility) (2026-07-01)

Fixes a regression the v1.6.0 pool-scoped ACL introduced: Datastore.Audit was placed in the per-storage Store role (granted only on local/local-lvm/felhom-pbs), which excluded the enrolled removable drives felhom-usb/felhom-flash. The agent enumerates storage via ListStorage/NodeStorage (both gated by Datastore.Auditinternal/storage/observe.go), so it could no longer SEE the drives → false "Meghajtó leválasztva" (drive detached) alerts + drives absent from the agent-view. (The v1.6.0 swap's "felhom-usb → 403" was mis-read as blast-radius success; felhom-usb is Felhom's OWN customer drive, not an out-of-scope object.)

  • Datastore.Audit moved from Store → Base (PVE_PRIVS_BASE now "Sys.Audit SDN.Use Datastore.Audit"; PVE_PRIVS_STORE now "Datastore.Allocate Datastore.AllocateSpace"). Audit is read-only metadata, so box-wide Audit restores visibility of ALL storages (incl. dynamically-enrolled drives — no per-drive grant ever needed) while the write privs (Allocate/AllocateSpace) stay per-storage → write/allocate blast-radius containment is UNCHANGED. Confirmed at source: the agent creates no PVE storage (no POST /storage/pvesm add); drives are dir-storages it observes + mounts via host ops, so they need only Audit, never Allocate.
  • apply_scoped_acl reordered Base-before-Store (role + grant) so a RE-APPLY on a live box adds Audit@/ before Store drops its per-storage Audit → gap-free (the agent never loses enumeration).
  • remove_scoped_acl / --uninstall / --rescope-acl operate by role NAME and inherit the corrected privs automatically (no other change).
  • Live-repaired felhom-pve (two pveum role modify, Base first — no agent stop/restart): drives reappeared (agent-view 3→5 storages), detach alerts cleared. Re-tested under the scoped token: drives readable (was 403), write-containment intact (vzdump→felhom-usb still 403; out-of-pool guest 403), PBS Store grant unchanged. bash -n + shellcheck clean (0 new warnings).
  • NOT physically run (source-confirmed, no Datastore.Allocate in the path): a brand-new-drive UI enrollment (needs a spare USB) — the host-ops/Audit path is unchanged from pre-3b.

felhom-host-install.sh v1.6.0 — pool-scoped token ACL (3-role) + --rescope-acl retrofit (2026-07-01)

Colleague-safety batch #4 phase b (script half; agent half = v0.53.0). Moves the agent token's dangerous privileges off / (which spanned every guest + storage) to /pool/felhom + /storage/<targets>, so on a shared box the token can only touch Felhom's own guests + storages. Validated by documentation/audits/SPIKE-pool-scoped-acl-2026-07-01.md (PASS) — implemented here.

  • 3-role scoped ACL (step_token rewrite). Replaces the single FelhomAgent role granted at / with three roles, each granted to BOTH the user AND the token (privsep intersection): FelhomAgentGuest (VM.* + Pool.Allocate) @ /pool/felhom; FelhomAgentStore (Datastore.*) @ each of PVE_STORAGES (default local local-lvm felhom-pbs — the offsite PBS MUST be included, SPIKE residual #1; --acl-storages overrides); FelhomAgentBase (Sys.Audit SDN.Use) @ /. Helpers apply_scoped_acl/remove_scoped_acl/_grant/_ensure_role.
  • Pool before token. ensure_felhom_pool runs at the top of step_token (always, incl. --skip-provision) so /pool/felhom exists before it's granted on.
  • Re-install safety. step_token also removes the pre-3b broad / grant + FelhomAgent role if present (remove_old_broad_acl, tolerate-absent), so a re-install can't leave the old grant unioned with the scoped one. The post-provision pool_add_guest is gone (the agent's restore --pool makes the guest a member atomically — v0.53.0).
  • --rescope-acl retrofit (new mode, mirrors --adopt-pool): migrate an existing install — ensure the pool + guest membership, apply the scoped grants, THEN remove the old broad grant (add-before- remove: the token is never grant-less mid-migration). Prints the "now deploy agent ≥ v0.53.0" ordering reminder. Idempotent + dry-run-aware. SUPERVISED (run with the agent stopped — the scoped ACL and the pool-param agent are mutually dependent; §13 of the task).
  • --uninstall now removes the scoped grants + 3 roles AND the pre-3b broad grant/role (both tolerate-absent → works on either shape), keeping the pool delete-if-empty (v1.5.0).
  • Validated on felhom-pve (dry-run): T-A fresh install (pool-before-token, 3 roles once, scoped grants incl. /storage/felhom-pbs), --rescope-acl (add scoped → remove old FelhomAgent), T-F uninstall (old-shape cleanup + pool not-empty skip). bash -n + shellcheck clean (0 new warnings). The live rescope + agent swap is the supervised STOP — not run here.

felhom-host-install.sh v1.5.0 — felhom pool by default + --adopt-pool retrofit + uninstall teardown (2026-07-01)

Colleague-safety batch #4 phase a. Every Felhom-managed guest now joins a dedicated felhom pool for fleet uniformity (and as the environment the later pool-scoped ACL — 3b — will spike against). All pool ops run as root@pam from the installer, so there is NO agent/token/ACL change and zero permission-model risk (PVE_PRIVS untouched; the FelhomAgent token stays scoped at /).

  • New felhom pool default. step_provision calls ensure_felhom_pool (create if absent, idempotent) and, after a successful provision, adds the guest via pveum pool modify felhom -vms <vmid> (skip-if-already-member). New helpers pool_exists / pool_members / ensure_felhom_pool / pool_add_guest; const PVE_POOL="felhom". PVE 9 syntax + /pools JSON shape confirmed live before wiring (pveum pool add|delete|modify; pvesh get /pools[{poolid,comment}], /pools/<id>{members:[{vmid,…}]}).
  • --adopt-pool retrofit mode. Non-destructive: adds an EXISTING Felhom guest to the pool (creating it if needed), resolving the guest from --vmid else the recorded provisioned_vmid. Reuses the ours-check (/etc/felhom-bootstrap mount) — refuses a non-Felhom guest unless --force. Touches ONLY pool membership: never reconfigures/restarts the guest, never contacts the hub. Idempotent (skip-if-member).
  • --uninstall pool teardown (step 5b). After the pveum removal, deletes the felhom pool only if empty (a destroyed guest is auto-removed from its pool); a pool that still has members is left with a log_skip naming them. Not reached on the Spec-1 safe-skip path (other Felhom guests remain).
  • Validated on felhom-pve (dry-run + SAFE live): T-A fresh-install dry-run shows the pool create + membership lines; T-B live adopt of guest 9201pvesh get /pools/felhom lists 9201, guest still running, config unchanged (the demo node is now pool-uniform); re-run = no-op; T-B' non-Felhom vmid → refusal; T-C uninstall dry-run → "pool felhom not empty (members: 9201) — leaving it". bash -n + shellcheck clean (0 new warnings; the 2 pre-existing SC2015 in step_verify unchanged).
  • NOT changed: PVE_PRIVS, the ACL grants, the agent, the provision-call args. 3b (pool-scoped ACL
    • agent restore-into-pool under a scoped token) is the separate spike-gated task.

felhom-host-install.sh v1.4.0 — appliance CPU/RAM cap passthrough (--cores / --memory) (2026-07-01)

Colleague-safety batch #3 (host-install half; the mechanism is agent v0.52.0). Lets an operator cap the provisioned guest so a trial appliance on a SHARED production Proxmox doesn't pressure the colleague's existing guests.

  • --cores N / --memory M (MiB) — optional; passed through to the agent's --selftest=provision as -cores/-memory. 0/unset = keep the golden's baked sizes (unchanged behaviour). New vars CPU_CORES/MEM_MIB; usage() header gains an "Appliance cap (optional)" group.
  • Conditional passthroughstep_provision builds a cap_args array and appends the flags to BOTH the dry-run log and the real agent call only when set. An agent < v0.52.0 would reject an unknown flag, so the flags are never sent unless the operator opts in (see the deploy dependency below).
  • Pre-flight sanity WARN (soft, provision only) — if --cores > host nproc or --memory > host MemTotal, log_warn "the cap won't protect other guests"; never die (the operator may know better).
  • Deploy dependency: a fresh install using --cores/--memory needs the hub artifact manifest to serve agent ≥ v0.52.0.
  • Validated dry-run on felhom-pve: --cores 2 --memory 4096 --dry-run → provision command shows -cores 2 -memory 4096; without the flags → neither present; --cores 64 --memory 65536 → both WARN lines (host 4 cores / ~15771 MiB). bash -n + shellcheck clean (0 new warnings; the 2 pre-existing SC2015 in step_verify unchanged).

felhom-host-install.sh v1.3.0 — --uninstall (clean revert) + pre-flight guards (2026-07-01)

Colleague-safety batch #1+#2. Adds a first-class, guarded --uninstall teardown so an operator can cleanly back out of a trial install, plus three provision pre-flight guards that stop common footguns. Script-only; no agent/hub/controller change.

  • --uninstall (local host teardown — no hub contact, no passphrase). Reverses an install in the install-order's reverse: guest → agent(unit/sudoers/binary/state/user) → pveum(ACL,token,user,role) → golden(opt-in) → state file. Every mutation goes through run() so --dry-run prints the full plan and executes nothing. Safety:
    • Ours-check: refuses to destroy a guest that lacks the /etc/felhom-bootstrap bind mount (matched by the constant guest path, not a hardcoded mpN slot — on the demo host it's mp9), unless --force.
    • Typed confirmation: must type the vmid to confirm PERMANENT destruction (read from /dev/tty; skipped only under --dry-run, where nothing is destroyed).
    • Other-guests guard: if any OTHER Felhom guest remains, destroys only the target and leaves the agent + PVE token + state in place (re-run with --force to remove host-level anyway — orphans the others).
    • Never removes the sudo package; never contacts the hub (the host record intentionally persists).
    • Presence-checked + idempotent: an already-absent guest/unit/sudoers/binary/user/ACL/token/role is a tolerated skip, not an error. The pveum role delete runs only after its ACL grants are gone (PVE refuses to delete a referenced role). Confirmed PVE 9 ACL-delete form: pveum acl delete / --users|--tokens <x> --roles FelhomAgent.
    • Target vmid resolves from --vmid, else the recorded provisioned_vmid (else dies). A --vmid that disagrees with the recorded one needs --force.
    • --remove-golden: with --uninstall, also delete the golden vzdump from the archive storage (pvesm free); otherwise it is left in place.
  • Install state now records customer_id + provisioned_vmid (new _state_put/_state_get helpers, dry-run-guarded like _state_mark; the completed[] shape is untouched) so a later --uninstall resolves its target automatically and safely.
  • Pre-flight guards (provision mode):
    • Multi-node guard — on a 2+-node cluster, die (naming the nodes) unless --node is explicit (new NODE_EXPLICIT); single-node keeps the current auto-pick. No-op under --skip-provision.
    • Archive-storage-exists guard — verify --archive-storage appears in pvesm status (else die); no-op under --skip-provision.
    • RAM floor (WARN, never fatal) — warn when MemAvailable < 2048 MiB. All three run inside step_preflight (before any mutation) so they also fire under --dry-run.
  • Validated dry-run-only on felhom-pve (single-node, live guest 9201): T-A full uninstall plan, T-C not-ours refusal (red-proof), archive-missing die, RAM line, other-guests detector, state round-trip; confirmed 9201 + agent + pveum + state untouched after all dry-runs. bash -n + shellcheck clean (0 new warnings vs. baseline; the 2 pre-existing SC2015 in step_verify are unchanged). NOT yet live-validated (awaiting a supervised run): a real live --uninstall (guest destroy + pveum removal) and the multi-node guard on an actual cluster.

felhom-host-install.sh v1.2.0 — /dev/tty passphrase read + vmid auto-detect (2026-07-01)

Two operator-experience fixes so a colleague can install online (via the hub's new "Option 1: Online install" one-liner) and onto a host that already runs a guest at 9201.

  • Passphrase prompt reads from /dev/tty, not stdin (read_passphrase). read -rsp … < /dev/tty makes the no-echo prompt work regardless of how stdin is wired — both download-then-run and curl … | sudo bash (where stdin is the pipe). Strictly more correct; the --passphrase-file path is unchanged. The passphrase is still never on argv / in logs / in the state file.
  • VMID auto-detect (--vmid now optional-smart). New VMID_EXPLICIT flag (set by --vmid). The pre-flight vmid guard now determines "in use" against the pct list + qm list id-set (LXC and VMs share the id space — more complete than the old pct status, which only knew LXC):
    • explicit --vmid → unchanged deterministic behavior: die if the id is in use unless --force (destructive over-provision).
    • default 9201, in use, no --forceauto-pick the next free id (scan upward from 9201 over the used-set) and ask to confirm from the terminal (read … < /dev/tty, [y/N]); proceed on yes, die "no free vmid confirmed" otherwise. Never a silent auto-pick.
    • default 9201 + --force → over-provision 9201 (destructive) without prompting, as before.
    • New helpers used_vmids / _vmid_in_use / next_free_vmid. --vmid help text + usage() updated.

felhom-host-install.sh v1.1.0 — self-install the agent + fetch the golden from Gitea (2026-06-28)

The script now installs the agent itself (the last big manual Day-0 prerequisite is gone). It fetches the agent binary + golden from Gitea generic packages and verifies each against the hub-vouched artifact manifest before installing/using it. BUNDLE slice; pairs with hub v0.16.0 (artifact manifest endpoint + operator UI) and felhom-agent v0.43.0 (canonical unit + publish).

  • New step 5/8 agent install (before agent-config): resolves the manifest (GET /api/v1/artifacts/{id}, passphrase) + the git fetch token (from the customer's controller.yaml via config-retrieve — NO new credential); fetches /api/packages/admin/generic/felhom-agent/<ver>/felhom-agent, verifies sha256 vs the hub manifest (aborts on mismatch — verify-before-use), backs up any existing binary, installs 0755 /usr/local/bin/felhom-agent; ensures the non-root felhom-agent system user; installs the canonical sudoers (0440, visudo -cf-validated) + systemd unit; daemon-reload + enable. Idempotent: same version already installed + service active → skip.
  • --skip-provision: install + configure + verify the agent (incl. golden fetch+verify) but do NOT provision a guest — the agent-only path for re-installing/upgrading the agent on a host that already has live guests. Adds an agent-only step_verify_agent (binary + non-root service active + a --selftest=hub collect-report).
  • New step 7/8 golden: local auto-discovery stays the default/fallback; otherwise fetches /api/packages/admin/generic/felhom-golden/<ver>/golden.tar.zst, verifies sha256, and imports it into the archive storage's dump dir for the restore. --force-gitea-golden forces the Gitea path.
  • Non-root agent model: the agent now runs as felhom-agent with privileged.mode: "sudo" (was the dev/CI direct+root shortcut). The config is chowned to the service user (0600) so the daemon can read it; systemctl is-active after restart is the real proof the non-root user can read the config.
  • Pre-flight relaxed: a missing agent binary is no longer fatal (step 5 installs it); the local golden requirement is deferred to step 7.
  • Trust model: checksum trust root = the hub (manifest), not Gitea; the fetch credential is the existing config-retrieve git token; artifacts are pinned to a version (never :latest).
  • Secrets: the git token is a never-logged runtime carrier (cleared on EXIT alongside the passphrase / pve-token / hub api_key); the sudoers is 0440 and visudo -cf-validated before install.
  • bash -n + shellcheck clean.

felhom-host-install.sh v1.0.0 — Day-0 host bootstrap (provision mode) (2026-06-26)

First release. A single operator-run script that automates Day-0 on a freshly-PVE-installed host: Proxmox API token → hub host enrollment (option C, single secret) → agent config → guest provision → verify. Composes proven mechanisms (the pveum role/token sequence, hub POST /host-enroll, felhom-agent --selftest=provision); grounded by documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md.

  • 7 steps, idempotent + resumable via /var/lib/felhom-install/state.json: pre-flight → Proxmox token → compute grows → host-enroll → agent config → provision → verify.
  • Single-secret (the retrieval passphrase): read no-echo or from a 0600 file, never on argv/logs/state. The global operator key never touches the box.
  • pveum automation: 16-priv FelhomAgent role (create-or-modify), felhom-agent@pve user, privsep token (reuse-if-working else rotate), and both ACL grants applied after the token exists (token-remove purges the token ACL).
  • Auto-discovery: golden archive (newest vzdump-lxc-<golden-vmid>), PVE node name, vmbr0 bridge IP for the local-api, and the served-leaf TLS fingerprint pin.
  • Safety: pre-flight fails fast (root, PVE 9.x, local-lvm headroom, hub reachable, customer+passphrase valid via read-only GET /config/{id}, golden resolvable); refuses to clobber an existing --vmid without --force; --dry-run previews every mutation; --preserve-from keeps operator infra (PBS/local_api/privileged/authz) on re-deploys.
  • --mode dr: documented 10D stub (restore customer PBS snapshot instead of golden) — not implemented.
  • Live-validated end-to-end on felhom-pve: authorized wipe of demo guest 9201 → re-provision from the golden → controller config-pull + public tunnel HTTP 200 → host-report of guest 9201 → idempotent --resume no-op. (One ordering bug — token ACL applied before rotation — was found and fixed during the live run.)