Files
felhom.eu/scripts/CHANGELOG.md
T
admin f7cc6a720e GL-4: uninstaller gap-closure + operator-key day-0 fold (host-install v1.11.0)
Uninstall: NEW 4b4 removes the self-update artifacts (guarded wrapper,
.prev/.new.* A/B slots, rollback unit, limits drop-in — derived from
configs/felhom-selfupdate-guarded); enrolled drives under /mnt/felhom-drives
unmounted plain-umount-only (busy = warn+guidance, never -l/-f; root bind
guarded); both modes end with a KEPT-vs-WIPED statement (drives/PBS/hub
record/escrow/vaulted recovery credential live on); guest-only mode prints
the vmid's bind-store drives.

Key-pin: OPERATOR_KEY_* constants (empty until the operator ceremony) +
--operator-pubkey-file (argv-validated; comment=key_id required; file
overrides constants) -> authz.signers per the agent SignerKey schema;
no-keys-resolved KEEPS preserved signers; verify reports armed/dormant WARN.

Harness: +13 static cases + GL4-D parity + GL4-INV + PVE-tier GL4 H-U
(full-uninstall dry transcript vs live 9201). 28/28 on felhom-pve;
red-proofs RP-1..3 run->fail->revert. GO-LIVE-PACKAGE.md ABSENT AGAIN ->
status in CONTEXT.md. Live teardown/armed-pin = GL-6.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
2026-07-08 08:14:07 +02:00

41 KiB
Raw Blame History

Felhom scripts — Changelog

felhom-host-install v1.11.0 — uninstaller gap-closure + operator-key day-0 fold (TASK GL-4, go-live G6 + G1-follow-up) (2026-07-08)

  • Uninstall gap-closure (G6):
    • NEW 4b4 block removes the self-update artifacts the install lays down but v1.10.0 never removed: /usr/local/sbin/felhom-selfupdate-guarded, the A/B slot files next to the live binary (felhom-agent.prev + orphaned .new.* temps), felhom-agent-rollback.service, and the felhom-agent.service.d/felhom-agent-limits.conf drop-in (+dir). Paths derived from the authoritative felhom-agent/configs/felhom-selfupdate-guarded; pending.json was already covered by the state-dir removal. Tolerate-absent, 4b2 shape.
    • Enrolled/network drives mounted under /mnt/felhom-drives/<name> are now unmounted (deepest first) before the root self-bind — plain umount ONLY, never -l/-f: a busy mount gets a warning + "eject via the dashboard or stop the apps and retry" and the root bind is then left alone (previously a child mount made the root umount die mid-teardown). The data is NEVER touched — no wipe/format path exists anywhere near /mnt/felhom-drives.
    • Both modes now END with an explicit KEPT-vs-WIPED statement: WIPED mirrors what the mode actually ran; KEPT names the drives + their data (physically removable; busy ones listed as retry), PBS backups + customer namespace (conditional on a pbs storage), the hub host/customer record, the escrow blob, and — when step 4b had vaulted one — the root@pam recovery credential.
    • Guest-only mode prints the drives the agent's bind store records for that vmid (best-effort from /var/lib/felhom-agent/guest-binds.json, generic note otherwise) + "eject BEFORE uninstall" guidance; no umounts there (remaining guests may use the drives).
  • Operator-key day-0 fold (the GL-1/G1 key-pin follow-up): OPERATOR_KEY_{OPERATIONAL,RECOVERY}_{ID,LINE} constants (EMPTY until the operator's offline pin ceremony) + --operator-pubkey-file PATH (lines operational|recovery <authorized_keys line>; the comment field IS the key_id, required; malformed/empty file dies at argv time naming the line; file OVERRIDES the constants with a notice). Resolved keys are written to authz.signers in the agent config (exact config.go SignerKey schema; nonce_store_path left to the agent default). Preserve rule: script/file keys are authoritative ONLY when non-empty — a reinstall with no keys KEEPS a preserved config's signers (never silently clobber a manually-pinned box; replacing preserved signers logs a pin-rotation notice). byo disclosure names the operational key_id (or "NONE — dormant"); verify reports authz signers: N (armed) or a dormant WARN (dormant = safe default).
  • Harness: +13 static cases (GL4-C1..C5 incl. 5 live key-file refusals + grep shapes, GL4-D disclosure↔uninstall parity, GL4-INV no-forced-unmount/no-format invariant) + PVE-tier GL4 H-U (full uninstall DRY transcript against the live guest, state-override protected). 28/28 PASS on felhom-pve; red-proofs RP-1 (4b4 dropped → parity FAILs), RP-2 (preserve guard dropped → C4 FAILs), RP-3 (role-die dropped → C3a FAILs) all run→fail→revert on scratch copies.
  • NOT live-validated (GL-6): the real (non-dry) uninstall with mounted/busy drives, the statement on a real teardown, and an armed key-pin install end-to-end.

felhom-host-install v1.10.0 — --mode appliance|byo install profile (TASK GL-2, go-live G2/G4/G5) (2026-07-07)

  • --mode appliance|byo is now REQUIRED for a fresh install / --resume (no default — the build-golden v2.0.0 precedent). appliance = the historical behavior, dry-transcript-preserved. byo = a Proxmox host the operator does NOT own (the pilot): step_break_glass is gated OFF at its call site (root@pam never reset, nothing vaulted — the one place the skip is auditable), --cores + --memory both mandatory (the only noisy-neighbor protection there), --enable-oob/--rotate-recovery refused at argv time, PVE-major gate (die on non-9.x unless --force), every --acl-storages entry must exist on the box (die naming ALL missing; appliance warns), informational host-DNS :53 line, full host-mutation disclosure + typed-hostname acknowledgement (/dev/tty, dry-run-aware), and post-write config asserts that refuse lan_resolver.enable/wg_tunnel.enabled/oob.enabled = true (byo also flips the config write's lan_resolver DEFAULT to off; appliance keeps default-on).
  • The old --mode provision value is retired (specific die message); the dr stub seam is kept. The mode is persisted in state.json; a --resume under the other mode refuses (C4).
  • Verify asserts (BOTH modes — campaign-2 R2 lesson): the verify step now asserts pool membership of the provisioned vmid + every scoped ACL grant (user AND token: Base@/, Guest@/pool/felhom, Store@each storage); any miss → error + "WITH WARNINGS" verdict.
  • --preflight-only: runs the full mode-aware step-1 checks, prints an explicit PRE-FLIGHT PASS/FAIL verdict, exits 0/1, writes NO state marks (never lets a later --resume skip the real preflight). Existing-guest count line added to preflight (both modes).
  • FELHOM_INSTALL_STATE_DIR env override (test-harness only) so harness cases can never touch a live state.json; usage() switched to a marker-based range (the numeric one had drifted).
  • NEW scripts/hostinstall-mode-harness.sh: static tier (bash -n, shellcheck, C1C4 refusal matrix asserting exact die messages, grep-invariants: one gated break-glass call site, chpasswd unreachable outside it, --mode in usage) + PVE tier (C5 live-refusal + Scenario A/B dry transcripts via FELHOM_TEST_CUSTOMER/FELHOM_TEST_PASSFILE). Red-proofs RP-1..RP-3 executed (gate removed / caps requirement dropped / mode-mismatch check dropped → matching case FAILs).
  • Live-validated read-only on felhom-pve: full static tier + C5 + byo dry transcript (no 4b/8/chpasswd/recovery-credential; disclosure+ack+DNS+storage lines present) + --preflight-only PASS and FAIL verdicts with zero state entries. Full A/B rc=0 transcripts + C7 await the supervised GL-6 drill (blocked today by demo-felhom's EMPTY git.username/git.token — any step-5 re-run for that customer dies until the operator sets real Gitea read creds).

felhom-host-install — --enable-oob installs the dedicated felhom-sshd OOB instance + belt (TASK H1) (2026-07-05)

  • --enable-oob + install_oob (in step 5): lays down the STATIC OOB parts — /etc/felhom-sshd tree
    • a dedicated ed25519 host key + the felhom-op login user + its scoped /etc/sudoers.d/felhom-op (visudo-validated) + the felhom-sshd.service unit (RuntimeDirectory-guarded — refuses the G1 incident cause) enabled-not-started (the agent renders the config) + the static inet felhom_oob nft table (nft -c-checked) + its boot loader felhom-oob-nft.service. The agent reconciles the DYNAMIC parts (port claim, config, operator authorized_keys, belt set elements) once oob.enabled.
  • Uninstall (step 4b3): stops/disables felhom-sshd + the belt loader, removes both units + the config tree + belt table + felhom-op user/sudoers; NEVER touches the stock sshd/:22//etc/ssh.

felhom-host-install — break-glass credential + management-plane watchdog install (TASK G1) (2026-07-05)

  • step_break_glass (new step 4b): generates a strong root@pam console password (openssl rand, NEVER logged/filed — stdin→chpasswd + stdin→curl), sets it, and vaults it to the hub over the enroll-authenticated host-key channel (PUT /hosts/{id}/recovery-credential). Idempotent (state marker) unless --rotate-recovery. The human break-glass path for the PVE web console.
  • install_mgmt_watchdog (in step 5): installs the G1 host artifacts from the agent repo configs — felhom-privsep.tmpfiles (boot-persistent /run/sshd), felhom-mgmt-watchdog.{sh,service,timer} (agent-independent ~60s privsep-dir auto-heal). HARD GUARD: refuses any fetched unit declaring RuntimeDirectory= (the SPIKE-felhom-sshd §8 incident cause). Non-fatal on a repo that predates the artifacts. Uninstall (step 4b2) stops+disables the timer and removes all four artifacts + the marker.

docs — 06-doc S3 row SHIPPED + agent-side revocation semantics (2026-07-04)

Docs-only companion to felhom-agent v0.64.0 (the S3 slice — keygen, registration, agent-managed wg-quick@wg-felhom, escrow join; live-validated on felhom-pve incl. revocation drill, reboot persistence, 30-min soak). 06-doc §3.5 now records: register-once marker, revoked-stays-revoked, re-add via the registration endpoint (the raw registry add doesn't bump the host generation — live finding), wg_tunnel.enabled default-FALSE rollout gate. S6 backlog notes added (hub poll constant configurable + first-adoption log; registry-add bump-or-label). CGNAT/mobile-hotspot appendix deferred (operator-assisted; §7's open validation stands).

felhom-peersync.sh v1.0.1 — strip out of process substitution (exit-swallow fix) (2026-07-04)

The S1 REPORT's exit-swallow class, fixed: wg syncconf wg0 <(wg-quick strip "$tmp") hid the strip exit code — a corrupt wg0.conf.head could feed syncconf empty/partial input that WIPES the live peer set while the script exits 0 (and the bad conf is then persisted). v1.0.1 runs strip as its own step into $tmpdir/stripped; a strip failure aborts BEFORE wg is invoked. Sandbox red-proof (stub wg-quick exiting 1 after partial output + recording stub wg): pre-fix shape invoked wg and returned rc=0; fixed shape errors first, wg never called. Redeployed to the dev endpoint (runbook step 5 install); shellcheck clean.

felhom-peersync.sh v1.0.0 — the offsite endpoint's WG reconcile script (2026-07-04)

S1 (doc 06 §5): the forced-command target the hub's wgsync pushes to (runbook offsite-endpoint.md step 5 installs it as /usr/local/bin/felhom-peersync, root:root 0755, invoked via a one-line sudoers grant from the felhom-peersync user's restrict,command= authorized_keys entry). Validate-FIRST design: jq contract check (version 1, interface wg0, 44-b64 pubkeys, 10.77.0.x/32 allowed_ips, never the endpoint's own .1) rejects on stderr with exit 1 before touching anything; then head-file + generated [Peer] blocks into a same-fs tmp, wg syncconf <(wg-quick strip …) from the TMP (exact-match: adds/removes without bouncing the interface), and only on success the atomic mv to /etc/wireguard/wg0.conf — runtime and boot config can never diverge in the failure direction. Zero-peer payload = valid wipe. Never reads or prints the private key; no wg-quick save; no second mode. shellcheck-clean. Live-proven on felhom-hetzner incl. the negatives (malformed JSON / bad pubkey / own-IP peer → exit 1, wg state byte-identical) and reboot persistence.

docs — architecture Part 06: offsite connectivity design-of-record (2026-07-03)

documentation/architecture/06-offsite-connectivity.md — the settled offsite-backup-transport design, authored from the spike verdict + operator-resolved forks (recorded, not re-litigated): plain WG (D1), host-side agent-managed wg-felhom as the agent-managed-unit pilot on the sudoers *.mount install pattern (D2), one shared hub-driven endpoint VM running WG + the offsite PBS with no agent (D3, CF-token pattern), hub source-of-truth with a wireguard block riding the existing WireDesiredState/DesiredGeneration channel (D4), one datastore + per-customer namespaces (D5), and PBS on the VM — relay-through-DooPlex rejected as non-scaling (D6). Includes the Day-0 join handshake, the robustness set (NOT-DynDNS roaming, endpoint DNS re-resolve watchdog, MTU 1420, per-/32 topological isolation, tunnel-health through the storage-target reachability model), trust conformance, the honest open ledger (CGNAT unmeasured → mobile-hotspot smoke test; peer-sync push-vs-pull = slice-1 design point), and the S1S6 slice roadmap (MVP = S1→S2→S3, then S4; S5 merges with DR-completeness). All claims cited at file:line against felhom.eu @ bf099f6 + felhom-agent @ 4ba1b14. day0-install.md backlog line now points at spike + design doc. Docs-only.

docs — SPIKE: offsite-backup connectivity — plain WireGuard WINS the ladder; Headscale = separable fleet layer (2026-07-03)

documentation/audits/SPIKE-connectivity-wireguard-2026-07-03.md — the offsite-backup transport decision, empirically grounded on both real ends (demo-felhom PVE host ⟷ throwaway Hetzner box, end state: powered off, secrets shredded). Headline results: the operator's line is plain-NAT with a fixed public IP, not CGNAT, and has zero IPv6 (P0 honesty — CGNAT confirmation deferred to Peti's VM 110); plain outbound WG held an 11.4-min idle window and carried a real 2 GiB worst-case PBS backup at 4.26 MiB/s = the full home uplink (~5% tunnel overhead), TLS pin intact through the tunnel (positive + negative proof); UDP 51820 and 443 both pass; kernel WG surprisingly works inside the unprivileged guest (P7 — host placement stands on architecture, not infeasibility). Recommendation: host-side agent-managed WG (cloudflared pattern), key in the escrowed IdentityBundle, small public endpoint VM (PBS-on-VM vs rendezvous-relay deferred to the spec). runbooks/day0-install.md backlog line resolved to point here; CONTEXT.md notes the DR-completeness task is unblocked (next: the production connectivity spec).

skills — NEW: felhom-app-catalog (4th skill) + SparkyFitness as its worked example (2026-07-03)

skills/felhom-app-catalog/SKILL.md — the catalog authoring workflow (research → inspect the image for the healthcheck family → write compose/.felhom.yml → deploy live through the dashboard → verify healthy → reconcile the app count). Deliberately points at app-catalog REUSE.md §12 + README.md §format for every field table (one-fact-one-place; no duplication). Unique content: the never-guess-the-healthcheck rule with the per-tool image-inspection loop (BusyBox ash command -v gotcha: it silently ignores all but its first argument — verified), the probe-container naming rule (controller probes the container named exactly like the stack — verified in felhom-controller/internal/stacks/healthprobe.go, row added to app-catalog REUSE.md), the Hungarian-quote YAML kill, and the deploy-is-the-test doctrine. No installer change needed — install_skills.py auto-discovers skills/*/SKILL.md; fresh-session discovery probe listed all 4. Proven by finalizing sparkyfitness end-to-end on demo (see app-catalog-felhom.eu CHANGELOG).

docs — golden 0.98.3 live: D.1b RETIRED, drill ledger B1/B5 → FIXED, backlog note resolved (2026-07-03)

Companion to felhom-agent's build-golden.sh v2.0.0 (@ ceca355): the golden now bakes the CURRENT controller (0.98.3, mandatory-tag convention — B5) and a felhom-controller-bootstrap.path unit (controller deploys on the bootstrap-mount hot-plug, no reboot — B1). Clean-room validated end-to-end BEFORE publish (bake → isolated hot-plug proof → local-golden Day-0 install → publish → vouch → --force-gitea-golden install); evidence: documentation/audits/DRILL-golden-098-2026-07-03.md.

  • documentation/runbooks/day0-install.md: D.1b reduced to a one-line version check (fresh boxes land current + self-manage); old manual-update procedure moved to Part F troubleshooting keyed on "golden older than 0.86.0 was vouched"; header validated-versions line → script v1.9.1 / agent v0.63.0 / golden v0.98.3; A.3 records the drilled-known-good pair + "vouch ≥ 0.98.3"; A.4 floor text rewritten (fresh boxes now floor-honoring; recommend raising the global floor after a golden rebuild — operator, 1 min).
  • documentation/audits/DRILL-day0-cleanroom-2026-07-03.md ledger: B1, B5 → FIXED (pointers); R6 note: the installer's post-provision reboot is retained as a belt, removal recorded as a candidate cleanup (not done).
  • documentation/backlog/FOLLOWUP-golden-default-controller-tag.md + backlog/README.md: RESOLVED per the M18/M19 convention (file kept + annotated; README entry marked FIXED).
  • New evidence doc: documentation/audits/DRILL-golden-098-2026-07-03.md (AD transcripts, unit states, resolution-order + fetch/sha proofs, cleanup, observations — incl. the SECURITY observation that the customer git.token has package-WRITE rights → scope-down + rotate follow-up).

docs — day0-install D.1b narrowed + drill ledger B2/B3 → FIXED agent v0.63.0 (2026-07-03)

Agent v0.63.0 fixed both drill findings (B3 token reload-on-miss + B2 guesthook snippets-dir mkdir; red-proofed, deployed on felhom-pve, Gitea-published sha256 b4a89c81…). Guide follow-through: the D.1b "restart the agent first" step is now CONDITIONAL (only for an installed agent < v0.63.0 — the Day-0 manifest still vouches 0.62.0, so today's fresh installs still hit it); the 401 troubleshooting row records the fix version; the drill ledger + go/no-go item 8 marked FIXED. Operator follow-up unchanged: vouch agent 0.63.0 in the Day-0 manifest UI, then the step is dead.

felhom-host-install.sh v1.9.1 — clean-room drill fixes: residue-free uninstall + post-provision reboot (2026-07-03)

Companion to the Day-0 go-live package (documentation/runbooks/day0-install.md + documentation/audits/DRILL-day0-cleanroom-2026-07-03.md). Every fix was found by the clean-room drill (virgin nested PVE 9.2.2) and re-verified there (v1.9.1 uninstall → zero-Felhom-residue diff vs the pre-install baseline; v1.9.1 install → controller up with no manual intervention).

  • Header/version sync (the header said v1.8.0 while SCRIPT_VERSION said 1.9.0); keep-in-sync note on SCRIPT_VERSION; usage sed range follows the header (2,95).
  • Uninstall now removes the drill-found residue (R1R5): the agent config (resolved from the unit's -config BEFORE the unit is removed — it holds the per-host hub api_key), the felhom-shared-parent unit + wants links + /usr/local/sbin/felhom-shared-parent.sh
    • the /mnt/felhom-drives self-bind/dir, /usr/local/sbin/felhom-mkfs-guarded, /var/lib/vz/snippets/felhom-guest-hook.sh, and /etc/dnsmasq.d/felhom-*.conf (+ dnsmasq restart when touched). All tolerate-absent; summary lines updated (sudo AND dnsmasq packages are the documented package remnants).
  • Post-provision guest reboot (R6): the golden's felhom-controller-bootstrap.service evaluates ConditionPathExists=/etc/felhom-bootstrap/bootstrap.json at BOOT, but the agent back-half hot-plugs the mount into the running guest — on slower hardware the first boot loses that race deterministically and the controller never deploys. step_provision now reboots the guest once (the agent's own output says "next: reboot the guest"); step_verify waits bounded (180 s) for the controller container instead of a momentary look.

felhom-host-install.sh v1.9.0 — Pool.Audit for the stale-lock reaper (A1) (2026-07-03)

Companion to felhom-agent v0.62.0 (audit A1: pool-membership ownership check). PVE_PRIVS_GUEST gains Pool.Audit (12 → 13 privs, granted at /pool/felhom via the existing FelhomAgentGuest role) so the agent can read GET /pools/felhom — its stale-lock reaper's ownership registry. Pool.Allocate does NOT satisfy the read (spike SPIKE-a1-pool-membership-read-2026-07-03 T2). No structural change: _ensure_role already role modifys to the exact priv set, so re-running --rescope-acl (or a fresh install) upgrades an existing box idempotently; remove_scoped_acl deletes by role name and needs nothing. Deploy order on a live box: rescope FIRST, then deploy agent v0.62.0 — the added read priv is harmless to an older agent, while the new agent on an old ACL fail-safes its reaper (skips) and reports pve:pool-read degraded until the rescope lands.

docs — SPIKE: A1 pool-membership read for the stale-lock reaper (2026-07-03)

Findings doc documentation/audits/SPIKE-a1-pool-membership-read-2026-07-03.md. Live-probed on felhom-pve under the PRODUCTION scoped token vs root: LXC enumeration IS already pool-filtered (token sees only 9201 of 4 guests); GET /pools/felhom 403s naming Pool.Audit; a throwaway token with ONLY Pool.Audit@/pool/felhom reads members (minimal delta proven, fully torn down); /cluster/resources withholds the pool field without Pool.Audit; local ownership records are all partial. Recommendation for the A1 impl spec: add Pool.Audit to PVE_PRIVS_GUEST in felhom-host-install.sh (L183) + a GET /pools/felhom cross-check in the agent's staleLockController.Guests(), fail-safe skip on read failure. No script/agent change in this commit — docs only. Appendix: committed-secrets (felhom.secret.yaml) rotation micro-runbook, operator follow-up.

install_skills.py — new: Claude Code skills installer (2026-07-03)

Installs skills/*/SKILL.md (felhom-build-deploy, felhom-ui-design, felhom-testing) into ~/.claude/skills/ as Windows junctions (mklink /J) so repo edits are live immediately; falls back to a full copy if junction creation fails or isn't followed (copy mode prints a re-run reminder). Idempotent — re-runs detect a correct junction and leave it. Verified: junctions ARE followed by Claude Code skill discovery (fresh-session probe found all three).

reuse_refs_check.py — new gate: REUSE.md citation checker (2026-07-03)

Staleness defense for the new per-repo REUSE.md reuse maps. Takes repo roots as argv, extracts every cited *.go/*.py/*.html/*.css/*.yml/*.yaml/*.sh path (slash-containing tokens only — bare filenames are conventions, not citations), verifies each exists; prints offenders, non-zero exit on any missing path. Symbols are spot-verified by the reviewer, not this script. Usage: python scripts/reuse_refs_check.py <repo-root> [...].

felhom-host-install.sh v1.8.0 — install the guarded-mkfs wrapper (Impl-1 Part B) (2026-07-01)

Companion to felhom-agent v0.54.0 (format-safety foundation). During agent install, fetch + install the guarded-mkfs wrapper so the agent's format path is safe on any box.

  • New step in step_agent_install: fetch configs/felhom-mkfs-guarded.sh from Gitea, bash -n validate, install -m0755 -o root -g root/usr/local/sbin/felhom-mkfs-guarded. Installed BEFORE the sudoers (which now allowlists ONLY the wrapper, not raw mkfs.*), so the ordering is gap-free.
  • The agent v0.54.0 sudoers (fetched by the same step) drops the raw mkfs.ext4 -F /dev/* / mkfs.xfs -f /dev/* allowlist and permits only felhom-mkfs-guarded /dev/* *, plus read-only pvs/zpool for the agent's unclaimed-disk guard. No other host-install change.
  • bash -n + shellcheck clean (0 new warnings). Live-validated on felhom-pve (agent v0.54.0 deploy): wrapper refuses the OS disk + an LVM-PV partition, raw mkfs is sudo-denied, an unclaimed throwaway disk formats; the agent guard's sudo reads (pvs/lsblk/zpool) all work as the felhom-agent user.

felhom-host-install.sh v1.7.0 — 3b-fix: Datastore.Audit box-wide (restore drive visibility) (2026-07-01)

Fixes a regression the v1.6.0 pool-scoped ACL introduced: Datastore.Audit was placed in the per-storage Store role (granted only on local/local-lvm/felhom-pbs), which excluded the enrolled removable drives felhom-usb/felhom-flash. The agent enumerates storage via ListStorage/NodeStorage (both gated by Datastore.Auditinternal/storage/observe.go), so it could no longer SEE the drives → false "Meghajtó leválasztva" (drive detached) alerts + drives absent from the agent-view. (The v1.6.0 swap's "felhom-usb → 403" was mis-read as blast-radius success; felhom-usb is Felhom's OWN customer drive, not an out-of-scope object.)

  • Datastore.Audit moved from Store → Base (PVE_PRIVS_BASE now "Sys.Audit SDN.Use Datastore.Audit"; PVE_PRIVS_STORE now "Datastore.Allocate Datastore.AllocateSpace"). Audit is read-only metadata, so box-wide Audit restores visibility of ALL storages (incl. dynamically-enrolled drives — no per-drive grant ever needed) while the write privs (Allocate/AllocateSpace) stay per-storage → write/allocate blast-radius containment is UNCHANGED. Confirmed at source: the agent creates no PVE storage (no POST /storage/pvesm add); drives are dir-storages it observes + mounts via host ops, so they need only Audit, never Allocate.
  • apply_scoped_acl reordered Base-before-Store (role + grant) so a RE-APPLY on a live box adds Audit@/ before Store drops its per-storage Audit → gap-free (the agent never loses enumeration).
  • remove_scoped_acl / --uninstall / --rescope-acl operate by role NAME and inherit the corrected privs automatically (no other change).
  • Live-repaired felhom-pve (two pveum role modify, Base first — no agent stop/restart): drives reappeared (agent-view 3→5 storages), detach alerts cleared. Re-tested under the scoped token: drives readable (was 403), write-containment intact (vzdump→felhom-usb still 403; out-of-pool guest 403), PBS Store grant unchanged. bash -n + shellcheck clean (0 new warnings).
  • NOT physically run (source-confirmed, no Datastore.Allocate in the path): a brand-new-drive UI enrollment (needs a spare USB) — the host-ops/Audit path is unchanged from pre-3b.

felhom-host-install.sh v1.6.0 — pool-scoped token ACL (3-role) + --rescope-acl retrofit (2026-07-01)

Colleague-safety batch #4 phase b (script half; agent half = v0.53.0). Moves the agent token's dangerous privileges off / (which spanned every guest + storage) to /pool/felhom + /storage/<targets>, so on a shared box the token can only touch Felhom's own guests + storages. Validated by documentation/audits/SPIKE-pool-scoped-acl-2026-07-01.md (PASS) — implemented here.

  • 3-role scoped ACL (step_token rewrite). Replaces the single FelhomAgent role granted at / with three roles, each granted to BOTH the user AND the token (privsep intersection): FelhomAgentGuest (VM.* + Pool.Allocate) @ /pool/felhom; FelhomAgentStore (Datastore.*) @ each of PVE_STORAGES (default local local-lvm felhom-pbs — the offsite PBS MUST be included, SPIKE residual #1; --acl-storages overrides); FelhomAgentBase (Sys.Audit SDN.Use) @ /. Helpers apply_scoped_acl/remove_scoped_acl/_grant/_ensure_role.
  • Pool before token. ensure_felhom_pool runs at the top of step_token (always, incl. --skip-provision) so /pool/felhom exists before it's granted on.
  • Re-install safety. step_token also removes the pre-3b broad / grant + FelhomAgent role if present (remove_old_broad_acl, tolerate-absent), so a re-install can't leave the old grant unioned with the scoped one. The post-provision pool_add_guest is gone (the agent's restore --pool makes the guest a member atomically — v0.53.0).
  • --rescope-acl retrofit (new mode, mirrors --adopt-pool): migrate an existing install — ensure the pool + guest membership, apply the scoped grants, THEN remove the old broad grant (add-before- remove: the token is never grant-less mid-migration). Prints the "now deploy agent ≥ v0.53.0" ordering reminder. Idempotent + dry-run-aware. SUPERVISED (run with the agent stopped — the scoped ACL and the pool-param agent are mutually dependent; §13 of the task).
  • --uninstall now removes the scoped grants + 3 roles AND the pre-3b broad grant/role (both tolerate-absent → works on either shape), keeping the pool delete-if-empty (v1.5.0).
  • Validated on felhom-pve (dry-run): T-A fresh install (pool-before-token, 3 roles once, scoped grants incl. /storage/felhom-pbs), --rescope-acl (add scoped → remove old FelhomAgent), T-F uninstall (old-shape cleanup + pool not-empty skip). bash -n + shellcheck clean (0 new warnings). The live rescope + agent swap is the supervised STOP — not run here.

felhom-host-install.sh v1.5.0 — felhom pool by default + --adopt-pool retrofit + uninstall teardown (2026-07-01)

Colleague-safety batch #4 phase a. Every Felhom-managed guest now joins a dedicated felhom pool for fleet uniformity (and as the environment the later pool-scoped ACL — 3b — will spike against). All pool ops run as root@pam from the installer, so there is NO agent/token/ACL change and zero permission-model risk (PVE_PRIVS untouched; the FelhomAgent token stays scoped at /).

  • New felhom pool default. step_provision calls ensure_felhom_pool (create if absent, idempotent) and, after a successful provision, adds the guest via pveum pool modify felhom -vms <vmid> (skip-if-already-member). New helpers pool_exists / pool_members / ensure_felhom_pool / pool_add_guest; const PVE_POOL="felhom". PVE 9 syntax + /pools JSON shape confirmed live before wiring (pveum pool add|delete|modify; pvesh get /pools[{poolid,comment}], /pools/<id>{members:[{vmid,…}]}).
  • --adopt-pool retrofit mode. Non-destructive: adds an EXISTING Felhom guest to the pool (creating it if needed), resolving the guest from --vmid else the recorded provisioned_vmid. Reuses the ours-check (/etc/felhom-bootstrap mount) — refuses a non-Felhom guest unless --force. Touches ONLY pool membership: never reconfigures/restarts the guest, never contacts the hub. Idempotent (skip-if-member).
  • --uninstall pool teardown (step 5b). After the pveum removal, deletes the felhom pool only if empty (a destroyed guest is auto-removed from its pool); a pool that still has members is left with a log_skip naming them. Not reached on the Spec-1 safe-skip path (other Felhom guests remain).
  • Validated on felhom-pve (dry-run + SAFE live): T-A fresh-install dry-run shows the pool create + membership lines; T-B live adopt of guest 9201pvesh get /pools/felhom lists 9201, guest still running, config unchanged (the demo node is now pool-uniform); re-run = no-op; T-B' non-Felhom vmid → refusal; T-C uninstall dry-run → "pool felhom not empty (members: 9201) — leaving it". bash -n + shellcheck clean (0 new warnings; the 2 pre-existing SC2015 in step_verify unchanged).
  • NOT changed: PVE_PRIVS, the ACL grants, the agent, the provision-call args. 3b (pool-scoped ACL
    • agent restore-into-pool under a scoped token) is the separate spike-gated task.

felhom-host-install.sh v1.4.0 — appliance CPU/RAM cap passthrough (--cores / --memory) (2026-07-01)

Colleague-safety batch #3 (host-install half; the mechanism is agent v0.52.0). Lets an operator cap the provisioned guest so a trial appliance on a SHARED production Proxmox doesn't pressure the colleague's existing guests.

  • --cores N / --memory M (MiB) — optional; passed through to the agent's --selftest=provision as -cores/-memory. 0/unset = keep the golden's baked sizes (unchanged behaviour). New vars CPU_CORES/MEM_MIB; usage() header gains an "Appliance cap (optional)" group.
  • Conditional passthroughstep_provision builds a cap_args array and appends the flags to BOTH the dry-run log and the real agent call only when set. An agent < v0.52.0 would reject an unknown flag, so the flags are never sent unless the operator opts in (see the deploy dependency below).
  • Pre-flight sanity WARN (soft, provision only) — if --cores > host nproc or --memory > host MemTotal, log_warn "the cap won't protect other guests"; never die (the operator may know better).
  • Deploy dependency: a fresh install using --cores/--memory needs the hub artifact manifest to serve agent ≥ v0.52.0.
  • Validated dry-run on felhom-pve: --cores 2 --memory 4096 --dry-run → provision command shows -cores 2 -memory 4096; without the flags → neither present; --cores 64 --memory 65536 → both WARN lines (host 4 cores / ~15771 MiB). bash -n + shellcheck clean (0 new warnings; the 2 pre-existing SC2015 in step_verify unchanged).

felhom-host-install.sh v1.3.0 — --uninstall (clean revert) + pre-flight guards (2026-07-01)

Colleague-safety batch #1+#2. Adds a first-class, guarded --uninstall teardown so an operator can cleanly back out of a trial install, plus three provision pre-flight guards that stop common footguns. Script-only; no agent/hub/controller change.

  • --uninstall (local host teardown — no hub contact, no passphrase). Reverses an install in the install-order's reverse: guest → agent(unit/sudoers/binary/state/user) → pveum(ACL,token,user,role) → golden(opt-in) → state file. Every mutation goes through run() so --dry-run prints the full plan and executes nothing. Safety:
    • Ours-check: refuses to destroy a guest that lacks the /etc/felhom-bootstrap bind mount (matched by the constant guest path, not a hardcoded mpN slot — on the demo host it's mp9), unless --force.
    • Typed confirmation: must type the vmid to confirm PERMANENT destruction (read from /dev/tty; skipped only under --dry-run, where nothing is destroyed).
    • Other-guests guard: if any OTHER Felhom guest remains, destroys only the target and leaves the agent + PVE token + state in place (re-run with --force to remove host-level anyway — orphans the others).
    • Never removes the sudo package; never contacts the hub (the host record intentionally persists).
    • Presence-checked + idempotent: an already-absent guest/unit/sudoers/binary/user/ACL/token/role is a tolerated skip, not an error. The pveum role delete runs only after its ACL grants are gone (PVE refuses to delete a referenced role). Confirmed PVE 9 ACL-delete form: pveum acl delete / --users|--tokens <x> --roles FelhomAgent.
    • Target vmid resolves from --vmid, else the recorded provisioned_vmid (else dies). A --vmid that disagrees with the recorded one needs --force.
    • --remove-golden: with --uninstall, also delete the golden vzdump from the archive storage (pvesm free); otherwise it is left in place.
  • Install state now records customer_id + provisioned_vmid (new _state_put/_state_get helpers, dry-run-guarded like _state_mark; the completed[] shape is untouched) so a later --uninstall resolves its target automatically and safely.
  • Pre-flight guards (provision mode):
    • Multi-node guard — on a 2+-node cluster, die (naming the nodes) unless --node is explicit (new NODE_EXPLICIT); single-node keeps the current auto-pick. No-op under --skip-provision.
    • Archive-storage-exists guard — verify --archive-storage appears in pvesm status (else die); no-op under --skip-provision.
    • RAM floor (WARN, never fatal) — warn when MemAvailable < 2048 MiB. All three run inside step_preflight (before any mutation) so they also fire under --dry-run.
  • Validated dry-run-only on felhom-pve (single-node, live guest 9201): T-A full uninstall plan, T-C not-ours refusal (red-proof), archive-missing die, RAM line, other-guests detector, state round-trip; confirmed 9201 + agent + pveum + state untouched after all dry-runs. bash -n + shellcheck clean (0 new warnings vs. baseline; the 2 pre-existing SC2015 in step_verify are unchanged). NOT yet live-validated (awaiting a supervised run): a real live --uninstall (guest destroy + pveum removal) and the multi-node guard on an actual cluster.

felhom-host-install.sh v1.2.0 — /dev/tty passphrase read + vmid auto-detect (2026-07-01)

Two operator-experience fixes so a colleague can install online (via the hub's new "Option 1: Online install" one-liner) and onto a host that already runs a guest at 9201.

  • Passphrase prompt reads from /dev/tty, not stdin (read_passphrase). read -rsp … < /dev/tty makes the no-echo prompt work regardless of how stdin is wired — both download-then-run and curl … | sudo bash (where stdin is the pipe). Strictly more correct; the --passphrase-file path is unchanged. The passphrase is still never on argv / in logs / in the state file.
  • VMID auto-detect (--vmid now optional-smart). New VMID_EXPLICIT flag (set by --vmid). The pre-flight vmid guard now determines "in use" against the pct list + qm list id-set (LXC and VMs share the id space — more complete than the old pct status, which only knew LXC):
    • explicit --vmid → unchanged deterministic behavior: die if the id is in use unless --force (destructive over-provision).
    • default 9201, in use, no --forceauto-pick the next free id (scan upward from 9201 over the used-set) and ask to confirm from the terminal (read … < /dev/tty, [y/N]); proceed on yes, die "no free vmid confirmed" otherwise. Never a silent auto-pick.
    • default 9201 + --force → over-provision 9201 (destructive) without prompting, as before.
    • New helpers used_vmids / _vmid_in_use / next_free_vmid. --vmid help text + usage() updated.

felhom-host-install.sh v1.1.0 — self-install the agent + fetch the golden from Gitea (2026-06-28)

The script now installs the agent itself (the last big manual Day-0 prerequisite is gone). It fetches the agent binary + golden from Gitea generic packages and verifies each against the hub-vouched artifact manifest before installing/using it. BUNDLE slice; pairs with hub v0.16.0 (artifact manifest endpoint + operator UI) and felhom-agent v0.43.0 (canonical unit + publish).

  • New step 5/8 agent install (before agent-config): resolves the manifest (GET /api/v1/artifacts/{id}, passphrase) + the git fetch token (from the customer's controller.yaml via config-retrieve — NO new credential); fetches /api/packages/admin/generic/felhom-agent/<ver>/felhom-agent, verifies sha256 vs the hub manifest (aborts on mismatch — verify-before-use), backs up any existing binary, installs 0755 /usr/local/bin/felhom-agent; ensures the non-root felhom-agent system user; installs the canonical sudoers (0440, visudo -cf-validated) + systemd unit; daemon-reload + enable. Idempotent: same version already installed + service active → skip.
  • --skip-provision: install + configure + verify the agent (incl. golden fetch+verify) but do NOT provision a guest — the agent-only path for re-installing/upgrading the agent on a host that already has live guests. Adds an agent-only step_verify_agent (binary + non-root service active + a --selftest=hub collect-report).
  • New step 7/8 golden: local auto-discovery stays the default/fallback; otherwise fetches /api/packages/admin/generic/felhom-golden/<ver>/golden.tar.zst, verifies sha256, and imports it into the archive storage's dump dir for the restore. --force-gitea-golden forces the Gitea path.
  • Non-root agent model: the agent now runs as felhom-agent with privileged.mode: "sudo" (was the dev/CI direct+root shortcut). The config is chowned to the service user (0600) so the daemon can read it; systemctl is-active after restart is the real proof the non-root user can read the config.
  • Pre-flight relaxed: a missing agent binary is no longer fatal (step 5 installs it); the local golden requirement is deferred to step 7.
  • Trust model: checksum trust root = the hub (manifest), not Gitea; the fetch credential is the existing config-retrieve git token; artifacts are pinned to a version (never :latest).
  • Secrets: the git token is a never-logged runtime carrier (cleared on EXIT alongside the passphrase / pve-token / hub api_key); the sudoers is 0440 and visudo -cf-validated before install.
  • bash -n + shellcheck clean.

felhom-host-install.sh v1.0.0 — Day-0 host bootstrap (provision mode) (2026-06-26)

First release. A single operator-run script that automates Day-0 on a freshly-PVE-installed host: Proxmox API token → hub host enrollment (option C, single secret) → agent config → guest provision → verify. Composes proven mechanisms (the pveum role/token sequence, hub POST /host-enroll, felhom-agent --selftest=provision); grounded by documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md.

  • 7 steps, idempotent + resumable via /var/lib/felhom-install/state.json: pre-flight → Proxmox token → compute grows → host-enroll → agent config → provision → verify.
  • Single-secret (the retrieval passphrase): read no-echo or from a 0600 file, never on argv/logs/state. The global operator key never touches the box.
  • pveum automation: 16-priv FelhomAgent role (create-or-modify), felhom-agent@pve user, privsep token (reuse-if-working else rotate), and both ACL grants applied after the token exists (token-remove purges the token ACL).
  • Auto-discovery: golden archive (newest vzdump-lxc-<golden-vmid>), PVE node name, vmbr0 bridge IP for the local-api, and the served-leaf TLS fingerprint pin.
  • Safety: pre-flight fails fast (root, PVE 9.x, local-lvm headroom, hub reachable, customer+passphrase valid via read-only GET /config/{id}, golden resolvable); refuses to clobber an existing --vmid without --force; --dry-run previews every mutation; --preserve-from keeps operator infra (PBS/local_api/privileged/authz) on re-deploys.
  • --mode dr: documented 10D stub (restore customer PBS snapshot instead of golden) — not implemented.
  • Live-validated end-to-end on felhom-pve: authorized wipe of demo guest 9201 → re-provision from the golden → controller config-pull + public tunnel HTTP 200 → host-report of guest 9201 → idempotent --resume no-op. (One ordering bug — token ACL applied before rotation — was found and fixed during the live run.)