# Felhom scripts — Changelog ## docs — 06-doc S3 row SHIPPED + agent-side revocation semantics (2026-07-04) Docs-only companion to **felhom-agent v0.64.0** (the S3 slice — keygen, registration, agent-managed `wg-quick@wg-felhom`, escrow join; live-validated on felhom-pve incl. revocation drill, reboot persistence, 30-min soak). 06-doc §3.5 now records: register-once marker, revoked-stays-revoked, re-add via the registration endpoint (the raw registry add doesn't bump the host generation — live finding), `wg_tunnel.enabled` default-FALSE rollout gate. S6 backlog notes added (hub poll constant configurable + first-adoption log; registry-add bump-or-label). CGNAT/mobile-hotspot appendix deferred (operator-assisted; §7's open validation stands). ## felhom-peersync.sh v1.0.1 — strip out of process substitution (exit-swallow fix) (2026-07-04) The S1 REPORT's exit-swallow class, fixed: `wg syncconf wg0 <(wg-quick strip "$tmp")` hid the strip exit code — a corrupt `wg0.conf.head` could feed syncconf empty/partial input that WIPES the live peer set while the script exits 0 (and the bad conf is then persisted). v1.0.1 runs strip as its own step into `$tmpdir/stripped`; a strip failure aborts BEFORE `wg` is invoked. Sandbox red-proof (stub `wg-quick` exiting 1 after partial output + recording stub `wg`): pre-fix shape invoked wg and returned rc=0; fixed shape errors first, wg never called. Redeployed to the dev endpoint (runbook step 5 install); shellcheck clean. ## felhom-peersync.sh v1.0.0 — the offsite endpoint's WG reconcile script (2026-07-04) S1 (doc 06 §5): the forced-command target the hub's wgsync pushes to (runbook `offsite-endpoint.md` step 5 installs it as `/usr/local/bin/felhom-peersync`, root:root 0755, invoked via a one-line sudoers grant from the `felhom-peersync` user's `restrict,command=` authorized_keys entry). Validate-FIRST design: jq contract check (version 1, interface wg0, 44-b64 pubkeys, `10.77.0.x/32` allowed_ips, never the endpoint's own .1) rejects on stderr with exit 1 before touching anything; then head-file + generated `[Peer]` blocks into a same-fs tmp, `wg syncconf <(wg-quick strip …)` from the TMP (exact-match: adds/removes without bouncing the interface), and only on success the atomic `mv` to `/etc/wireguard/wg0.conf` — runtime and boot config can never diverge in the failure direction. Zero-peer payload = valid wipe. Never reads or prints the private key; no `wg-quick save`; no second mode. shellcheck-clean. Live-proven on felhom-hetzner incl. the negatives (malformed JSON / bad pubkey / own-IP peer → exit 1, wg state byte-identical) and reboot persistence. ## docs — architecture Part 06: offsite connectivity design-of-record (2026-07-03) `documentation/architecture/06-offsite-connectivity.md` — the settled offsite-backup-transport design, authored from the spike verdict + operator-resolved forks (recorded, not re-litigated): plain WG (D1), host-side **agent-managed** `wg-felhom` as the agent-managed-unit pilot on the sudoers `*.mount` install pattern (D2), one shared hub-driven endpoint VM running WG + the offsite PBS with no agent (D3, CF-token pattern), hub source-of-truth with a `wireguard` block riding the existing `WireDesiredState`/DesiredGeneration channel (D4), one datastore + per-customer namespaces (D5), and PBS **on** the VM — relay-through-DooPlex rejected as non-scaling (D6). Includes the Day-0 join handshake, the robustness set (NOT-DynDNS roaming, endpoint DNS re-resolve watchdog, MTU 1420, per-/32 topological isolation, tunnel-health through the storage-target reachability model), trust conformance, the honest open ledger (CGNAT unmeasured → mobile-hotspot smoke test; peer-sync push-vs-pull = slice-1 design point), and the S1–S6 slice roadmap (MVP = S1→S2→S3, then S4; S5 merges with DR-completeness). All claims cited at file:line against felhom.eu @ bf099f6 + felhom-agent @ 4ba1b14. `day0-install.md` backlog line now points at spike + design doc. Docs-only. ## docs — SPIKE: offsite-backup connectivity — plain WireGuard WINS the ladder; Headscale = separable fleet layer (2026-07-03) `documentation/audits/SPIKE-connectivity-wireguard-2026-07-03.md` — the offsite-backup transport decision, empirically grounded on both real ends (demo-felhom PVE host ⟷ throwaway Hetzner box, end state: powered off, secrets shredded). Headline results: the operator's line is **plain-NAT with a fixed public IP, not CGNAT, and has zero IPv6** (P0 honesty — CGNAT confirmation deferred to Peti's VM 110); plain outbound WG held an 11.4-min idle window and carried a **real 2 GiB worst-case PBS backup at 4.26 MiB/s = the full home uplink** (~5% tunnel overhead), TLS pin intact through the tunnel (positive + negative proof); UDP 51820 **and** 443 both pass; kernel WG surprisingly *works* inside the unprivileged guest (P7 — host placement stands on architecture, not infeasibility). Recommendation: host-side agent-managed WG (cloudflared pattern), key in the escrowed IdentityBundle, small public endpoint VM (PBS-on-VM vs rendezvous-relay deferred to the spec). `runbooks/day0-install.md` backlog line resolved to point here; `CONTEXT.md` notes the DR-completeness task is unblocked (next: the production connectivity spec). ## skills — NEW: felhom-app-catalog (4th skill) + SparkyFitness as its worked example (2026-07-03) `skills/felhom-app-catalog/SKILL.md` — the catalog **authoring workflow** (research → inspect the image for the healthcheck family → write compose/.felhom.yml → deploy live through the dashboard → verify healthy → reconcile the app count). Deliberately points at app-catalog `REUSE.md` §1–2 + `README.md` §format for every field table (one-fact-one-place; no duplication). Unique content: the never-guess-the-healthcheck rule with the per-tool image-inspection loop (BusyBox `ash` `command -v` gotcha: it silently ignores all but its first argument — verified), the probe-container naming rule (controller probes the container named exactly like the stack — verified in `felhom-controller/internal/stacks/healthprobe.go`, row added to app-catalog REUSE.md), the Hungarian-quote YAML kill, and the deploy-is-the-test doctrine. No installer change needed — `install_skills.py` auto-discovers `skills/*/SKILL.md`; fresh-session discovery probe listed all 4. Proven by finalizing `sparkyfitness` end-to-end on demo (see app-catalog-felhom.eu CHANGELOG). ## docs — golden 0.98.3 live: D.1b RETIRED, drill ledger B1/B5 → FIXED, backlog note resolved (2026-07-03) Companion to felhom-agent's `build-golden.sh` v2.0.0 (@ `ceca355`): the golden now bakes the CURRENT controller (0.98.3, mandatory-tag convention — B5) and a `felhom-controller-bootstrap.path` unit (controller deploys on the bootstrap-mount hot-plug, no reboot — B1). Clean-room validated end-to-end BEFORE publish (bake → isolated hot-plug proof → local-golden Day-0 install → publish → vouch → `--force-gitea-golden` install); evidence: `documentation/audits/DRILL-golden-098-2026-07-03.md`. - `documentation/runbooks/day0-install.md`: **D.1b reduced to a one-line version check** (fresh boxes land current + self-manage); old manual-update procedure moved to Part F troubleshooting keyed on "golden older than 0.86.0 was vouched"; header validated-versions line → script v1.9.1 / agent v0.63.0 / golden v0.98.3; A.3 records the drilled-known-good pair + "vouch ≥ 0.98.3"; A.4 floor text rewritten (fresh boxes now floor-honoring; recommend raising the global floor after a golden rebuild — operator, 1 min). - `documentation/audits/DRILL-day0-cleanroom-2026-07-03.md` ledger: **B1, B5 → FIXED** (pointers); R6 note: the installer's post-provision reboot is retained as a belt, removal recorded as a candidate cleanup (not done). - `documentation/backlog/FOLLOWUP-golden-default-controller-tag.md` + `backlog/README.md`: **RESOLVED** per the M18/M19 convention (file kept + annotated; README entry marked FIXED). - New evidence doc: `documentation/audits/DRILL-golden-098-2026-07-03.md` (A–D transcripts, unit states, resolution-order + fetch/sha proofs, cleanup, observations — incl. the SECURITY observation that the customer `git.token` has package-WRITE rights → scope-down + rotate follow-up). ## docs — day0-install D.1b narrowed + drill ledger B2/B3 → FIXED agent v0.63.0 (2026-07-03) Agent v0.63.0 fixed both drill findings (B3 token reload-on-miss + B2 guesthook snippets-dir mkdir; red-proofed, deployed on felhom-pve, Gitea-published sha256 b4a89c81…). Guide follow-through: the D.1b "restart the agent first" step is now CONDITIONAL (only for an installed agent < v0.63.0 — the Day-0 manifest still vouches 0.62.0, so today's fresh installs still hit it); the 401 troubleshooting row records the fix version; the drill ledger + go/no-go item 8 marked FIXED. Operator follow-up unchanged: vouch agent 0.63.0 in the Day-0 manifest UI, then the step is dead. ## felhom-host-install.sh v1.9.1 — clean-room drill fixes: residue-free uninstall + post-provision reboot (2026-07-03) Companion to the Day-0 go-live package (`documentation/runbooks/day0-install.md` + `documentation/audits/DRILL-day0-cleanroom-2026-07-03.md`). Every fix was found by the clean-room drill (virgin nested PVE 9.2.2) and re-verified there (v1.9.1 uninstall → **zero-Felhom-residue diff vs the pre-install baseline**; v1.9.1 install → controller up with no manual intervention). - **Header/version sync** (the header said v1.8.0 while `SCRIPT_VERSION` said 1.9.0); keep-in-sync note on `SCRIPT_VERSION`; usage sed range follows the header (2,95). - **Uninstall now removes the drill-found residue (R1–R5):** the agent **config** (resolved from the unit's `-config` BEFORE the unit is removed — it holds the per-host hub api_key), the `felhom-shared-parent` unit + wants links + `/usr/local/sbin/felhom-shared-parent.sh` + the `/mnt/felhom-drives` self-bind/dir, `/usr/local/sbin/felhom-mkfs-guarded`, `/var/lib/vz/snippets/felhom-guest-hook.sh`, and `/etc/dnsmasq.d/felhom-*.conf` (+ dnsmasq restart when touched). All tolerate-absent; summary lines updated (`sudo` AND `dnsmasq` packages are the documented package remnants). - **Post-provision guest reboot (R6):** the golden's `felhom-controller-bootstrap.service` evaluates `ConditionPathExists=/etc/felhom-bootstrap/bootstrap.json` at BOOT, but the agent back-half hot-plugs the mount into the running guest — on slower hardware the first boot loses that race deterministically and the controller never deploys. `step_provision` now reboots the guest once (the agent's own output says "next: reboot the guest"); `step_verify` waits bounded (180 s) for the controller container instead of a momentary look. ## felhom-host-install.sh v1.9.0 — Pool.Audit for the stale-lock reaper (A1) (2026-07-03) Companion to felhom-agent v0.62.0 (audit A1: pool-membership ownership check). `PVE_PRIVS_GUEST` gains **`Pool.Audit`** (12 → 13 privs, granted at `/pool/felhom` via the existing FelhomAgentGuest role) so the agent can read `GET /pools/felhom` — its stale-lock reaper's ownership registry. `Pool.Allocate` does NOT satisfy the read (spike SPIKE-a1-pool-membership-read-2026-07-03 T2). No structural change: `_ensure_role` already `role modify`s to the exact priv set, so re-running `--rescope-acl` (or a fresh install) upgrades an existing box idempotently; `remove_scoped_acl` deletes by role name and needs nothing. **Deploy order on a live box: rescope FIRST, then deploy agent v0.62.0** — the added read priv is harmless to an older agent, while the new agent on an old ACL fail-safes its reaper (skips) and reports `pve:pool-read` degraded until the rescope lands. ## docs — SPIKE: A1 pool-membership read for the stale-lock reaper (2026-07-03) Findings doc `documentation/audits/SPIKE-a1-pool-membership-read-2026-07-03.md`. Live-probed on felhom-pve under the PRODUCTION scoped token vs root: LXC enumeration IS already pool-filtered (token sees only 9201 of 4 guests); `GET /pools/felhom` 403s naming `Pool.Audit`; a throwaway token with ONLY `Pool.Audit`@`/pool/felhom` reads members (minimal delta proven, fully torn down); `/cluster/resources` withholds the `pool` field without `Pool.Audit`; local ownership records are all partial. Recommendation for the A1 impl spec: add `Pool.Audit` to `PVE_PRIVS_GUEST` in `felhom-host-install.sh` (L183) + a `GET /pools/felhom` cross-check in the agent's `staleLockController.Guests()`, fail-safe skip on read failure. No script/agent change in this commit — docs only. Appendix: committed-secrets (felhom.secret.yaml) rotation micro-runbook, operator follow-up. ## install_skills.py — new: Claude Code skills installer (2026-07-03) Installs `skills/*/SKILL.md` (felhom-build-deploy, felhom-ui-design, felhom-testing) into `~/.claude/skills/` as Windows junctions (`mklink /J`) so repo edits are live immediately; falls back to a full copy if junction creation fails or isn't followed (copy mode prints a re-run reminder). Idempotent — re-runs detect a correct junction and leave it. Verified: junctions ARE followed by Claude Code skill discovery (fresh-session probe found all three). ## reuse_refs_check.py — new gate: REUSE.md citation checker (2026-07-03) Staleness defense for the new per-repo `REUSE.md` reuse maps. Takes repo roots as argv, extracts every cited `*.go/*.py/*.html/*.css/*.yml/*.yaml/*.sh` path (slash-containing tokens only — bare filenames are conventions, not citations), verifies each exists; prints offenders, non-zero exit on any missing path. Symbols are spot-verified by the reviewer, not this script. Usage: `python scripts/reuse_refs_check.py [...]`. ## felhom-host-install.sh v1.8.0 — install the guarded-mkfs wrapper (Impl-1 Part B) (2026-07-01) Companion to felhom-agent v0.54.0 (format-safety foundation). During agent install, fetch + install the guarded-mkfs wrapper so the agent's format path is safe on any box. - **New step in `step_agent_install`:** fetch `configs/felhom-mkfs-guarded.sh` from Gitea, `bash -n` validate, `install -m0755 -o root -g root` → `/usr/local/sbin/felhom-mkfs-guarded`. Installed BEFORE the sudoers (which now allowlists ONLY the wrapper, not raw `mkfs.*`), so the ordering is gap-free. - The agent v0.54.0 sudoers (fetched by the same step) drops the raw `mkfs.ext4 -F /dev/* / mkfs.xfs -f /dev/*` allowlist and permits only `felhom-mkfs-guarded /dev/* *`, plus read-only `pvs`/`zpool` for the agent's unclaimed-disk guard. No other host-install change. - `bash -n` + `shellcheck` clean (0 new warnings). Live-validated on felhom-pve (agent v0.54.0 deploy): wrapper refuses the OS disk + an LVM-PV partition, raw mkfs is sudo-denied, an unclaimed throwaway disk formats; the agent guard's sudo reads (pvs/lsblk/zpool) all work as the felhom-agent user. ## felhom-host-install.sh v1.7.0 — 3b-fix: `Datastore.Audit` box-wide (restore drive visibility) (2026-07-01) Fixes a regression the v1.6.0 pool-scoped ACL introduced: `Datastore.Audit` was placed in the per-storage `Store` role (granted only on `local`/`local-lvm`/`felhom-pbs`), which **excluded the enrolled removable drives** `felhom-usb`/`felhom-flash`. The agent enumerates storage via `ListStorage`/`NodeStorage` (both gated by `Datastore.Audit` — `internal/storage/observe.go`), so it could no longer SEE the drives → false "Meghajtó leválasztva" (drive detached) alerts + drives absent from the agent-view. (The v1.6.0 swap's "felhom-usb → 403" was mis-read as blast-radius success; felhom-usb is Felhom's OWN customer drive, not an out-of-scope object.) - **`Datastore.Audit` moved from Store → Base** (`PVE_PRIVS_BASE` now `"Sys.Audit SDN.Use Datastore.Audit"`; `PVE_PRIVS_STORE` now `"Datastore.Allocate Datastore.AllocateSpace"`). Audit is read-only metadata, so box-wide Audit restores visibility of ALL storages (incl. dynamically-enrolled drives — no per-drive grant ever needed) while the **write** privs (`Allocate`/`AllocateSpace`) stay per-storage → write/allocate blast-radius containment is UNCHANGED. Confirmed at source: the agent creates no PVE storage (no `POST /storage`/`pvesm add`); drives are dir-storages it observes + mounts via host ops, so they need only Audit, never Allocate. - **`apply_scoped_acl` reordered** Base-before-Store (role + grant) so a RE-APPLY on a live box adds `Audit@/` before Store drops its per-storage Audit → gap-free (the agent never loses enumeration). - `remove_scoped_acl` / `--uninstall` / `--rescope-acl` operate by role NAME and inherit the corrected privs automatically (no other change). - **Live-repaired felhom-pve** (two `pveum role modify`, Base first — no agent stop/restart): drives reappeared (agent-view 3→5 storages), detach alerts cleared. Re-tested under the scoped token: drives readable (was 403), write-containment intact (vzdump→felhom-usb still 403; out-of-pool guest 403), PBS Store grant unchanged. `bash -n` + `shellcheck` clean (0 new warnings). - **NOT physically run** (source-confirmed, no `Datastore.Allocate` in the path): a brand-new-drive UI enrollment (needs a spare USB) — the host-ops/Audit path is unchanged from pre-3b. ## felhom-host-install.sh v1.6.0 — pool-scoped token ACL (3-role) + `--rescope-acl` retrofit (2026-07-01) Colleague-safety batch #4 phase b (script half; agent half = v0.53.0). Moves the agent token's dangerous privileges off `/` (which spanned every guest + storage) to `/pool/felhom` + `/storage/`, so on a shared box the token can only touch Felhom's own guests + storages. Validated by `documentation/audits/SPIKE-pool-scoped-acl-2026-07-01.md` (PASS) — implemented here. - **3-role scoped ACL (`step_token` rewrite).** Replaces the single `FelhomAgent` role granted at `/` with three roles, each granted to BOTH the user AND the token (privsep intersection): `FelhomAgentGuest` (`VM.*` + `Pool.Allocate`) @ `/pool/felhom`; `FelhomAgentStore` (`Datastore.*`) @ each of `PVE_STORAGES` (default `local local-lvm felhom-pbs` — the offsite PBS MUST be included, SPIKE residual #1; `--acl-storages` overrides); `FelhomAgentBase` (`Sys.Audit SDN.Use`) @ `/`. Helpers `apply_scoped_acl`/`remove_scoped_acl`/`_grant`/`_ensure_role`. - **Pool before token.** `ensure_felhom_pool` runs at the top of `step_token` (always, incl. `--skip-provision`) so `/pool/felhom` exists before it's granted on. - **Re-install safety.** `step_token` also removes the pre-3b broad `/` grant + `FelhomAgent` role if present (`remove_old_broad_acl`, tolerate-absent), so a re-install can't leave the old grant unioned with the scoped one. The post-provision `pool_add_guest` is gone (the agent's `restore --pool` makes the guest a member atomically — v0.53.0). - **`--rescope-acl` retrofit** (new mode, mirrors `--adopt-pool`): migrate an existing install — ensure the pool + guest membership, apply the scoped grants, THEN remove the old broad grant (add-before- remove: the token is never grant-less mid-migration). Prints the "now deploy agent ≥ v0.53.0" ordering reminder. Idempotent + dry-run-aware. **SUPERVISED** (run with the agent stopped — the scoped ACL and the pool-param agent are mutually dependent; §13 of the task). - **`--uninstall`** now removes the scoped grants + 3 roles AND the pre-3b broad grant/role (both tolerate-absent → works on either shape), keeping the pool delete-if-empty (v1.5.0). - **Validated on felhom-pve** (dry-run): T-A fresh install (pool-before-token, 3 roles once, scoped grants incl. `/storage/felhom-pbs`), `--rescope-acl` (add scoped → remove old `FelhomAgent`), T-F uninstall (old-shape cleanup + pool not-empty skip). `bash -n` + `shellcheck` clean (0 new warnings). **The live rescope + agent swap is the supervised STOP** — not run here. ## felhom-host-install.sh v1.5.0 — `felhom` pool by default + `--adopt-pool` retrofit + uninstall teardown (2026-07-01) Colleague-safety batch #4 phase a. Every Felhom-managed guest now joins a dedicated **`felhom` pool** for fleet uniformity (and as the environment the later pool-scoped ACL — 3b — will spike against). All pool ops run as `root@pam` from the installer, so there is **NO agent/token/ACL change** and zero permission-model risk (`PVE_PRIVS` untouched; the `FelhomAgent` token stays scoped at `/`). - **New `felhom` pool default.** `step_provision` calls `ensure_felhom_pool` (create if absent, idempotent) and, after a successful provision, adds the guest via `pveum pool modify felhom -vms ` (skip-if-already-member). New helpers `pool_exists` / `pool_members` / `ensure_felhom_pool` / `pool_add_guest`; const `PVE_POOL="felhom"`. PVE 9 syntax + `/pools` JSON shape confirmed live before wiring (`pveum pool add|delete|modify`; `pvesh get /pools` → `[{poolid,comment}]`, `/pools/` → `{members:[{vmid,…}]}`). - **`--adopt-pool` retrofit mode.** Non-destructive: adds an EXISTING Felhom guest to the pool (creating it if needed), resolving the guest from `--vmid` else the recorded `provisioned_vmid`. Reuses the ours-check (`/etc/felhom-bootstrap` mount) — refuses a non-Felhom guest unless `--force`. Touches ONLY pool membership: never reconfigures/restarts the guest, never contacts the hub. Idempotent (skip-if-member). - **`--uninstall` pool teardown (step 5b).** After the pveum removal, deletes the `felhom` pool **only if empty** (a destroyed guest is auto-removed from its pool); a pool that still has members is left with a `log_skip` naming them. Not reached on the Spec-1 safe-skip path (other Felhom guests remain). - **Validated on felhom-pve** (dry-run + SAFE live): T-A fresh-install dry-run shows the pool create + membership lines; T-B **live adopt of guest 9201** → `pvesh get /pools/felhom` lists 9201, guest still running, config unchanged (the demo node is now pool-uniform); re-run = no-op; T-B' non-Felhom vmid → refusal; T-C uninstall dry-run → "pool felhom not empty (members: 9201) — leaving it". `bash -n` + `shellcheck` clean (0 new warnings; the 2 pre-existing SC2015 in `step_verify` unchanged). - **NOT changed:** `PVE_PRIVS`, the ACL grants, the agent, the provision-call args. 3b (pool-scoped ACL + agent restore-into-pool under a scoped token) is the separate spike-gated task. ## felhom-host-install.sh v1.4.0 — appliance CPU/RAM cap passthrough (`--cores` / `--memory`) (2026-07-01) Colleague-safety batch #3 (host-install half; the mechanism is agent v0.52.0). Lets an operator cap the provisioned guest so a trial appliance on a SHARED production Proxmox doesn't pressure the colleague's existing guests. - **`--cores N` / `--memory M` (MiB)** — optional; passed through to the agent's `--selftest=provision` as `-cores`/`-memory`. `0`/unset = keep the golden's baked sizes (unchanged behaviour). New vars `CPU_CORES`/`MEM_MIB`; `usage()` header gains an "Appliance cap (optional)" group. - **Conditional passthrough** — `step_provision` builds a `cap_args` array and appends the flags to BOTH the dry-run log and the real agent call **only when set**. An agent < v0.52.0 would reject an unknown flag, so the flags are never sent unless the operator opts in (see the deploy dependency below). - **Pre-flight sanity WARN (soft, provision only)** — if `--cores` > host `nproc` or `--memory` > host `MemTotal`, `log_warn` "the cap won't protect other guests"; never `die` (the operator may know better). - **Deploy dependency:** a fresh install using `--cores`/`--memory` needs the hub artifact manifest to serve **agent ≥ v0.52.0**. - **Validated dry-run on felhom-pve:** `--cores 2 --memory 4096 --dry-run` → provision command shows `-cores 2 -memory 4096`; without the flags → neither present; `--cores 64 --memory 65536` → both WARN lines (host 4 cores / ~15771 MiB). `bash -n` + `shellcheck` clean (0 new warnings; the 2 pre-existing SC2015 in `step_verify` unchanged). ## felhom-host-install.sh v1.3.0 — `--uninstall` (clean revert) + pre-flight guards (2026-07-01) Colleague-safety batch #1+#2. Adds a first-class, guarded **`--uninstall`** teardown so an operator can cleanly back out of a trial install, plus three provision pre-flight guards that stop common footguns. Script-only; no agent/hub/controller change. - **`--uninstall` (local host teardown — no hub contact, no passphrase).** Reverses an install in the install-order's reverse: **guest → agent(unit/sudoers/binary/state/user) → pveum(ACL,token,user,role) → golden(opt-in) → state file.** Every mutation goes through `run()` so `--dry-run` prints the full plan and executes nothing. Safety: - **Ours-check:** refuses to destroy a guest that lacks the `/etc/felhom-bootstrap` bind mount (matched by the constant guest *path*, not a hardcoded `mpN` slot — on the demo host it's `mp9`), unless `--force`. - **Typed confirmation:** must type the vmid to confirm PERMANENT destruction (read from `/dev/tty`; skipped only under `--dry-run`, where nothing is destroyed). - **Other-guests guard:** if any OTHER Felhom guest remains, destroys only the target and **leaves the agent + PVE token + state in place** (re-run with `--force` to remove host-level anyway — orphans the others). - **Never removes the `sudo` package**; never contacts the hub (the host record intentionally persists). - Presence-checked + idempotent: an already-absent guest/unit/sudoers/binary/user/ACL/token/role is a tolerated skip, not an error. The `pveum role delete` runs only after its ACL grants are gone (PVE refuses to delete a referenced role). Confirmed PVE 9 ACL-delete form: `pveum acl delete / --users|--tokens --roles FelhomAgent`. - Target vmid resolves from `--vmid`, else the recorded `provisioned_vmid` (else dies). A `--vmid` that disagrees with the recorded one needs `--force`. - **`--remove-golden`:** with `--uninstall`, also delete the golden vzdump from the archive storage (`pvesm free`); otherwise it is left in place. - **Install state now records `customer_id` + `provisioned_vmid`** (new `_state_put`/`_state_get` helpers, dry-run-guarded like `_state_mark`; the `completed[]` shape is untouched) so a later `--uninstall` resolves its target automatically and safely. - **Pre-flight guards (provision mode):** - **Multi-node guard** — on a 2+-node cluster, `die` (naming the nodes) unless `--node` is explicit (new `NODE_EXPLICIT`); single-node keeps the current auto-pick. No-op under `--skip-provision`. - **Archive-storage-exists guard** — verify `--archive-storage` appears in `pvesm status` (else `die`); no-op under `--skip-provision`. - **RAM floor (WARN, never fatal)** — warn when `MemAvailable < 2048 MiB`. All three run inside `step_preflight` (before any mutation) so they also fire under `--dry-run`. - **Validated dry-run-only on felhom-pve** (single-node, live guest 9201): T-A full uninstall plan, T-C not-ours refusal (red-proof), archive-missing `die`, RAM line, other-guests detector, state round-trip; confirmed 9201 + agent + pveum + state untouched after all dry-runs. `bash -n` + `shellcheck` clean (0 new warnings vs. baseline; the 2 pre-existing SC2015 in `step_verify` are unchanged). **NOT yet live-validated (awaiting a supervised run):** a real live `--uninstall` (guest destroy + pveum removal) and the multi-node guard on an actual cluster. ## felhom-host-install.sh v1.2.0 — /dev/tty passphrase read + vmid auto-detect (2026-07-01) Two operator-experience fixes so a colleague can install online (via the hub's new "Option 1: Online install" one-liner) and onto a host that already runs a guest at 9201. - **Passphrase prompt reads from `/dev/tty`, not stdin** (`read_passphrase`). `read -rsp … < /dev/tty` makes the no-echo prompt work regardless of how stdin is wired — both download-then-run **and** `curl … | sudo bash` (where stdin is the pipe). Strictly more correct; the `--passphrase-file` path is unchanged. The passphrase is still never on argv / in logs / in the state file. - **VMID auto-detect (`--vmid` now optional-smart).** New `VMID_EXPLICIT` flag (set by `--vmid`). The pre-flight vmid guard now determines "in use" against the **`pct list` + `qm list`** id-set (LXC and VMs share the id space — more complete than the old `pct status`, which only knew LXC): - **explicit `--vmid`** → unchanged deterministic behavior: die if the id is in use unless `--force` (destructive over-provision). - **default 9201, in use, no `--force`** → **auto-pick the next free id** (scan upward from 9201 over the used-set) and **ask to confirm** from the terminal (`read … < /dev/tty`, `[y/N]`); proceed on yes, `die "no free vmid confirmed"` otherwise. Never a silent auto-pick. - **default 9201 + `--force`** → over-provision 9201 (destructive) without prompting, as before. - New helpers `used_vmids` / `_vmid_in_use` / `next_free_vmid`. `--vmid` help text + `usage()` updated. ## felhom-host-install.sh v1.1.0 — self-install the agent + fetch the golden from Gitea (2026-06-28) The script now **installs the agent itself** (the last big manual Day-0 prerequisite is gone). It fetches the agent binary + golden from Gitea generic packages and **verifies each against the hub-vouched artifact manifest** before installing/using it. BUNDLE slice; pairs with hub v0.16.0 (artifact manifest endpoint + operator UI) and felhom-agent v0.43.0 (canonical unit + publish). - **New step `5/8 agent install`** (before agent-config): resolves the manifest (`GET /api/v1/artifacts/{id}`, passphrase) + the git fetch token (from the customer's `controller.yaml` via config-retrieve — **NO new credential**); fetches `/api/packages/admin/generic/felhom-agent//felhom-agent`, **verifies sha256 vs the hub manifest** (aborts on mismatch — verify-before-use), backs up any existing binary, installs `0755 /usr/local/bin/felhom-agent`; ensures the non-root `felhom-agent` system user; installs the canonical sudoers (`0440`, `visudo -cf`-validated) + systemd unit; `daemon-reload` + enable. Idempotent: same version already installed + service active → skip. - **`--skip-provision`:** install + configure + verify the agent (incl. golden fetch+verify) but do NOT provision a guest — the agent-only path for re-installing/upgrading the agent on a host that already has live guests. Adds an agent-only `step_verify_agent` (binary + non-root service active + a `--selftest=hub` collect-report). - **New step `7/8 golden`:** local auto-discovery stays the default/fallback; otherwise fetches `/api/packages/admin/generic/felhom-golden//golden.tar.zst`, **verifies sha256**, and imports it into the archive storage's dump dir for the restore. `--force-gitea-golden` forces the Gitea path. - **Non-root agent model:** the agent now runs as `felhom-agent` with `privileged.mode: "sudo"` (was the dev/CI `direct`+root shortcut). The config is `chown`ed to the service user (0600) so the daemon can read it; `systemctl is-active` after restart is the real proof the non-root user can read the config. - **Pre-flight relaxed:** a missing agent binary is no longer fatal (step 5 installs it); the local golden requirement is deferred to step 7. - **Trust model:** checksum **trust root = the hub** (manifest), not Gitea; the fetch credential is the existing config-retrieve git token; artifacts are pinned to a version (never `:latest`). - **Secrets:** the git token is a never-logged runtime carrier (cleared on EXIT alongside the passphrase / pve-token / hub api_key); the sudoers is `0440` and `visudo -cf`-validated before install. - `bash -n` + `shellcheck` clean. ## felhom-host-install.sh v1.0.0 — Day-0 host bootstrap (provision mode) (2026-06-26) First release. A single operator-run script that automates Day-0 on a freshly-PVE-installed host: Proxmox API token → hub host enrollment (option C, single secret) → agent config → guest provision → verify. Composes proven mechanisms (the `pveum` role/token sequence, hub `POST /host-enroll`, `felhom-agent --selftest=provision`); grounded by `documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md`. - **7 steps, idempotent + resumable** via `/var/lib/felhom-install/state.json`: pre-flight → Proxmox token → compute grows → host-enroll → agent config → provision → verify. - **Single-secret** (the retrieval passphrase): read no-echo or from a 0600 file, never on argv/logs/state. The global operator key never touches the box. - **pveum automation:** 16-priv `FelhomAgent` role (create-or-modify), `felhom-agent@pve` user, privsep token (reuse-if-working else rotate), and **both** ACL grants applied **after** the token exists (token-remove purges the token ACL). - **Auto-discovery:** golden archive (newest `vzdump-lxc-`), PVE node name, vmbr0 bridge IP for the local-api, and the served-leaf TLS fingerprint pin. - **Safety:** pre-flight fails fast (root, PVE 9.x, local-lvm headroom, hub reachable, customer+passphrase valid via read-only `GET /config/{id}`, golden resolvable); refuses to clobber an existing `--vmid` without `--force`; `--dry-run` previews every mutation; `--preserve-from` keeps operator infra (PBS/local_api/privileged/authz) on re-deploys. - **`--mode dr`:** documented 10D stub (restore customer PBS snapshot instead of golden) — not implemented. - **Live-validated** end-to-end on `felhom-pve`: authorized wipe of demo guest 9201 → re-provision from the golden → controller config-pull + public tunnel `HTTP 200` → host-report of guest 9201 → idempotent `--resume` no-op. (One ordering bug — token ACL applied before rotation — was found and fixed during the live run.)