Files
felhom.eu/scripts/CHANGELOG.md
T

409 lines
32 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Felhom scripts — Changelog
## felhom-peersync.sh v1.0.0 — the offsite endpoint's WG reconcile script (2026-07-04)
S1 (doc 06 §5): the forced-command target the hub's wgsync pushes to (runbook
`offsite-endpoint.md` step 5 installs it as `/usr/local/bin/felhom-peersync`, root:root 0755,
invoked via a one-line sudoers grant from the `felhom-peersync` user's `restrict,command=`
authorized_keys entry). Validate-FIRST design: jq contract check (version 1, interface wg0,
44-b64 pubkeys, `10.77.0.x/32` allowed_ips, never the endpoint's own .1) rejects on stderr with
exit 1 before touching anything; then head-file + generated `[Peer]` blocks into a same-fs tmp,
`wg syncconf <(wg-quick strip …)` from the TMP (exact-match: adds/removes without bouncing the
interface), and only on success the atomic `mv` to `/etc/wireguard/wg0.conf` — runtime and boot
config can never diverge in the failure direction. Zero-peer payload = valid wipe. Never reads
or prints the private key; no `wg-quick save`; no second mode. shellcheck-clean. Live-proven on
felhom-hetzner incl. the negatives (malformed JSON / bad pubkey / own-IP peer → exit 1, wg state
byte-identical) and reboot persistence.
## docs — architecture Part 06: offsite connectivity design-of-record (2026-07-03)
`documentation/architecture/06-offsite-connectivity.md` — the settled offsite-backup-transport
design, authored from the spike verdict + operator-resolved forks (recorded, not re-litigated):
plain WG (D1), host-side **agent-managed** `wg-felhom` as the agent-managed-unit pilot on the
sudoers `*.mount` install pattern (D2), one shared hub-driven endpoint VM running WG + the
offsite PBS with no agent (D3, CF-token pattern), hub source-of-truth with a `wireguard` block
riding the existing `WireDesiredState`/DesiredGeneration channel (D4), one datastore +
per-customer namespaces (D5), and PBS **on** the VM — relay-through-DooPlex rejected as
non-scaling (D6). Includes the Day-0 join handshake, the robustness set (NOT-DynDNS roaming,
endpoint DNS re-resolve watchdog, MTU 1420, per-/32 topological isolation, tunnel-health through
the storage-target reachability model), trust conformance, the honest open ledger (CGNAT
unmeasured → mobile-hotspot smoke test; peer-sync push-vs-pull = slice-1 design point), and the
S1S6 slice roadmap (MVP = S1→S2→S3, then S4; S5 merges with DR-completeness). All claims cited
at file:line against felhom.eu @ bf099f6 + felhom-agent @ 4ba1b14. `day0-install.md` backlog line
now points at spike + design doc. Docs-only.
## docs — SPIKE: offsite-backup connectivity — plain WireGuard WINS the ladder; Headscale = separable fleet layer (2026-07-03)
`documentation/audits/SPIKE-connectivity-wireguard-2026-07-03.md` — the offsite-backup transport
decision, empirically grounded on both real ends (demo-felhom PVE host ⟷ throwaway Hetzner box,
end state: powered off, secrets shredded). Headline results: the operator's line is **plain-NAT
with a fixed public IP, not CGNAT, and has zero IPv6** (P0 honesty — CGNAT confirmation deferred
to Peti's VM 110); plain outbound WG held an 11.4-min idle window and carried a **real 2 GiB
worst-case PBS backup at 4.26 MiB/s = the full home uplink** (~5% tunnel overhead), TLS pin
intact through the tunnel (positive + negative proof); UDP 51820 **and** 443 both pass; kernel WG
surprisingly *works* inside the unprivileged guest (P7 — host placement stands on architecture,
not infeasibility). Recommendation: host-side agent-managed WG (cloudflared pattern), key in the
escrowed IdentityBundle, small public endpoint VM (PBS-on-VM vs rendezvous-relay deferred to the
spec). `runbooks/day0-install.md` backlog line resolved to point here; `CONTEXT.md` notes the
DR-completeness task is unblocked (next: the production connectivity spec).
## skills — NEW: felhom-app-catalog (4th skill) + SparkyFitness as its worked example (2026-07-03)
`skills/felhom-app-catalog/SKILL.md` — the catalog **authoring workflow** (research → inspect the
image for the healthcheck family → write compose/.felhom.yml → deploy live through the dashboard →
verify healthy → reconcile the app count). Deliberately points at app-catalog `REUSE.md` §12 +
`README.md` §format for every field table (one-fact-one-place; no duplication). Unique content:
the never-guess-the-healthcheck rule with the per-tool image-inspection loop (BusyBox `ash`
`command -v` gotcha: it silently ignores all but its first argument — verified), the
probe-container naming rule (controller probes the container named exactly like the stack —
verified in `felhom-controller/internal/stacks/healthprobe.go`, row added to app-catalog REUSE.md),
the Hungarian-quote YAML kill, and the deploy-is-the-test doctrine. No installer change needed —
`install_skills.py` auto-discovers `skills/*/SKILL.md`; fresh-session discovery probe listed all 4.
Proven by finalizing `sparkyfitness` end-to-end on demo (see app-catalog-felhom.eu CHANGELOG).
## docs — golden 0.98.3 live: D.1b RETIRED, drill ledger B1/B5 → FIXED, backlog note resolved (2026-07-03)
Companion to felhom-agent's `build-golden.sh` v2.0.0 (@ `ceca355`): the golden now bakes the CURRENT
controller (0.98.3, mandatory-tag convention — B5) and a `felhom-controller-bootstrap.path` unit
(controller deploys on the bootstrap-mount hot-plug, no reboot — B1). Clean-room validated end-to-end
BEFORE publish (bake → isolated hot-plug proof → local-golden Day-0 install → publish → vouch →
`--force-gitea-golden` install); evidence: `documentation/audits/DRILL-golden-098-2026-07-03.md`.
- `documentation/runbooks/day0-install.md`: **D.1b reduced to a one-line version check** (fresh boxes
land current + self-manage); old manual-update procedure moved to Part F troubleshooting keyed on
"golden older than 0.86.0 was vouched"; header validated-versions line → script v1.9.1 / agent
v0.63.0 / golden v0.98.3; A.3 records the drilled-known-good pair + "vouch ≥ 0.98.3"; A.4 floor
text rewritten (fresh boxes now floor-honoring; recommend raising the global floor after a golden
rebuild — operator, 1 min).
- `documentation/audits/DRILL-day0-cleanroom-2026-07-03.md` ledger: **B1, B5 → FIXED** (pointers);
R6 note: the installer's post-provision reboot is retained as a belt, removal recorded as a
candidate cleanup (not done).
- `documentation/backlog/FOLLOWUP-golden-default-controller-tag.md` + `backlog/README.md`:
**RESOLVED** per the M18/M19 convention (file kept + annotated; README entry marked FIXED).
- New evidence doc: `documentation/audits/DRILL-golden-098-2026-07-03.md` (AD transcripts, unit
states, resolution-order + fetch/sha proofs, cleanup, observations — incl. the SECURITY
observation that the customer `git.token` has package-WRITE rights → scope-down + rotate
follow-up).
## docs — day0-install D.1b narrowed + drill ledger B2/B3 → FIXED agent v0.63.0 (2026-07-03)
Agent v0.63.0 fixed both drill findings (B3 token reload-on-miss + B2 guesthook snippets-dir
mkdir; red-proofed, deployed on felhom-pve, Gitea-published sha256 b4a89c81…). Guide follow-through:
the D.1b "restart the agent first" step is now CONDITIONAL (only for an installed agent < v0.63.0 —
the Day-0 manifest still vouches 0.62.0, so today's fresh installs still hit it); the 401
troubleshooting row records the fix version; the drill ledger + go/no-go item 8 marked FIXED.
Operator follow-up unchanged: vouch agent 0.63.0 in the Day-0 manifest UI, then the step is dead.
## felhom-host-install.sh v1.9.1 — clean-room drill fixes: residue-free uninstall + post-provision reboot (2026-07-03)
Companion to the Day-0 go-live package (`documentation/runbooks/day0-install.md` +
`documentation/audits/DRILL-day0-cleanroom-2026-07-03.md`). Every fix was found by the clean-room
drill (virgin nested PVE 9.2.2) and re-verified there (v1.9.1 uninstall → **zero-Felhom-residue
diff vs the pre-install baseline**; v1.9.1 install → controller up with no manual intervention).
- **Header/version sync** (the header said v1.8.0 while `SCRIPT_VERSION` said 1.9.0); keep-in-sync
note on `SCRIPT_VERSION`; usage sed range follows the header (2,95).
- **Uninstall now removes the drill-found residue (R1R5):** the agent **config**
(resolved from the unit's `-config` BEFORE the unit is removed — it holds the per-host hub
api_key), the `felhom-shared-parent` unit + wants links + `/usr/local/sbin/felhom-shared-parent.sh`
+ the `/mnt/felhom-drives` self-bind/dir, `/usr/local/sbin/felhom-mkfs-guarded`,
`/var/lib/vz/snippets/felhom-guest-hook.sh`, and `/etc/dnsmasq.d/felhom-*.conf`
(+ dnsmasq restart when touched). All tolerate-absent; summary lines updated (`sudo` AND
`dnsmasq` packages are the documented package remnants).
- **Post-provision guest reboot (R6):** the golden's `felhom-controller-bootstrap.service`
evaluates `ConditionPathExists=/etc/felhom-bootstrap/bootstrap.json` at BOOT, but the agent
back-half hot-plugs the mount into the running guest — on slower hardware the first boot loses
that race deterministically and the controller never deploys. `step_provision` now reboots the
guest once (the agent's own output says "next: reboot the guest"); `step_verify` waits bounded
(180 s) for the controller container instead of a momentary look.
## felhom-host-install.sh v1.9.0 — Pool.Audit for the stale-lock reaper (A1) (2026-07-03)
Companion to felhom-agent v0.62.0 (audit A1: pool-membership ownership check). `PVE_PRIVS_GUEST`
gains **`Pool.Audit`** (12 → 13 privs, granted at `/pool/felhom` via the existing FelhomAgentGuest
role) so the agent can read `GET /pools/felhom` — its stale-lock reaper's ownership registry.
`Pool.Allocate` does NOT satisfy the read (spike SPIKE-a1-pool-membership-read-2026-07-03 T2).
No structural change: `_ensure_role` already `role modify`s to the exact priv set, so re-running
`--rescope-acl` (or a fresh install) upgrades an existing box idempotently; `remove_scoped_acl`
deletes by role name and needs nothing. **Deploy order on a live box: rescope FIRST, then deploy
agent v0.62.0** — the added read priv is harmless to an older agent, while the new agent on an old
ACL fail-safes its reaper (skips) and reports `pve:pool-read` degraded until the rescope lands.
## docs — SPIKE: A1 pool-membership read for the stale-lock reaper (2026-07-03)
Findings doc `documentation/audits/SPIKE-a1-pool-membership-read-2026-07-03.md`. Live-probed on
felhom-pve under the PRODUCTION scoped token vs root: LXC enumeration IS already pool-filtered
(token sees only 9201 of 4 guests); `GET /pools/felhom` 403s naming `Pool.Audit`; a throwaway
token with ONLY `Pool.Audit`@`/pool/felhom` reads members (minimal delta proven, fully torn down);
`/cluster/resources` withholds the `pool` field without `Pool.Audit`; local ownership records are
all partial. Recommendation for the A1 impl spec: add `Pool.Audit` to `PVE_PRIVS_GUEST` in
`felhom-host-install.sh` (L183) + a `GET /pools/felhom` cross-check in the agent's
`staleLockController.Guests()`, fail-safe skip on read failure. No script/agent change in this
commit — docs only. Appendix: committed-secrets (felhom.secret.yaml) rotation micro-runbook,
operator follow-up.
## install_skills.py — new: Claude Code skills installer (2026-07-03)
Installs `skills/*/SKILL.md` (felhom-build-deploy, felhom-ui-design, felhom-testing) into
`~/.claude/skills/` as Windows junctions (`mklink /J`) so repo edits are live immediately; falls
back to a full copy if junction creation fails or isn't followed (copy mode prints a re-run
reminder). Idempotent — re-runs detect a correct junction and leave it. Verified: junctions ARE
followed by Claude Code skill discovery (fresh-session probe found all three).
## reuse_refs_check.py — new gate: REUSE.md citation checker (2026-07-03)
Staleness defense for the new per-repo `REUSE.md` reuse maps. Takes repo roots as argv, extracts
every cited `*.go/*.py/*.html/*.css/*.yml/*.yaml/*.sh` path (slash-containing tokens only — bare
filenames are conventions, not citations), verifies each exists; prints offenders, non-zero exit on
any missing path. Symbols are spot-verified by the reviewer, not this script.
Usage: `python scripts/reuse_refs_check.py <repo-root> [...]`.
## felhom-host-install.sh v1.8.0 — install the guarded-mkfs wrapper (Impl-1 Part B) (2026-07-01)
Companion to felhom-agent v0.54.0 (format-safety foundation). During agent install, fetch + install the
guarded-mkfs wrapper so the agent's format path is safe on any box.
- **New step in `step_agent_install`:** fetch `configs/felhom-mkfs-guarded.sh` from Gitea, `bash -n`
validate, `install -m0755 -o root -g root``/usr/local/sbin/felhom-mkfs-guarded`. Installed BEFORE
the sudoers (which now allowlists ONLY the wrapper, not raw `mkfs.*`), so the ordering is gap-free.
- The agent v0.54.0 sudoers (fetched by the same step) drops the raw `mkfs.ext4 -F /dev/* / mkfs.xfs -f
/dev/*` allowlist and permits only `felhom-mkfs-guarded /dev/* *`, plus read-only `pvs`/`zpool` for
the agent's unclaimed-disk guard. No other host-install change.
- `bash -n` + `shellcheck` clean (0 new warnings). Live-validated on felhom-pve (agent v0.54.0 deploy):
wrapper refuses the OS disk + an LVM-PV partition, raw mkfs is sudo-denied, an unclaimed throwaway
disk formats; the agent guard's sudo reads (pvs/lsblk/zpool) all work as the felhom-agent user.
## felhom-host-install.sh v1.7.0 — 3b-fix: `Datastore.Audit` box-wide (restore drive visibility) (2026-07-01)
Fixes a regression the v1.6.0 pool-scoped ACL introduced: `Datastore.Audit` was placed in the
per-storage `Store` role (granted only on `local`/`local-lvm`/`felhom-pbs`), which **excluded the
enrolled removable drives** `felhom-usb`/`felhom-flash`. The agent enumerates storage via
`ListStorage`/`NodeStorage` (both gated by `Datastore.Audit` — `internal/storage/observe.go`), so it
could no longer SEE the drives → false "Meghajtó leválasztva" (drive detached) alerts + drives absent
from the agent-view. (The v1.6.0 swap's "felhom-usb → 403" was mis-read as blast-radius success;
felhom-usb is Felhom's OWN customer drive, not an out-of-scope object.)
- **`Datastore.Audit` moved from Store → Base** (`PVE_PRIVS_BASE` now `"Sys.Audit SDN.Use
Datastore.Audit"`; `PVE_PRIVS_STORE` now `"Datastore.Allocate Datastore.AllocateSpace"`). Audit is
read-only metadata, so box-wide Audit restores visibility of ALL storages (incl. dynamically-enrolled
drives — no per-drive grant ever needed) while the **write** privs (`Allocate`/`AllocateSpace`) stay
per-storage → write/allocate blast-radius containment is UNCHANGED. Confirmed at source: the agent
creates no PVE storage (no `POST /storage`/`pvesm add`); drives are dir-storages it observes + mounts
via host ops, so they need only Audit, never Allocate.
- **`apply_scoped_acl` reordered** Base-before-Store (role + grant) so a RE-APPLY on a live box adds
`Audit@/` before Store drops its per-storage Audit → gap-free (the agent never loses enumeration).
- `remove_scoped_acl` / `--uninstall` / `--rescope-acl` operate by role NAME and inherit the corrected
privs automatically (no other change).
- **Live-repaired felhom-pve** (two `pveum role modify`, Base first — no agent stop/restart): drives
reappeared (agent-view 3→5 storages), detach alerts cleared. Re-tested under the scoped token: drives
readable (was 403), write-containment intact (vzdump→felhom-usb still 403; out-of-pool guest 403),
PBS Store grant unchanged. `bash -n` + `shellcheck` clean (0 new warnings).
- **NOT physically run** (source-confirmed, no `Datastore.Allocate` in the path): a brand-new-drive UI
enrollment (needs a spare USB) — the host-ops/Audit path is unchanged from pre-3b.
## felhom-host-install.sh v1.6.0 — pool-scoped token ACL (3-role) + `--rescope-acl` retrofit (2026-07-01)
Colleague-safety batch #4 phase b (script half; agent half = v0.53.0). Moves the agent token's dangerous
privileges off `/` (which spanned every guest + storage) to `/pool/felhom` + `/storage/<targets>`, so on
a shared box the token can only touch Felhom's own guests + storages. Validated by
`documentation/audits/SPIKE-pool-scoped-acl-2026-07-01.md` (PASS) — implemented here.
- **3-role scoped ACL (`step_token` rewrite).** Replaces the single `FelhomAgent` role granted at `/`
with three roles, each granted to BOTH the user AND the token (privsep intersection): `FelhomAgentGuest`
(`VM.*` + `Pool.Allocate`) @ `/pool/felhom`; `FelhomAgentStore` (`Datastore.*`) @ each of
`PVE_STORAGES` (default `local local-lvm felhom-pbs` — the offsite PBS MUST be included, SPIKE
residual #1; `--acl-storages` overrides); `FelhomAgentBase` (`Sys.Audit SDN.Use`) @ `/`. Helpers
`apply_scoped_acl`/`remove_scoped_acl`/`_grant`/`_ensure_role`.
- **Pool before token.** `ensure_felhom_pool` runs at the top of `step_token` (always, incl.
`--skip-provision`) so `/pool/felhom` exists before it's granted on.
- **Re-install safety.** `step_token` also removes the pre-3b broad `/` grant + `FelhomAgent` role if
present (`remove_old_broad_acl`, tolerate-absent), so a re-install can't leave the old grant unioned
with the scoped one. The post-provision `pool_add_guest` is gone (the agent's `restore --pool` makes
the guest a member atomically — v0.53.0).
- **`--rescope-acl` retrofit** (new mode, mirrors `--adopt-pool`): migrate an existing install — ensure
the pool + guest membership, apply the scoped grants, THEN remove the old broad grant (add-before-
remove: the token is never grant-less mid-migration). Prints the "now deploy agent ≥ v0.53.0"
ordering reminder. Idempotent + dry-run-aware. **SUPERVISED** (run with the agent stopped — the scoped
ACL and the pool-param agent are mutually dependent; §13 of the task).
- **`--uninstall`** now removes the scoped grants + 3 roles AND the pre-3b broad grant/role (both
tolerate-absent → works on either shape), keeping the pool delete-if-empty (v1.5.0).
- **Validated on felhom-pve** (dry-run): T-A fresh install (pool-before-token, 3 roles once, scoped
grants incl. `/storage/felhom-pbs`), `--rescope-acl` (add scoped → remove old `FelhomAgent`), T-F
uninstall (old-shape cleanup + pool not-empty skip). `bash -n` + `shellcheck` clean (0 new warnings).
**The live rescope + agent swap is the supervised STOP** — not run here.
## felhom-host-install.sh v1.5.0 — `felhom` pool by default + `--adopt-pool` retrofit + uninstall teardown (2026-07-01)
Colleague-safety batch #4 phase a. Every Felhom-managed guest now joins a dedicated **`felhom` pool**
for fleet uniformity (and as the environment the later pool-scoped ACL — 3b — will spike against). All
pool ops run as `root@pam` from the installer, so there is **NO agent/token/ACL change** and zero
permission-model risk (`PVE_PRIVS` untouched; the `FelhomAgent` token stays scoped at `/`).
- **New `felhom` pool default.** `step_provision` calls `ensure_felhom_pool` (create if absent,
idempotent) and, after a successful provision, adds the guest via `pveum pool modify felhom -vms
<vmid>` (skip-if-already-member). New helpers `pool_exists` / `pool_members` / `ensure_felhom_pool` /
`pool_add_guest`; const `PVE_POOL="felhom"`. PVE 9 syntax + `/pools` JSON shape confirmed live before
wiring (`pveum pool add|delete|modify`; `pvesh get /pools` → `[{poolid,comment}]`, `/pools/<id>` →
`{members:[{vmid,…}]}`).
- **`--adopt-pool` retrofit mode.** Non-destructive: adds an EXISTING Felhom guest to the pool (creating
it if needed), resolving the guest from `--vmid` else the recorded `provisioned_vmid`. Reuses the
ours-check (`/etc/felhom-bootstrap` mount) — refuses a non-Felhom guest unless `--force`. Touches ONLY
pool membership: never reconfigures/restarts the guest, never contacts the hub. Idempotent
(skip-if-member).
- **`--uninstall` pool teardown (step 5b).** After the pveum removal, deletes the `felhom` pool **only
if empty** (a destroyed guest is auto-removed from its pool); a pool that still has members is left
with a `log_skip` naming them. Not reached on the Spec-1 safe-skip path (other Felhom guests remain).
- **Validated on felhom-pve** (dry-run + SAFE live): T-A fresh-install dry-run shows the pool create +
membership lines; T-B **live adopt of guest 9201** → `pvesh get /pools/felhom` lists 9201, guest still
running, config unchanged (the demo node is now pool-uniform); re-run = no-op; T-B' non-Felhom vmid →
refusal; T-C uninstall dry-run → "pool felhom not empty (members: 9201) — leaving it". `bash -n` +
`shellcheck` clean (0 new warnings; the 2 pre-existing SC2015 in `step_verify` unchanged).
- **NOT changed:** `PVE_PRIVS`, the ACL grants, the agent, the provision-call args. 3b (pool-scoped ACL
+ agent restore-into-pool under a scoped token) is the separate spike-gated task.
## felhom-host-install.sh v1.4.0 — appliance CPU/RAM cap passthrough (`--cores` / `--memory`) (2026-07-01)
Colleague-safety batch #3 (host-install half; the mechanism is agent v0.52.0). Lets an operator cap the
provisioned guest so a trial appliance on a SHARED production Proxmox doesn't pressure the colleague's
existing guests.
- **`--cores N` / `--memory M` (MiB)** — optional; passed through to the agent's `--selftest=provision`
as `-cores`/`-memory`. `0`/unset = keep the golden's baked sizes (unchanged behaviour). New vars
`CPU_CORES`/`MEM_MIB`; `usage()` header gains an "Appliance cap (optional)" group.
- **Conditional passthrough** — `step_provision` builds a `cap_args` array and appends the flags to BOTH
the dry-run log and the real agent call **only when set**. An agent < v0.52.0 would reject an unknown
flag, so the flags are never sent unless the operator opts in (see the deploy dependency below).
- **Pre-flight sanity WARN (soft, provision only)** — if `--cores` > host `nproc` or `--memory` > host
`MemTotal`, `log_warn` "the cap won't protect other guests"; never `die` (the operator may know better).
- **Deploy dependency:** a fresh install using `--cores`/`--memory` needs the hub artifact manifest to
serve **agent ≥ v0.52.0**.
- **Validated dry-run on felhom-pve:** `--cores 2 --memory 4096 --dry-run` → provision command shows
`-cores 2 -memory 4096`; without the flags → neither present; `--cores 64 --memory 65536` → both WARN
lines (host 4 cores / ~15771 MiB). `bash -n` + `shellcheck` clean (0 new warnings; the 2 pre-existing
SC2015 in `step_verify` unchanged).
## felhom-host-install.sh v1.3.0 — `--uninstall` (clean revert) + pre-flight guards (2026-07-01)
Colleague-safety batch #1+#2. Adds a first-class, guarded **`--uninstall`** teardown so an operator can
cleanly back out of a trial install, plus three provision pre-flight guards that stop common footguns.
Script-only; no agent/hub/controller change.
- **`--uninstall` (local host teardown — no hub contact, no passphrase).** Reverses an install in the
install-order's reverse: **guest → agent(unit/sudoers/binary/state/user) → pveum(ACL,token,user,role)
→ golden(opt-in) → state file.** Every mutation goes through `run()` so `--dry-run` prints the full
plan and executes nothing. Safety:
- **Ours-check:** refuses to destroy a guest that lacks the `/etc/felhom-bootstrap` bind mount (matched
by the constant guest *path*, not a hardcoded `mpN` slot — on the demo host it's `mp9`), unless
`--force`.
- **Typed confirmation:** must type the vmid to confirm PERMANENT destruction (read from `/dev/tty`;
skipped only under `--dry-run`, where nothing is destroyed).
- **Other-guests guard:** if any OTHER Felhom guest remains, destroys only the target and **leaves the
agent + PVE token + state in place** (re-run with `--force` to remove host-level anyway — orphans the
others).
- **Never removes the `sudo` package**; never contacts the hub (the host record intentionally persists).
- Presence-checked + idempotent: an already-absent guest/unit/sudoers/binary/user/ACL/token/role is a
tolerated skip, not an error. The `pveum role delete` runs only after its ACL grants are gone (PVE
refuses to delete a referenced role). Confirmed PVE 9 ACL-delete form:
`pveum acl delete / --users|--tokens <x> --roles FelhomAgent`.
- Target vmid resolves from `--vmid`, else the recorded `provisioned_vmid` (else dies). A `--vmid` that
disagrees with the recorded one needs `--force`.
- **`--remove-golden`:** with `--uninstall`, also delete the golden vzdump from the archive storage
(`pvesm free`); otherwise it is left in place.
- **Install state now records `customer_id` + `provisioned_vmid`** (new `_state_put`/`_state_get` helpers,
dry-run-guarded like `_state_mark`; the `completed[]` shape is untouched) so a later `--uninstall`
resolves its target automatically and safely.
- **Pre-flight guards (provision mode):**
- **Multi-node guard** — on a 2+-node cluster, `die` (naming the nodes) unless `--node` is explicit
(new `NODE_EXPLICIT`); single-node keeps the current auto-pick. No-op under `--skip-provision`.
- **Archive-storage-exists guard** — verify `--archive-storage` appears in `pvesm status` (else `die`);
no-op under `--skip-provision`.
- **RAM floor (WARN, never fatal)** — warn when `MemAvailable < 2048 MiB`.
All three run inside `step_preflight` (before any mutation) so they also fire under `--dry-run`.
- **Validated dry-run-only on felhom-pve** (single-node, live guest 9201): T-A full uninstall plan, T-C
not-ours refusal (red-proof), archive-missing `die`, RAM line, other-guests detector, state round-trip;
confirmed 9201 + agent + pveum + state untouched after all dry-runs. `bash -n` + `shellcheck` clean
(0 new warnings vs. baseline; the 2 pre-existing SC2015 in `step_verify` are unchanged). **NOT yet
live-validated (awaiting a supervised run):** a real live `--uninstall` (guest destroy + pveum removal)
and the multi-node guard on an actual cluster.
## felhom-host-install.sh v1.2.0 — /dev/tty passphrase read + vmid auto-detect (2026-07-01)
Two operator-experience fixes so a colleague can install online (via the hub's new "Option 1: Online
install" one-liner) and onto a host that already runs a guest at 9201.
- **Passphrase prompt reads from `/dev/tty`, not stdin** (`read_passphrase`). `read -rsp … < /dev/tty`
makes the no-echo prompt work regardless of how stdin is wired — both download-then-run **and**
`curl … | sudo bash` (where stdin is the pipe). Strictly more correct; the `--passphrase-file` path is
unchanged. The passphrase is still never on argv / in logs / in the state file.
- **VMID auto-detect (`--vmid` now optional-smart).** New `VMID_EXPLICIT` flag (set by `--vmid`). The
pre-flight vmid guard now determines "in use" against the **`pct list` + `qm list`** id-set (LXC and
VMs share the id space — more complete than the old `pct status`, which only knew LXC):
- **explicit `--vmid`** → unchanged deterministic behavior: die if the id is in use unless `--force`
(destructive over-provision).
- **default 9201, in use, no `--force`** → **auto-pick the next free id** (scan upward from 9201 over
the used-set) and **ask to confirm** from the terminal (`read … < /dev/tty`, `[y/N]`); proceed on
yes, `die "no free vmid confirmed"` otherwise. Never a silent auto-pick.
- **default 9201 + `--force`** → over-provision 9201 (destructive) without prompting, as before.
- New helpers `used_vmids` / `_vmid_in_use` / `next_free_vmid`. `--vmid` help text + `usage()` updated.
## felhom-host-install.sh v1.1.0 — self-install the agent + fetch the golden from Gitea (2026-06-28)
The script now **installs the agent itself** (the last big manual Day-0 prerequisite is gone). It
fetches the agent binary + golden from Gitea generic packages and **verifies each against the
hub-vouched artifact manifest** before installing/using it. BUNDLE slice; pairs with hub v0.16.0
(artifact manifest endpoint + operator UI) and felhom-agent v0.43.0 (canonical unit + publish).
- **New step `5/8 agent install`** (before agent-config): resolves the manifest
(`GET /api/v1/artifacts/{id}`, passphrase) + the git fetch token (from the customer's
`controller.yaml` via config-retrieve — **NO new credential**); fetches
`/api/packages/admin/generic/felhom-agent/<ver>/felhom-agent`, **verifies sha256 vs the hub
manifest** (aborts on mismatch — verify-before-use), backs up any existing binary, installs
`0755 /usr/local/bin/felhom-agent`; ensures the non-root `felhom-agent` system user; installs the
canonical sudoers (`0440`, `visudo -cf`-validated) + systemd unit; `daemon-reload` + enable. Idempotent:
same version already installed + service active → skip.
- **`--skip-provision`:** install + configure + verify the agent (incl. golden fetch+verify) but do NOT
provision a guest — the agent-only path for re-installing/upgrading the agent on a host that already
has live guests. Adds an agent-only `step_verify_agent` (binary + non-root service active + a
`--selftest=hub` collect-report).
- **New step `7/8 golden`:** local auto-discovery stays the default/fallback; otherwise fetches
`/api/packages/admin/generic/felhom-golden/<ver>/golden.tar.zst`, **verifies sha256**, and imports it
into the archive storage's dump dir for the restore. `--force-gitea-golden` forces the Gitea path.
- **Non-root agent model:** the agent now runs as `felhom-agent` with `privileged.mode: "sudo"` (was the
dev/CI `direct`+root shortcut). The config is `chown`ed to the service user (0600) so the daemon can
read it; `systemctl is-active` after restart is the real proof the non-root user can read the config.
- **Pre-flight relaxed:** a missing agent binary is no longer fatal (step 5 installs it); the local
golden requirement is deferred to step 7.
- **Trust model:** checksum **trust root = the hub** (manifest), not Gitea; the fetch credential is the
existing config-retrieve git token; artifacts are pinned to a version (never `:latest`).
- **Secrets:** the git token is a never-logged runtime carrier (cleared on EXIT alongside the passphrase
/ pve-token / hub api_key); the sudoers is `0440` and `visudo -cf`-validated before install.
- `bash -n` + `shellcheck` clean.
## felhom-host-install.sh v1.0.0 — Day-0 host bootstrap (provision mode) (2026-06-26)
First release. A single operator-run script that automates Day-0 on a freshly-PVE-installed
host: Proxmox API token → hub host enrollment (option C, single secret) → agent config →
guest provision → verify. Composes proven mechanisms (the `pveum` role/token sequence, hub
`POST /host-enroll`, `felhom-agent --selftest=provision`); grounded by
`documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md`.
- **7 steps, idempotent + resumable** via `/var/lib/felhom-install/state.json`: pre-flight →
Proxmox token → compute grows → host-enroll → agent config → provision → verify.
- **Single-secret** (the retrieval passphrase): read no-echo or from a 0600 file, never on
argv/logs/state. The global operator key never touches the box.
- **pveum automation:** 16-priv `FelhomAgent` role (create-or-modify), `felhom-agent@pve` user,
privsep token (reuse-if-working else rotate), and **both** ACL grants applied **after** the
token exists (token-remove purges the token ACL).
- **Auto-discovery:** golden archive (newest `vzdump-lxc-<golden-vmid>`), PVE node name, vmbr0
bridge IP for the local-api, and the served-leaf TLS fingerprint pin.
- **Safety:** pre-flight fails fast (root, PVE 9.x, local-lvm headroom, hub reachable,
customer+passphrase valid via read-only `GET /config/{id}`, golden resolvable); refuses to
clobber an existing `--vmid` without `--force`; `--dry-run` previews every mutation;
`--preserve-from` keeps operator infra (PBS/local_api/privileged/authz) on re-deploys.
- **`--mode dr`:** documented 10D stub (restore customer PBS snapshot instead of golden) — not
implemented.
- **Live-validated** end-to-end on `felhom-pve`: authorized wipe of demo guest 9201 →
re-provision from the golden → controller config-pull + public tunnel `HTTP 200` →
host-report of guest 9201 → idempotent `--resume` no-op. (One ordering bug — token ACL
applied before rotation — was found and fixed during the live run.)