Files
felhom.eu/scripts/CHANGELOG.md
T
admin 592818492c hub v0.66.0 + ISO v1.20.0: customer self-bind (R-27 slice 1)
Let a customer bind their own freshly-installed appliance without the
operator: operator "Send self-bind link" mints a 7-day tokenized
capability link, emailed (Hungarian, sibling sender) to the customer, who
opens a public /bind/<token> page and proves two factors — the console
pairing code shown on the box screen + their retrieval passphrase — and
the hub stages the bind via the same BindAppliance (provenance
customer_selfbind). The box's ~30s appliance poll delivers.

Viktor's three rulings verbatim: console pairing code (no appliance list
ever rendered), operator-sent tokenized link, 5-attempt lockout ->
"call support". Wrong code == wrong passphrase (one generic failure, no
oracle, both factors compared unconditionally); expiry falls back to
operator-bind unchanged.

THE TRAP: one public prefix /bind/, exempt from auth+CSRF at both /login
gate sites via a single isPublicBindPath predicate (tight trailing-slash
match; ServeMux ..-cleans; handler rejects '/' in token). 9 tests
(Scenarios A-F + F1/F2); 4 red-proofs verified red-then-green (lockout,
oracle, widened-prefix, single-active). GC verdict: no appliance GC ->
the 7-day TTL stands alone. Controller/agent untouched; R-27b deferred.

Green: full hub build/vet/test (17 ok) + bash -n + hub confirm gate.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017qDiBqKKQ5vPB5fXBqu7Kp
2026-07-17 23:56:53 +02:00

58 KiB
Raw Blame History

Felhom scripts — Changelog

build-felhom-iso.sh v1.20.0 — console pairing-code banner for customer self-bind (R-27 slice 1) (2026-07-17)

Supports the hub's customer self-bind flow (hub v0.66.0). In PAIRING mode, felhom-bootstrap.sh now reads the additive pairing_code from the POST /api/v1/appliance/register response, persists it at /etc/felhom/appliance-pairing-code, and prints a Hungarian console banner (to /dev/console, stdout fallback) each pairing cycle so the customer can read the code off the physical screen and type it — together with their retrieval passphrase — on the hub's public /bind/<token> page. The code is non-secret (possession proof only; the passphrase is the second factor), so it is safe on the console. Graceful degradation both ways: a hub older than v0.66.0 omits pairing_code → the banner prints nothing and register/poll are unchanged; an old ISO against a v0.66.0 hub simply ignores the new field. No change to DIRECT mode. Green: bash -n clean on both scripts.

build-felhom-iso.sh v1.19.0 — the universal secret-free ISO: --pairing mode (R-21 slice C) (2026-07-17)

The scripts half of the universal ISO. felhom-bootstrap.sh gains a PAIRING mode — one unit, two modes, decided by the env:

  • DIRECT (env has FELHOM_CUSTOMER_ID + FELHOM_RETRIEVAL_PASSPHRASE): the slice-A path, byte-identical — only branched around. Scenario D regression proves the pairing code is provably NOT entered (a fake hub records ZERO /appliance/* calls).
  • PAIRING (generic ISO, no customer/passphrase baked in): gather identity (SMBIOS uuid + physical MAC set + SSH host pubkeys + hw), POST /api/v1/appliance/register → persist the appliance token (0600), then ONE GET /api/v1/appliance/poll per invocation (the existing Restart=on-failure/RestartSec=30 IS the poll timer — no long-running-oneshot timeout). On the bind's 200 delivery, WRITE the delivered credentials into the env (0600) and fall through to the DIRECT path — so every later retry is a plain direct install (the delivery is one-shot; a second poll → 410). Delivery-received secrets + the appliance token are shredded on host-install success.
  • build-felhom-iso.sh --pairing builds the GENERIC ISO: no --bootstrap-env, a secret-free env carrying only the hub URL, -generic filename marker, manifest mode: pairing + secret-bearing: no. Direct mode (secret-bearing) is unchanged. New profiles/generic.profile.
  • Validated: bash -n + shellcheck clean; the test/bootstrap-modes.sh harness (Scenario D + pairing register/poll + the delivery→env→host-install handoff) all green in a debian container.

build-felhom-iso.sh v1.18.0 — firmware loader option --loader shim|mkimage (R-21 slice B, F1) (2026-07-17)

Closes N100 finding F1 (HIGH): cheap AMI (AN3PLUS 0.01-class) UEFI firmware can't relocate the ISO's stock signed GRUB from USB (relocation 0x0 is not implemented yet; SB-off + shim-bypass don't help). The run's live grub-mkimage workaround is now a first-class pipeline mode.

  • --loader shim|mkimage (default shim; profile-settable via FELHOM_LOADER, --loader wins). shim = the stock output, byte-for-byte unchanged (the surgery is skipped) — keeps the MS-signed shim→GRUB chain and Secure Boot on compliant firmware (S2b). mkimage replaces the ISO's UEFI boot path with a monolithic grub-mkimage loader built from the ISO's own GRUB modules (module set from the ISO's grub.cfg; embedded search --fs-uuidconfigfile the real menu). The recipe is reproduced from the N100 run evidence, not re-derived.
  • mkimage-surgery.sh (new): runs in the assistant container after prepare-iso, so the answer/first-boot payload is provably untouched except the loader. Swaps the loader into the ISO9660 tree (real lowercase Rock-Ridge path) and inside the efi.img ESP (FAT), then re-masters with xorriso preserving the hybrid BIOS boot (grub2-mbr + El Torito) and the GPT EFI System Partition, dropping only the irrelevant Apple HFS+/APM map (which trips xorriso's overlap check on re-emit).
  • Loud + unmistakable (rule 4): build banner, manifest loader: + grub-mkimage: fields, and a -mkimage filename suffix when non-default. mkimage ISOs are unsigned → Secure Boot must be OFF on the target board.
  • Dockerfile.assistant: adds grub-common + grub-efi-amd64-bin + mtools + dosfstools (grub 2.12 == the PVE 9.x ISO's 2.12-9+pmx2 generation). New profiles/n100.profile (cheap-board, FELHOM_LOADER="mkimage", SB-off note).
  • Validated (RUNBOOK-B legs, nested VM 311 on felhom-pve): leg 1 shim boots + zero-touch installs under OVMF Secure Boot enforcing (+ SeaBIOS boot — hybrid path survived); leg 2 --loader mkimage boots + installs under OVMF SB-off (the monolith's embedded fs-uuid search found the ISO + chained its menu); leg 3 (red-proof) the mkimage ISO under SB-enforcing FAILS — firmware Access Denied, no boot; leg 4 the surgery is surgical (only the loader + regenerated El Torito artifacts differ; the full install payload is byte-identical). bash -n + shellcheck clean. Physical N100 closure folds into the supervised rehearsal (an n100-safety match-nothing ISO built
    • sha-recorded, unbooted, for an optional zero-risk pre-flight).

felhom-tenantsync.sh v1.1.0 — deprovision op (customer RESET teardown) (2026-07-17)

Adds the {"op":"deprovision","customer_id":"<id>"} op the slice-1 header explicitly reserved ("namespace/data deletion is a deliberate, separate decision"). It is exactly that deliberate, hub-side ack-gated decision (the customer RESET, hub v0.61.0): delete the token (its ACLs purge with it) → delete the residual namespace ACLs → destroy the namespace AND all its backup groups (proxmox-backup-client namespace delete <ns> --delete-groups true, via the transient admin token). IDEMPOTENT — a missing token / namespace is success (deleted:false), so a re-run after a partial reset converges. The shared felhom@pbs user is NEVER touched (co-tenants ride it). Returns {"status":"ok","namespace","datastore","deleted":<bool>}. Secret hygiene unchanged (no secrets in this path). Client seam: tenantsync.Deprovision(ctx, customerID) (existed bool, err error).

felhom-host-install.sh v1.17.0 — appliance guest auto-sizing (F5) + doc-drift fix (2026-07-17)

Closes VALIDATION-n100-baremetal-2026-07-16.md F5 (MEDIUM): appliance mode provisioned the golden default (2 GB RAM) on a 16 GB host and never surfaced the --memory/--cores caps. Now, in appliance mode with no explicit cap, the guest is auto-sized from the host (LXC limits are cheap → err generous): RAM = clamp(host-4096, min 4096, max host-2048) then a hard ceiling host-1024 (never over-commit); cores = host-1, min 2. An explicit --memory/--cores ALWAYS wins untouched; byo mode still requires explicit caps (never auto-sized). Host reads (MemTotal/nproc) are overridable via FELHOM_FAKE_MEMTOTAL_MIB/FELHOM_FAKE_NPROC for the mode harness. Sizing table: 8 GB→4096, 16 GB→12288, 32 GB→28672 MiB; a 4 GB host→3072 (min capped at host-1024). Red-proof: hostinstall-mode-harness.sh new F5 section (8/16/32 GB + small-host edge + explicit-flag precedence) — pre-fix (auto-size disabled) FAILs, post-fix PASSes. Same commit (R-16 doc-drift leftover): the operator-signing-keys "EMPTY by default" comment was stale (the keys are PINNED to the real felhom-op-1/felhom-rec-1 ceremony keypairs) — corrected, comment-only, no behavior.

felhom-poke.sh v1.0.0 — NEW (agent-plane immediate-sync, Direction-2a, 2026-07-16)

The offsite endpoint's THIRD hub forced-command surface (documentation/runbooks/offsite-endpoint.md §11): a NON-root (felhom-peersync, no sudoers grant — a datagram needs no privilege) forced command that reads a box's WireGuard /32 from $SSH_ORIGINAL_COMMAND, validates it to 10.77.0.0/24, and sends ONE contentless (empty) UDP datagram from wg0 to <ip>:51822. That is the last hop of the hub→ep0→box "sync now" poke; the box's felhom-agent (v0.89.0) poke listener then runs an immediate desired-state cycle. Contentless + confined (the WG kernel independently refuses non-peer /32s — SPIKE-immediate-sync-transport-2026-07-16 P1 EKEYREJECTED). Port 51822 is a shared cross-repo constant (felhom-agent internal/poke.Port, hub/internal/poke, this script's POKE_PORT). Hub client: hub/internal/poke. Installed + live-validated on ep0 2026-07-16 (poke-fired; non-WG target and arbitrary SSH command both refused non-WG target — the forced command never executes client input). Peersync/tenantsync surfaces untouched.

felhom bare-metal ISO pipeline v1.16.0 — NEW (R-21 slice A, 2026-07-16)

New operator tooling under scripts/iso/: a reproducible DooPlex pipeline that turns the official PVE ISO into a Felhom auto-install ISO whose first-boot chain unattended-fetches-and-runs felhom-host-install.sh until the host is enrolled + a guest provisioned. felhom-host-install.sh is unchanged (invoked, not modified). Tagged v1.16.0 to align the ISO with the current host-install release (they must stay coherent; the ISO's baked invocation targets that flag surface).

  • Dockerfile.assistant (assistant paired to the ISO's PVE generation by Debian codename + shellcheck), build-felhom-iso.sh (render answer from answer.toml.tmpl + a profile, mint a fresh THROWAWAY root hash per build, gate validate-answer on OUTPUT TEXT not $? — spike S1 exit-0-on-failure trap, prepare-iso --fetch-from iso --on-first-boot, emit ISO + sha256 + manifest).
  • In-ISO chain: stub-first-boot.sh (dumb, exactly-once, from-iso/fully-up) installs felhom-bootstrap.sh + felhom-bootstrap.service (retry-forever Type=oneshot+Restart=on-failure)
    • a 0600 bootstrap.env. Retry-vs-resume encoded once (plain first; --resume when /var/lib/felhom-install/state.json exists — safe per host-install v1.11.3).
  • Secret-bearing: the unattended install needs the customer retrieval passphrase, so the ISO embeds it (the distribution URL itself carries no secret). Supervised/single-use; the bootstrap scrubs the env on success. Slice C removes the secret from the distributed ISO via claim-pairing.
  • Validated on nested VM 310: build gate + red-proof; disk-filter fail-safe through the pipeline; stub→unit→real public-channel fetch→invocation→retry; resume decision; exactly-once; no-net retry + recovery. Terminal host-install success (hub-enrolled + guest-provisioned) is operator-gated (drill customer needs the password-gated create-UI; live-DB insert unsafe). Full report: root REPORT.md, documentation/audits/SPIKE-baremetal-iso-2026-07-16.md.

felhom-host-install v1.16.0 — FELHOM_ESCROW rides the canonical sudoers (2026-07-13)

Companion to agent v0.88.0 + controller v0.127.0 (the controller-driven escrow ceremony wizard). No new installer STEP: the FELHOM_ESCROW sudoers alias (the wizard's fixed-argv sudo self-invocation grant) ships via the existing canonical-sudoers fetch (configs/felhom-agent.sudoers from the agent repo, visudo -cf-gated before install) — the bump documents that every ≥1.16.0 install/update carries it. Hub hostInstallVersion bumped in the same commit (hostinstall_gates.py asserts equality); the hub deploy rides the next hub train (display-only lag, accepted). EXISTING hosts get the refreshed sudoers with the agent-0.88.0 deploy (ship-configs step), not from this script.

manifest_bearer_gate.py — new gate: no bearer-shaped literals in manifests/ (2026-07-13)

Part of the hub v0.53.0 bearer de-git. Greps every manifests/*.yaml for 64-hex literals (the openssl rand -hex 32 shape all felhom bearer/API keys use), comments included; exit 1 on any hit. felhom.secret.yaml's pre-existing committed secrets (umami APP_SECRET is 64-hex) print as a visible non-fatal KNOWN-BACKLOG line until that file's own de-git (secrets.md backlog). Run after ANY manifests/ change. Named "bearer" not "secret" — the repo .gitignore *secret* pattern silently un-tracks any "secret"-named file (learned the hard way: the gate's first landing never reached git). Red-proven: reintroducing a 64-hex literal → exit 1.

felhom-host-install v1.15.0 — DR-tier-by-default: uniform DR plumbing on every install (2026-07-12)

Sweeps the DRILL-day0-vm-2026-07-12 installer findings so the next Day-0 runs the drill's §5 sequence with zero fix-and-continue stops. Operator decisions 2026-07-12: capability is BAKED uniformly; activation is a hub flag; WG is base infrastructure.

  • F-7 — ships /usr/local/sbin/felhom-pbs-apply (fetch configs/felhom-pbs-apply from the agent repo, bash -n, install 0755 — the mkfs/selfupdate wrapper pattern). The FELHOM_PBSDR sudoers alias already shipped; the 3 pbsdr-* capabilities are no longer born DEGRADED. Uninstall removes it (disclosure↔uninstall parity kept, harness GL4-D green).
  • F-10 — installs the age package (escrow-ceremony identity-wrap dependency; a fresh-box ceremony died on it live).
  • F-9 / decision 5 — the rendered agent.json now defaults wg_tunnel.enabled: true in BOTH modes (WG is base infrastructure like cloudflared, also the future OOB transport; setdefault — a deliberately-disabled preserved config is respected). The byo wg-off config assert is RETIRED; byo consent = the typed disclosure ack (the WG tunnel + pbs-apply wrapper + age are now on the disclosed list) + the tester-agreement disclosure line.
  • ACL — the default --acl-storages set keeps felhom-pbs (drill-proven: narrowing it 403s the PBS-DR apply-bridge at /storage/felhom-pbs); a missing felhom-pbs storage is now an expected-pre-tier INFO, and no longer a byo refusal (the tier creates the storage later).
  • F-1 — single version source: SCRIPT_VERSION alone feeds -h, the run banners, and (via a gate cross-check) the hub Setup-tab const hostInstallVersion (bumped 1.12.0→1.15.0 in the same commit). The header no longer carries its own version literal.
  • F-2 — dry-run fetch lines print the REAL auth mode (curl (anonymous …) vs curl -u <git>) instead of the hardcoded -u <git> lie on the anonymous branch.
  • F-8 — step 4b + the final summary now say out loud that root@pam was rotated + vaulted and where to retrieve it (pointer only, never the credential).
  • New gate: scripts/hostinstall_gates.py — mechanical grep-assertions (version single-source
    • hub-const sync, age, pbs-apply ship+uninstall, wg default-on + byo-assert retirement, felhom-pbs in the default ACL). All red-proofed; the age assertion was hardened TWICE by its red-proof (a prefix regex matched agekit, then a loose one matched the dry-run print line).
  • F-3 (root-owned guests/ parents) is fixed at its creator — the AGENT (v0.86.0); this pass applied the drill's live-fix (chown the two parent dirs) to felhom-pve, which had the same latent state. Peti's host: unreachable, check deferred.

felhom-host-install v1.14.0 — deployment_mode from --mode (appliance self-heal gate) (2026-07-12)

  • The agent-config writer now templates the top-level deployment_mode field from --mode (appliance | byo), authoritative even over a --preserve-from config. It gates the agent's node self-heal (host networking recovery at boot — F12-class defense in depth, agent v0.85.0): only deployment_mode:"appliance" unlocks the remedy; anything else, including a stale preserved value or a typo, is byo (the check runs + WARNs, the remedy is unreachable). The byo config assert additionally refuses deployment_mode=appliance (a byo box must never carry a host-service self-heal). Companion to CAMPAIGN-3 Task A (documentation/audits/CAMPAIGN-3-2026-07-11.md).

felhom-host-install v1.13.0 — systemd-journal group for the agent user (NAS verify) (2026-07-11)

  • The agent-user setup now adds felhom-agent to the systemd-journal group (idempotent usermod -aG): the NAS verify-before-commit pipeline (agent v0.81.0) classifies mount failures from an UNPRIVILEGED journalctl -u <unit> read — group membership instead of a sudoers grant (rule: no journalctl via sudo). Existing installs get the one-liner usermod -aG systemd-journal felhom-agent && systemctl restart felhom-agent (applied to the demo host during this task's live validation; Peti's box only when the NAS feature reaches him).
  • Version-header drift fixed: the file header still said v1.11.0 while SCRIPT_VERSION was 1.12.0 (the v1.12.0 bump shipped without a changelog entry or header sync); both now 1.13.0.

felhom-host-install v1.11.0 — uninstaller gap-closure + operator-key day-0 fold (TASK GL-4, go-live G6 + G1-follow-up) (2026-07-08)

  • Uninstall gap-closure (G6):
    • NEW 4b4 block removes the self-update artifacts the install lays down but v1.10.0 never removed: /usr/local/sbin/felhom-selfupdate-guarded, the A/B slot files next to the live binary (felhom-agent.prev + orphaned .new.* temps), felhom-agent-rollback.service, and the felhom-agent.service.d/felhom-agent-limits.conf drop-in (+dir). Paths derived from the authoritative felhom-agent/configs/felhom-selfupdate-guarded; pending.json was already covered by the state-dir removal. Tolerate-absent, 4b2 shape.
    • Enrolled/network drives mounted under /mnt/felhom-drives/<name> are now unmounted (deepest first) before the root self-bind — plain umount ONLY, never -l/-f: a busy mount gets a warning + "eject via the dashboard or stop the apps and retry" and the root bind is then left alone (previously a child mount made the root umount die mid-teardown). The data is NEVER touched — no wipe/format path exists anywhere near /mnt/felhom-drives.
    • Both modes now END with an explicit KEPT-vs-WIPED statement: WIPED mirrors what the mode actually ran; KEPT names the drives + their data (physically removable; busy ones listed as retry), PBS backups + customer namespace (conditional on a pbs storage), the hub host/customer record, the escrow blob, and — when step 4b had vaulted one — the root@pam recovery credential.
    • Guest-only mode prints the drives the agent's bind store records for that vmid (best-effort from /var/lib/felhom-agent/guest-binds.json, generic note otherwise) + "eject BEFORE uninstall" guidance; no umounts there (remaining guests may use the drives).
  • Operator-key day-0 fold (the GL-1/G1 key-pin follow-up): OPERATOR_KEY_{OPERATIONAL,RECOVERY}_{ID,LINE} constants (EMPTY until the operator's offline pin ceremony) + --operator-pubkey-file PATH (lines operational|recovery <authorized_keys line>; the comment field IS the key_id, required; malformed/empty file dies at argv time naming the line; file OVERRIDES the constants with a notice). Resolved keys are written to authz.signers in the agent config (exact config.go SignerKey schema; nonce_store_path left to the agent default). Preserve rule: script/file keys are authoritative ONLY when non-empty — a reinstall with no keys KEEPS a preserved config's signers (never silently clobber a manually-pinned box; replacing preserved signers logs a pin-rotation notice). byo disclosure names the operational key_id (or "NONE — dormant"); verify reports authz signers: N (armed) or a dormant WARN (dormant = safe default).
  • Harness: +13 static cases (GL4-C1..C5 incl. 5 live key-file refusals + grep shapes, GL4-D disclosure↔uninstall parity, GL4-INV no-forced-unmount/no-format invariant) + PVE-tier GL4 H-U (full uninstall DRY transcript against the live guest, state-override protected). 28/28 PASS on felhom-pve; red-proofs RP-1 (4b4 dropped → parity FAILs), RP-2 (preserve guard dropped → C4 FAILs), RP-3 (role-die dropped → C3a FAILs) all run→fail→revert on scratch copies.
  • NOT live-validated (GL-6): the real (non-dry) uninstall with mounted/busy drives, the statement on a real teardown, and an armed key-pin install end-to-end.

felhom-host-install v1.10.0 — --mode appliance|byo install profile (TASK GL-2, go-live G2/G4/G5) (2026-07-07)

  • --mode appliance|byo is now REQUIRED for a fresh install / --resume (no default — the build-golden v2.0.0 precedent). appliance = the historical behavior, dry-transcript-preserved. byo = a Proxmox host the operator does NOT own (the pilot): step_break_glass is gated OFF at its call site (root@pam never reset, nothing vaulted — the one place the skip is auditable), --cores + --memory both mandatory (the only noisy-neighbor protection there), --enable-oob/--rotate-recovery refused at argv time, PVE-major gate (die on non-9.x unless --force), every --acl-storages entry must exist on the box (die naming ALL missing; appliance warns), informational host-DNS :53 line, full host-mutation disclosure + typed-hostname acknowledgement (/dev/tty, dry-run-aware), and post-write config asserts that refuse lan_resolver.enable/wg_tunnel.enabled/oob.enabled = true (byo also flips the config write's lan_resolver DEFAULT to off; appliance keeps default-on).
  • The old --mode provision value is retired (specific die message); the dr stub seam is kept. The mode is persisted in state.json; a --resume under the other mode refuses (C4).
  • Verify asserts (BOTH modes — campaign-2 R2 lesson): the verify step now asserts pool membership of the provisioned vmid + every scoped ACL grant (user AND token: Base@/, Guest@/pool/felhom, Store@each storage); any miss → error + "WITH WARNINGS" verdict.
  • --preflight-only: runs the full mode-aware step-1 checks, prints an explicit PRE-FLIGHT PASS/FAIL verdict, exits 0/1, writes NO state marks (never lets a later --resume skip the real preflight). Existing-guest count line added to preflight (both modes).
  • FELHOM_INSTALL_STATE_DIR env override (test-harness only) so harness cases can never touch a live state.json; usage() switched to a marker-based range (the numeric one had drifted).
  • NEW scripts/hostinstall-mode-harness.sh: static tier (bash -n, shellcheck, C1C4 refusal matrix asserting exact die messages, grep-invariants: one gated break-glass call site, chpasswd unreachable outside it, --mode in usage) + PVE tier (C5 live-refusal + Scenario A/B dry transcripts via FELHOM_TEST_CUSTOMER/FELHOM_TEST_PASSFILE). Red-proofs RP-1..RP-3 executed (gate removed / caps requirement dropped / mode-mismatch check dropped → matching case FAILs).
  • Live-validated read-only on felhom-pve: full static tier + C5 + byo dry transcript (no 4b/8/chpasswd/recovery-credential; disclosure+ack+DNS+storage lines present) + --preflight-only PASS and FAIL verdicts with zero state entries. Full A/B rc=0 transcripts + C7 await the supervised GL-6 drill (blocked today by demo-felhom's EMPTY git.username/git.token — any step-5 re-run for that customer dies until the operator sets real Gitea read creds).

felhom-host-install — --enable-oob installs the dedicated felhom-sshd OOB instance + belt (TASK H1) (2026-07-05)

  • --enable-oob + install_oob (in step 5): lays down the STATIC OOB parts — /etc/felhom-sshd tree
    • a dedicated ed25519 host key + the felhom-op login user + its scoped /etc/sudoers.d/felhom-op (visudo-validated) + the felhom-sshd.service unit (RuntimeDirectory-guarded — refuses the G1 incident cause) enabled-not-started (the agent renders the config) + the static inet felhom_oob nft table (nft -c-checked) + its boot loader felhom-oob-nft.service. The agent reconciles the DYNAMIC parts (port claim, config, operator authorized_keys, belt set elements) once oob.enabled.
  • Uninstall (step 4b3): stops/disables felhom-sshd + the belt loader, removes both units + the config tree + belt table + felhom-op user/sudoers; NEVER touches the stock sshd/:22//etc/ssh.

felhom-host-install — break-glass credential + management-plane watchdog install (TASK G1) (2026-07-05)

  • step_break_glass (new step 4b): generates a strong root@pam console password (openssl rand, NEVER logged/filed — stdin→chpasswd + stdin→curl), sets it, and vaults it to the hub over the enroll-authenticated host-key channel (PUT /hosts/{id}/recovery-credential). Idempotent (state marker) unless --rotate-recovery. The human break-glass path for the PVE web console.
  • install_mgmt_watchdog (in step 5): installs the G1 host artifacts from the agent repo configs — felhom-privsep.tmpfiles (boot-persistent /run/sshd), felhom-mgmt-watchdog.{sh,service,timer} (agent-independent ~60s privsep-dir auto-heal). HARD GUARD: refuses any fetched unit declaring RuntimeDirectory= (the SPIKE-felhom-sshd §8 incident cause). Non-fatal on a repo that predates the artifacts. Uninstall (step 4b2) stops+disables the timer and removes all four artifacts + the marker.

docs — 06-doc S3 row SHIPPED + agent-side revocation semantics (2026-07-04)

Docs-only companion to felhom-agent v0.64.0 (the S3 slice — keygen, registration, agent-managed wg-quick@wg-felhom, escrow join; live-validated on felhom-pve incl. revocation drill, reboot persistence, 30-min soak). 06-doc §3.5 now records: register-once marker, revoked-stays-revoked, re-add via the registration endpoint (the raw registry add doesn't bump the host generation — live finding), wg_tunnel.enabled default-FALSE rollout gate. S6 backlog notes added (hub poll constant configurable + first-adoption log; registry-add bump-or-label). CGNAT/mobile-hotspot appendix deferred (operator-assisted; §7's open validation stands).

felhom-peersync.sh v1.0.1 — strip out of process substitution (exit-swallow fix) (2026-07-04)

The S1 REPORT's exit-swallow class, fixed: wg syncconf wg0 <(wg-quick strip "$tmp") hid the strip exit code — a corrupt wg0.conf.head could feed syncconf empty/partial input that WIPES the live peer set while the script exits 0 (and the bad conf is then persisted). v1.0.1 runs strip as its own step into $tmpdir/stripped; a strip failure aborts BEFORE wg is invoked. Sandbox red-proof (stub wg-quick exiting 1 after partial output + recording stub wg): pre-fix shape invoked wg and returned rc=0; fixed shape errors first, wg never called. Redeployed to the dev endpoint (runbook step 5 install); shellcheck clean.

felhom-peersync.sh v1.0.0 — the offsite endpoint's WG reconcile script (2026-07-04)

S1 (doc 06 §5): the forced-command target the hub's wgsync pushes to (runbook offsite-endpoint.md step 5 installs it as /usr/local/bin/felhom-peersync, root:root 0755, invoked via a one-line sudoers grant from the felhom-peersync user's restrict,command= authorized_keys entry). Validate-FIRST design: jq contract check (version 1, interface wg0, 44-b64 pubkeys, 10.77.0.x/32 allowed_ips, never the endpoint's own .1) rejects on stderr with exit 1 before touching anything; then head-file + generated [Peer] blocks into a same-fs tmp, wg syncconf <(wg-quick strip …) from the TMP (exact-match: adds/removes without bouncing the interface), and only on success the atomic mv to /etc/wireguard/wg0.conf — runtime and boot config can never diverge in the failure direction. Zero-peer payload = valid wipe. Never reads or prints the private key; no wg-quick save; no second mode. shellcheck-clean. Live-proven on felhom-hetzner incl. the negatives (malformed JSON / bad pubkey / own-IP peer → exit 1, wg state byte-identical) and reboot persistence.

docs — architecture Part 06: offsite connectivity design-of-record (2026-07-03)

documentation/architecture/06-offsite-connectivity.md — the settled offsite-backup-transport design, authored from the spike verdict + operator-resolved forks (recorded, not re-litigated): plain WG (D1), host-side agent-managed wg-felhom as the agent-managed-unit pilot on the sudoers *.mount install pattern (D2), one shared hub-driven endpoint VM running WG + the offsite PBS with no agent (D3, CF-token pattern), hub source-of-truth with a wireguard block riding the existing WireDesiredState/DesiredGeneration channel (D4), one datastore + per-customer namespaces (D5), and PBS on the VM — relay-through-DooPlex rejected as non-scaling (D6). Includes the Day-0 join handshake, the robustness set (NOT-DynDNS roaming, endpoint DNS re-resolve watchdog, MTU 1420, per-/32 topological isolation, tunnel-health through the storage-target reachability model), trust conformance, the honest open ledger (CGNAT unmeasured → mobile-hotspot smoke test; peer-sync push-vs-pull = slice-1 design point), and the S1S6 slice roadmap (MVP = S1→S2→S3, then S4; S5 merges with DR-completeness). All claims cited at file:line against felhom.eu @ bf099f6 + felhom-agent @ 4ba1b14. day0-install.md backlog line now points at spike + design doc. Docs-only.

docs — SPIKE: offsite-backup connectivity — plain WireGuard WINS the ladder; Headscale = separable fleet layer (2026-07-03)

documentation/audits/SPIKE-connectivity-wireguard-2026-07-03.md — the offsite-backup transport decision, empirically grounded on both real ends (demo-felhom PVE host ⟷ throwaway Hetzner box, end state: powered off, secrets shredded). Headline results: the operator's line is plain-NAT with a fixed public IP, not CGNAT, and has zero IPv6 (P0 honesty — CGNAT confirmation deferred to Peti's VM 110); plain outbound WG held an 11.4-min idle window and carried a real 2 GiB worst-case PBS backup at 4.26 MiB/s = the full home uplink (~5% tunnel overhead), TLS pin intact through the tunnel (positive + negative proof); UDP 51820 and 443 both pass; kernel WG surprisingly works inside the unprivileged guest (P7 — host placement stands on architecture, not infeasibility). Recommendation: host-side agent-managed WG (cloudflared pattern), key in the escrowed IdentityBundle, small public endpoint VM (PBS-on-VM vs rendezvous-relay deferred to the spec). runbooks/day0-install.md backlog line resolved to point here; CONTEXT.md notes the DR-completeness task is unblocked (next: the production connectivity spec).

skills — NEW: felhom-app-catalog (4th skill) + SparkyFitness as its worked example (2026-07-03)

skills/felhom-app-catalog/SKILL.md — the catalog authoring workflow (research → inspect the image for the healthcheck family → write compose/.felhom.yml → deploy live through the dashboard → verify healthy → reconcile the app count). Deliberately points at app-catalog REUSE.md §12 + README.md §format for every field table (one-fact-one-place; no duplication). Unique content: the never-guess-the-healthcheck rule with the per-tool image-inspection loop (BusyBox ash command -v gotcha: it silently ignores all but its first argument — verified), the probe-container naming rule (controller probes the container named exactly like the stack — verified in felhom-controller/internal/stacks/healthprobe.go, row added to app-catalog REUSE.md), the Hungarian-quote YAML kill, and the deploy-is-the-test doctrine. No installer change needed — install_skills.py auto-discovers skills/*/SKILL.md; fresh-session discovery probe listed all 4. Proven by finalizing sparkyfitness end-to-end on demo (see app-catalog-felhom.eu CHANGELOG).

docs — golden 0.98.3 live: D.1b RETIRED, drill ledger B1/B5 → FIXED, backlog note resolved (2026-07-03)

Companion to felhom-agent's build-golden.sh v2.0.0 (@ ceca355): the golden now bakes the CURRENT controller (0.98.3, mandatory-tag convention — B5) and a felhom-controller-bootstrap.path unit (controller deploys on the bootstrap-mount hot-plug, no reboot — B1). Clean-room validated end-to-end BEFORE publish (bake → isolated hot-plug proof → local-golden Day-0 install → publish → vouch → --force-gitea-golden install); evidence: documentation/audits/DRILL-golden-098-2026-07-03.md.

  • documentation/runbooks/day0-install.md: D.1b reduced to a one-line version check (fresh boxes land current + self-manage); old manual-update procedure moved to Part F troubleshooting keyed on "golden older than 0.86.0 was vouched"; header validated-versions line → script v1.9.1 / agent v0.63.0 / golden v0.98.3; A.3 records the drilled-known-good pair + "vouch ≥ 0.98.3"; A.4 floor text rewritten (fresh boxes now floor-honoring; recommend raising the global floor after a golden rebuild — operator, 1 min).
  • documentation/audits/DRILL-day0-cleanroom-2026-07-03.md ledger: B1, B5 → FIXED (pointers); R6 note: the installer's post-provision reboot is retained as a belt, removal recorded as a candidate cleanup (not done).
  • documentation/backlog/FOLLOWUP-golden-default-controller-tag.md + backlog/README.md: RESOLVED per the M18/M19 convention (file kept + annotated; README entry marked FIXED).
  • New evidence doc: documentation/audits/DRILL-golden-098-2026-07-03.md (AD transcripts, unit states, resolution-order + fetch/sha proofs, cleanup, observations — incl. the SECURITY observation that the customer git.token has package-WRITE rights → scope-down + rotate follow-up).

docs — day0-install D.1b narrowed + drill ledger B2/B3 → FIXED agent v0.63.0 (2026-07-03)

Agent v0.63.0 fixed both drill findings (B3 token reload-on-miss + B2 guesthook snippets-dir mkdir; red-proofed, deployed on felhom-pve, Gitea-published sha256 b4a89c81…). Guide follow-through: the D.1b "restart the agent first" step is now CONDITIONAL (only for an installed agent < v0.63.0 — the Day-0 manifest still vouches 0.62.0, so today's fresh installs still hit it); the 401 troubleshooting row records the fix version; the drill ledger + go/no-go item 8 marked FIXED. Operator follow-up unchanged: vouch agent 0.63.0 in the Day-0 manifest UI, then the step is dead.

felhom-host-install.sh v1.9.1 — clean-room drill fixes: residue-free uninstall + post-provision reboot (2026-07-03)

Companion to the Day-0 go-live package (documentation/runbooks/day0-install.md + documentation/audits/DRILL-day0-cleanroom-2026-07-03.md). Every fix was found by the clean-room drill (virgin nested PVE 9.2.2) and re-verified there (v1.9.1 uninstall → zero-Felhom-residue diff vs the pre-install baseline; v1.9.1 install → controller up with no manual intervention).

  • Header/version sync (the header said v1.8.0 while SCRIPT_VERSION said 1.9.0); keep-in-sync note on SCRIPT_VERSION; usage sed range follows the header (2,95).
  • Uninstall now removes the drill-found residue (R1R5): the agent config (resolved from the unit's -config BEFORE the unit is removed — it holds the per-host hub api_key), the felhom-shared-parent unit + wants links + /usr/local/sbin/felhom-shared-parent.sh
    • the /mnt/felhom-drives self-bind/dir, /usr/local/sbin/felhom-mkfs-guarded, /var/lib/vz/snippets/felhom-guest-hook.sh, and /etc/dnsmasq.d/felhom-*.conf (+ dnsmasq restart when touched). All tolerate-absent; summary lines updated (sudo AND dnsmasq packages are the documented package remnants).
  • Post-provision guest reboot (R6): the golden's felhom-controller-bootstrap.service evaluates ConditionPathExists=/etc/felhom-bootstrap/bootstrap.json at BOOT, but the agent back-half hot-plugs the mount into the running guest — on slower hardware the first boot loses that race deterministically and the controller never deploys. step_provision now reboots the guest once (the agent's own output says "next: reboot the guest"); step_verify waits bounded (180 s) for the controller container instead of a momentary look.

felhom-host-install.sh v1.9.0 — Pool.Audit for the stale-lock reaper (A1) (2026-07-03)

Companion to felhom-agent v0.62.0 (audit A1: pool-membership ownership check). PVE_PRIVS_GUEST gains Pool.Audit (12 → 13 privs, granted at /pool/felhom via the existing FelhomAgentGuest role) so the agent can read GET /pools/felhom — its stale-lock reaper's ownership registry. Pool.Allocate does NOT satisfy the read (spike SPIKE-a1-pool-membership-read-2026-07-03 T2). No structural change: _ensure_role already role modifys to the exact priv set, so re-running --rescope-acl (or a fresh install) upgrades an existing box idempotently; remove_scoped_acl deletes by role name and needs nothing. Deploy order on a live box: rescope FIRST, then deploy agent v0.62.0 — the added read priv is harmless to an older agent, while the new agent on an old ACL fail-safes its reaper (skips) and reports pve:pool-read degraded until the rescope lands.

docs — SPIKE: A1 pool-membership read for the stale-lock reaper (2026-07-03)

Findings doc documentation/audits/SPIKE-a1-pool-membership-read-2026-07-03.md. Live-probed on felhom-pve under the PRODUCTION scoped token vs root: LXC enumeration IS already pool-filtered (token sees only 9201 of 4 guests); GET /pools/felhom 403s naming Pool.Audit; a throwaway token with ONLY Pool.Audit@/pool/felhom reads members (minimal delta proven, fully torn down); /cluster/resources withholds the pool field without Pool.Audit; local ownership records are all partial. Recommendation for the A1 impl spec: add Pool.Audit to PVE_PRIVS_GUEST in felhom-host-install.sh (L183) + a GET /pools/felhom cross-check in the agent's staleLockController.Guests(), fail-safe skip on read failure. No script/agent change in this commit — docs only. Appendix: committed-secrets (felhom.secret.yaml) rotation micro-runbook, operator follow-up.

install_skills.py — new: Claude Code skills installer (2026-07-03)

Installs skills/*/SKILL.md (felhom-build-deploy, felhom-ui-design, felhom-testing) into ~/.claude/skills/ as Windows junctions (mklink /J) so repo edits are live immediately; falls back to a full copy if junction creation fails or isn't followed (copy mode prints a re-run reminder). Idempotent — re-runs detect a correct junction and leave it. Verified: junctions ARE followed by Claude Code skill discovery (fresh-session probe found all three).

reuse_refs_check.py — new gate: REUSE.md citation checker (2026-07-03)

Staleness defense for the new per-repo REUSE.md reuse maps. Takes repo roots as argv, extracts every cited *.go/*.py/*.html/*.css/*.yml/*.yaml/*.sh path (slash-containing tokens only — bare filenames are conventions, not citations), verifies each exists; prints offenders, non-zero exit on any missing path. Symbols are spot-verified by the reviewer, not this script. Usage: python scripts/reuse_refs_check.py <repo-root> [...].

felhom-host-install.sh v1.8.0 — install the guarded-mkfs wrapper (Impl-1 Part B) (2026-07-01)

Companion to felhom-agent v0.54.0 (format-safety foundation). During agent install, fetch + install the guarded-mkfs wrapper so the agent's format path is safe on any box.

  • New step in step_agent_install: fetch configs/felhom-mkfs-guarded.sh from Gitea, bash -n validate, install -m0755 -o root -g root/usr/local/sbin/felhom-mkfs-guarded. Installed BEFORE the sudoers (which now allowlists ONLY the wrapper, not raw mkfs.*), so the ordering is gap-free.
  • The agent v0.54.0 sudoers (fetched by the same step) drops the raw mkfs.ext4 -F /dev/* / mkfs.xfs -f /dev/* allowlist and permits only felhom-mkfs-guarded /dev/* *, plus read-only pvs/zpool for the agent's unclaimed-disk guard. No other host-install change.
  • bash -n + shellcheck clean (0 new warnings). Live-validated on felhom-pve (agent v0.54.0 deploy): wrapper refuses the OS disk + an LVM-PV partition, raw mkfs is sudo-denied, an unclaimed throwaway disk formats; the agent guard's sudo reads (pvs/lsblk/zpool) all work as the felhom-agent user.

felhom-host-install.sh v1.7.0 — 3b-fix: Datastore.Audit box-wide (restore drive visibility) (2026-07-01)

Fixes a regression the v1.6.0 pool-scoped ACL introduced: Datastore.Audit was placed in the per-storage Store role (granted only on local/local-lvm/felhom-pbs), which excluded the enrolled removable drives felhom-usb/felhom-flash. The agent enumerates storage via ListStorage/NodeStorage (both gated by Datastore.Auditinternal/storage/observe.go), so it could no longer SEE the drives → false "Meghajtó leválasztva" (drive detached) alerts + drives absent from the agent-view. (The v1.6.0 swap's "felhom-usb → 403" was mis-read as blast-radius success; felhom-usb is Felhom's OWN customer drive, not an out-of-scope object.)

  • Datastore.Audit moved from Store → Base (PVE_PRIVS_BASE now "Sys.Audit SDN.Use Datastore.Audit"; PVE_PRIVS_STORE now "Datastore.Allocate Datastore.AllocateSpace"). Audit is read-only metadata, so box-wide Audit restores visibility of ALL storages (incl. dynamically-enrolled drives — no per-drive grant ever needed) while the write privs (Allocate/AllocateSpace) stay per-storage → write/allocate blast-radius containment is UNCHANGED. Confirmed at source: the agent creates no PVE storage (no POST /storage/pvesm add); drives are dir-storages it observes + mounts via host ops, so they need only Audit, never Allocate.
  • apply_scoped_acl reordered Base-before-Store (role + grant) so a RE-APPLY on a live box adds Audit@/ before Store drops its per-storage Audit → gap-free (the agent never loses enumeration).
  • remove_scoped_acl / --uninstall / --rescope-acl operate by role NAME and inherit the corrected privs automatically (no other change).
  • Live-repaired felhom-pve (two pveum role modify, Base first — no agent stop/restart): drives reappeared (agent-view 3→5 storages), detach alerts cleared. Re-tested under the scoped token: drives readable (was 403), write-containment intact (vzdump→felhom-usb still 403; out-of-pool guest 403), PBS Store grant unchanged. bash -n + shellcheck clean (0 new warnings).
  • NOT physically run (source-confirmed, no Datastore.Allocate in the path): a brand-new-drive UI enrollment (needs a spare USB) — the host-ops/Audit path is unchanged from pre-3b.

felhom-host-install.sh v1.6.0 — pool-scoped token ACL (3-role) + --rescope-acl retrofit (2026-07-01)

Colleague-safety batch #4 phase b (script half; agent half = v0.53.0). Moves the agent token's dangerous privileges off / (which spanned every guest + storage) to /pool/felhom + /storage/<targets>, so on a shared box the token can only touch Felhom's own guests + storages. Validated by documentation/audits/SPIKE-pool-scoped-acl-2026-07-01.md (PASS) — implemented here.

  • 3-role scoped ACL (step_token rewrite). Replaces the single FelhomAgent role granted at / with three roles, each granted to BOTH the user AND the token (privsep intersection): FelhomAgentGuest (VM.* + Pool.Allocate) @ /pool/felhom; FelhomAgentStore (Datastore.*) @ each of PVE_STORAGES (default local local-lvm felhom-pbs — the offsite PBS MUST be included, SPIKE residual #1; --acl-storages overrides); FelhomAgentBase (Sys.Audit SDN.Use) @ /. Helpers apply_scoped_acl/remove_scoped_acl/_grant/_ensure_role.
  • Pool before token. ensure_felhom_pool runs at the top of step_token (always, incl. --skip-provision) so /pool/felhom exists before it's granted on.
  • Re-install safety. step_token also removes the pre-3b broad / grant + FelhomAgent role if present (remove_old_broad_acl, tolerate-absent), so a re-install can't leave the old grant unioned with the scoped one. The post-provision pool_add_guest is gone (the agent's restore --pool makes the guest a member atomically — v0.53.0).
  • --rescope-acl retrofit (new mode, mirrors --adopt-pool): migrate an existing install — ensure the pool + guest membership, apply the scoped grants, THEN remove the old broad grant (add-before- remove: the token is never grant-less mid-migration). Prints the "now deploy agent ≥ v0.53.0" ordering reminder. Idempotent + dry-run-aware. SUPERVISED (run with the agent stopped — the scoped ACL and the pool-param agent are mutually dependent; §13 of the task).
  • --uninstall now removes the scoped grants + 3 roles AND the pre-3b broad grant/role (both tolerate-absent → works on either shape), keeping the pool delete-if-empty (v1.5.0).
  • Validated on felhom-pve (dry-run): T-A fresh install (pool-before-token, 3 roles once, scoped grants incl. /storage/felhom-pbs), --rescope-acl (add scoped → remove old FelhomAgent), T-F uninstall (old-shape cleanup + pool not-empty skip). bash -n + shellcheck clean (0 new warnings). The live rescope + agent swap is the supervised STOP — not run here.

felhom-host-install.sh v1.5.0 — felhom pool by default + --adopt-pool retrofit + uninstall teardown (2026-07-01)

Colleague-safety batch #4 phase a. Every Felhom-managed guest now joins a dedicated felhom pool for fleet uniformity (and as the environment the later pool-scoped ACL — 3b — will spike against). All pool ops run as root@pam from the installer, so there is NO agent/token/ACL change and zero permission-model risk (PVE_PRIVS untouched; the FelhomAgent token stays scoped at /).

  • New felhom pool default. step_provision calls ensure_felhom_pool (create if absent, idempotent) and, after a successful provision, adds the guest via pveum pool modify felhom -vms <vmid> (skip-if-already-member). New helpers pool_exists / pool_members / ensure_felhom_pool / pool_add_guest; const PVE_POOL="felhom". PVE 9 syntax + /pools JSON shape confirmed live before wiring (pveum pool add|delete|modify; pvesh get /pools[{poolid,comment}], /pools/<id>{members:[{vmid,…}]}).
  • --adopt-pool retrofit mode. Non-destructive: adds an EXISTING Felhom guest to the pool (creating it if needed), resolving the guest from --vmid else the recorded provisioned_vmid. Reuses the ours-check (/etc/felhom-bootstrap mount) — refuses a non-Felhom guest unless --force. Touches ONLY pool membership: never reconfigures/restarts the guest, never contacts the hub. Idempotent (skip-if-member).
  • --uninstall pool teardown (step 5b). After the pveum removal, deletes the felhom pool only if empty (a destroyed guest is auto-removed from its pool); a pool that still has members is left with a log_skip naming them. Not reached on the Spec-1 safe-skip path (other Felhom guests remain).
  • Validated on felhom-pve (dry-run + SAFE live): T-A fresh-install dry-run shows the pool create + membership lines; T-B live adopt of guest 9201pvesh get /pools/felhom lists 9201, guest still running, config unchanged (the demo node is now pool-uniform); re-run = no-op; T-B' non-Felhom vmid → refusal; T-C uninstall dry-run → "pool felhom not empty (members: 9201) — leaving it". bash -n + shellcheck clean (0 new warnings; the 2 pre-existing SC2015 in step_verify unchanged).
  • NOT changed: PVE_PRIVS, the ACL grants, the agent, the provision-call args. 3b (pool-scoped ACL
    • agent restore-into-pool under a scoped token) is the separate spike-gated task.

felhom-host-install.sh v1.4.0 — appliance CPU/RAM cap passthrough (--cores / --memory) (2026-07-01)

Colleague-safety batch #3 (host-install half; the mechanism is agent v0.52.0). Lets an operator cap the provisioned guest so a trial appliance on a SHARED production Proxmox doesn't pressure the colleague's existing guests.

  • --cores N / --memory M (MiB) — optional; passed through to the agent's --selftest=provision as -cores/-memory. 0/unset = keep the golden's baked sizes (unchanged behaviour). New vars CPU_CORES/MEM_MIB; usage() header gains an "Appliance cap (optional)" group.
  • Conditional passthroughstep_provision builds a cap_args array and appends the flags to BOTH the dry-run log and the real agent call only when set. An agent < v0.52.0 would reject an unknown flag, so the flags are never sent unless the operator opts in (see the deploy dependency below).
  • Pre-flight sanity WARN (soft, provision only) — if --cores > host nproc or --memory > host MemTotal, log_warn "the cap won't protect other guests"; never die (the operator may know better).
  • Deploy dependency: a fresh install using --cores/--memory needs the hub artifact manifest to serve agent ≥ v0.52.0.
  • Validated dry-run on felhom-pve: --cores 2 --memory 4096 --dry-run → provision command shows -cores 2 -memory 4096; without the flags → neither present; --cores 64 --memory 65536 → both WARN lines (host 4 cores / ~15771 MiB). bash -n + shellcheck clean (0 new warnings; the 2 pre-existing SC2015 in step_verify unchanged).

felhom-host-install.sh v1.3.0 — --uninstall (clean revert) + pre-flight guards (2026-07-01)

Colleague-safety batch #1+#2. Adds a first-class, guarded --uninstall teardown so an operator can cleanly back out of a trial install, plus three provision pre-flight guards that stop common footguns. Script-only; no agent/hub/controller change.

  • --uninstall (local host teardown — no hub contact, no passphrase). Reverses an install in the install-order's reverse: guest → agent(unit/sudoers/binary/state/user) → pveum(ACL,token,user,role) → golden(opt-in) → state file. Every mutation goes through run() so --dry-run prints the full plan and executes nothing. Safety:
    • Ours-check: refuses to destroy a guest that lacks the /etc/felhom-bootstrap bind mount (matched by the constant guest path, not a hardcoded mpN slot — on the demo host it's mp9), unless --force.
    • Typed confirmation: must type the vmid to confirm PERMANENT destruction (read from /dev/tty; skipped only under --dry-run, where nothing is destroyed).
    • Other-guests guard: if any OTHER Felhom guest remains, destroys only the target and leaves the agent + PVE token + state in place (re-run with --force to remove host-level anyway — orphans the others).
    • Never removes the sudo package; never contacts the hub (the host record intentionally persists).
    • Presence-checked + idempotent: an already-absent guest/unit/sudoers/binary/user/ACL/token/role is a tolerated skip, not an error. The pveum role delete runs only after its ACL grants are gone (PVE refuses to delete a referenced role). Confirmed PVE 9 ACL-delete form: pveum acl delete / --users|--tokens <x> --roles FelhomAgent.
    • Target vmid resolves from --vmid, else the recorded provisioned_vmid (else dies). A --vmid that disagrees with the recorded one needs --force.
    • --remove-golden: with --uninstall, also delete the golden vzdump from the archive storage (pvesm free); otherwise it is left in place.
  • Install state now records customer_id + provisioned_vmid (new _state_put/_state_get helpers, dry-run-guarded like _state_mark; the completed[] shape is untouched) so a later --uninstall resolves its target automatically and safely.
  • Pre-flight guards (provision mode):
    • Multi-node guard — on a 2+-node cluster, die (naming the nodes) unless --node is explicit (new NODE_EXPLICIT); single-node keeps the current auto-pick. No-op under --skip-provision.
    • Archive-storage-exists guard — verify --archive-storage appears in pvesm status (else die); no-op under --skip-provision.
    • RAM floor (WARN, never fatal) — warn when MemAvailable < 2048 MiB. All three run inside step_preflight (before any mutation) so they also fire under --dry-run.
  • Validated dry-run-only on felhom-pve (single-node, live guest 9201): T-A full uninstall plan, T-C not-ours refusal (red-proof), archive-missing die, RAM line, other-guests detector, state round-trip; confirmed 9201 + agent + pveum + state untouched after all dry-runs. bash -n + shellcheck clean (0 new warnings vs. baseline; the 2 pre-existing SC2015 in step_verify are unchanged). NOT yet live-validated (awaiting a supervised run): a real live --uninstall (guest destroy + pveum removal) and the multi-node guard on an actual cluster.

felhom-host-install.sh v1.2.0 — /dev/tty passphrase read + vmid auto-detect (2026-07-01)

Two operator-experience fixes so a colleague can install online (via the hub's new "Option 1: Online install" one-liner) and onto a host that already runs a guest at 9201.

  • Passphrase prompt reads from /dev/tty, not stdin (read_passphrase). read -rsp … < /dev/tty makes the no-echo prompt work regardless of how stdin is wired — both download-then-run and curl … | sudo bash (where stdin is the pipe). Strictly more correct; the --passphrase-file path is unchanged. The passphrase is still never on argv / in logs / in the state file.
  • VMID auto-detect (--vmid now optional-smart). New VMID_EXPLICIT flag (set by --vmid). The pre-flight vmid guard now determines "in use" against the pct list + qm list id-set (LXC and VMs share the id space — more complete than the old pct status, which only knew LXC):
    • explicit --vmid → unchanged deterministic behavior: die if the id is in use unless --force (destructive over-provision).
    • default 9201, in use, no --forceauto-pick the next free id (scan upward from 9201 over the used-set) and ask to confirm from the terminal (read … < /dev/tty, [y/N]); proceed on yes, die "no free vmid confirmed" otherwise. Never a silent auto-pick.
    • default 9201 + --force → over-provision 9201 (destructive) without prompting, as before.
    • New helpers used_vmids / _vmid_in_use / next_free_vmid. --vmid help text + usage() updated.

felhom-host-install.sh v1.1.0 — self-install the agent + fetch the golden from Gitea (2026-06-28)

The script now installs the agent itself (the last big manual Day-0 prerequisite is gone). It fetches the agent binary + golden from Gitea generic packages and verifies each against the hub-vouched artifact manifest before installing/using it. BUNDLE slice; pairs with hub v0.16.0 (artifact manifest endpoint + operator UI) and felhom-agent v0.43.0 (canonical unit + publish).

  • New step 5/8 agent install (before agent-config): resolves the manifest (GET /api/v1/artifacts/{id}, passphrase) + the git fetch token (from the customer's controller.yaml via config-retrieve — NO new credential); fetches /api/packages/admin/generic/felhom-agent/<ver>/felhom-agent, verifies sha256 vs the hub manifest (aborts on mismatch — verify-before-use), backs up any existing binary, installs 0755 /usr/local/bin/felhom-agent; ensures the non-root felhom-agent system user; installs the canonical sudoers (0440, visudo -cf-validated) + systemd unit; daemon-reload + enable. Idempotent: same version already installed + service active → skip.
  • --skip-provision: install + configure + verify the agent (incl. golden fetch+verify) but do NOT provision a guest — the agent-only path for re-installing/upgrading the agent on a host that already has live guests. Adds an agent-only step_verify_agent (binary + non-root service active + a --selftest=hub collect-report).
  • New step 7/8 golden: local auto-discovery stays the default/fallback; otherwise fetches /api/packages/admin/generic/felhom-golden/<ver>/golden.tar.zst, verifies sha256, and imports it into the archive storage's dump dir for the restore. --force-gitea-golden forces the Gitea path.
  • Non-root agent model: the agent now runs as felhom-agent with privileged.mode: "sudo" (was the dev/CI direct+root shortcut). The config is chowned to the service user (0600) so the daemon can read it; systemctl is-active after restart is the real proof the non-root user can read the config.
  • Pre-flight relaxed: a missing agent binary is no longer fatal (step 5 installs it); the local golden requirement is deferred to step 7.
  • Trust model: checksum trust root = the hub (manifest), not Gitea; the fetch credential is the existing config-retrieve git token; artifacts are pinned to a version (never :latest).
  • Secrets: the git token is a never-logged runtime carrier (cleared on EXIT alongside the passphrase / pve-token / hub api_key); the sudoers is 0440 and visudo -cf-validated before install.
  • bash -n + shellcheck clean.

felhom-host-install.sh v1.0.0 — Day-0 host bootstrap (provision mode) (2026-06-26)

First release. A single operator-run script that automates Day-0 on a freshly-PVE-installed host: Proxmox API token → hub host enrollment (option C, single secret) → agent config → guest provision → verify. Composes proven mechanisms (the pveum role/token sequence, hub POST /host-enroll, felhom-agent --selftest=provision); grounded by documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md.

  • 7 steps, idempotent + resumable via /var/lib/felhom-install/state.json: pre-flight → Proxmox token → compute grows → host-enroll → agent config → provision → verify.
  • Single-secret (the retrieval passphrase): read no-echo or from a 0600 file, never on argv/logs/state. The global operator key never touches the box.
  • pveum automation: 16-priv FelhomAgent role (create-or-modify), felhom-agent@pve user, privsep token (reuse-if-working else rotate), and both ACL grants applied after the token exists (token-remove purges the token ACL).
  • Auto-discovery: golden archive (newest vzdump-lxc-<golden-vmid>), PVE node name, vmbr0 bridge IP for the local-api, and the served-leaf TLS fingerprint pin.
  • Safety: pre-flight fails fast (root, PVE 9.x, local-lvm headroom, hub reachable, customer+passphrase valid via read-only GET /config/{id}, golden resolvable); refuses to clobber an existing --vmid without --force; --dry-run previews every mutation; --preserve-from keeps operator infra (PBS/local_api/privileged/authz) on re-deploys.
  • --mode dr: documented 10D stub (restore customer PBS snapshot instead of golden) — not implemented.
  • Live-validated end-to-end on felhom-pve: authorized wipe of demo guest 9201 → re-provision from the golden → controller config-pull + public tunnel HTTP 200 → host-report of guest 9201 → idempotent --resume no-op. (One ordering bug — token ACL applied before rotation — was found and fixed during the live run.)