Files
felhom.eu/scripts/felhom-host-install.sh
T
admin d319ae573e hub: delete the host-install version label (R-94) + invert hostinstall gate 1
The Setup tab said 'host-install 1.19.0' while the served script was 1.22.0, and had
been wrong since 2026-07-14. Deriving the number honestly is not possible: the Option-1
command downloads felhom-host-install.sh from the website at RUN TIME and the website
git-syncs main every 30s (R-110), so no build-time value in the hub can be true. R-94(a)
offered derive-or-delete; deleted, which removes the drift class instead of automating it.

- configs.go: hostInstallVersion const, pageData.ScriptVersion field and its assignment
  all removed; a NOTE in their place records why there is no constant here.
- customer_unified.html: the sentence now says the command always fetches the current
  installer, and renders no version.
- hostinstall_gates.py gate 1: the third assertion INVERTS — it used to require the hub
  const to equal SCRIPT_VERSION, it now asserts the hub carries no host-install version
  literal at all, matched in six code shapes across every .go/.html under hub/ (comments
  are deliberately not stripped: a // inside a URL literal would blind the scan).
- render_test.go: the assertion 'html contains hostInstallVersion' compared the constant
  to itself and passed at ANY value — demonstrated green with the const at 9.9.9 while the
  script was 1.22.0. Deleted, not replaced: there is no longer a version to assert.
- felhom-host-install.sh: COMMENT ONLY (SCRIPT_VERSION untouched) — it claimed the gate
  keeps the hub copy equal, an invariant that no longer exists.

Red-proofs: restoring the const fails the rewritten gate 1 (3 shapes hit); the old
render_test assertion passes at 9.9.9.
2026-08-02 15:16:01 +02:00

2805 lines
170 KiB
Bash

#!/bin/bash
#===============================================================================
# felhom-host-install.sh — version: the SCRIPT_VERSION variable below (single source;
# `-h` and the run banner both read it — the v1.14.0/v1.13.0 header/banner drift was drill F-1).
# Day-0 host-bootstrap for a Felhom Proxmox host (operator-deploy model).
#
# Run by the operator on a FRESHLY-PVE-INSTALLED box (after a manual PVE install
# + SSH in). Given a customer-id + retrieval passphrase, it fully automates
# Day-0: Proxmox API token -> hub host enrollment -> AGENT INSTALL (fetch from
# Gitea + verify sha256 + install) -> agent config -> golden -> guest provision
# -> verify. It composes already-proven mechanisms (the pveum role/token
# sequence, hub POST /host-enroll [option C], felhom-agent --selftest=provision).
# The agent renders bootstrap.json and the controller pulls its own
# controller.yaml in-guest; this script does NOT fetch that.
#
# v1.1.0 (BUNDLE slice): the agent binary + golden are now fetched from Gitea
# generic packages and VERIFIED against the hub-vouched artifact manifest
# (GET /api/v1/artifacts/{id}) before install/use. The fetch credential is the
# git token already inside the customer's controller.yaml (config-retrieve) — NO
# new credential. The checksum trust root is the HUB, not Gitea. This removes the
# old prerequisite "install the agent binary + unit manually".
#
# v1.15.0 (DR-tier-by-default, DRILL-day0-vm-2026-07-12 sweep): every install now deploys the FULL
# DR plumbing uniformly — capability is baked, activation is a hub flag. F-7: ships the
# `felhom-pbs-apply` wrapper (like mkfs/selfupdate — pbsdr capabilities no longer born DEGRADED).
# F-10: installs `age` (escrow-ceremony identity-wrap dependency). F-9: renders
# `wg_tunnel.enabled: true` (WG is BASE infrastructure like cloudflared, not a DR component —
# operator decision 2026-07-12; byo consent moves to the disclosure ack + tester agreement, the
# byo wg-off config assert is retired). The default --acl-storages set keeps `felhom-pbs`
# (drill-proven: narrowing it 403s the PBS-DR apply-bridge later); a missing felhom-pbs storage
# is expected pre-tier and no longer a byo refusal. F-1: single VERSION source (SCRIPT_VERSION).
# F-2: dry-run fetch lines show the real auth mode (anonymous vs -u). F-8: step 4b + the final
# summary now say OUT LOUD that root@pam was rotated + where to retrieve it.
# v1.14.0 (CAMPAIGN-3 Part 6): --mode now also templates the agent's top-level `deployment_mode`
# ("appliance" | "byo") — authoritative from --mode, written even over a --preserve-from config, so a
# byo box can never inherit an appliance host-service self-heal. Gates the agent's node self-heal
# (host networking recovery at boot, F12-class; agent v0.85.0). The byo config assert now also refuses
# deployment_mode=appliance. absent/typo → byo (fail-safe).
# v1.10.0 (GL-2, go-live G2/G4/G5): explicit --mode appliance|byo install profile — the flag is now
# REQUIRED for a fresh install (no default; build-golden v2.0.0 precedent: defaults rot). byo =
# BYO-host hardening for a Proxmox host the operator does NOT own: break-glass (step 4b) gated OFF
# at its call site (root@pam is never touched), mandatory --cores/--memory noisy-neighbor caps,
# strict preflight (PVE-major gate, --acl-storages existence, host-DNS :53 info line, host-mutation
# disclosure + typed-hostname acknowledgement), byo config asserts (lan_resolver/wg_tunnel/oob must
# stay off), pool+ACL verify asserts (BOTH modes — campaign-2 R2 lesson), and --preflight-only.
# Test harness: scripts/hostinstall-mode-harness.sh (static tier runs anywhere; PVE tier dry-only).
#
# v1.11.0 (GL-4, go-live G6 + the G1 key-pin fold): (A) uninstall gap-closure — the teardown now
# also removes the self-update artifacts (guarded wrapper, .prev/.new.* A/B slots, rollback unit,
# start-limit drop-in), unmounts every enrolled drive under /mnt/felhom-drives (plain umount only,
# NEVER -l/-f; busy = warn + guidance, data always stays on the drive), and ends with an explicit
# KEPT-vs-WIPED statement (PBS backups, hub record, escrow, vaulted recovery credential live on).
# (B) operator-key day-0 fold — OPERATOR_KEY_* constants (empty until the pin ceremony) and
# --operator-pubkey-file write authz.signers into the agent config at step 6; empty keys keep a
# preserved config's signers (never clobber a manually-pinned box); verify reports armed/dormant.
#
# v1.11.2 (GL-6 Gate-0 operator ruling, G3): ANONYMOUS artifact fetch supported — empty
# git.username/git.token in the customer config now WARNS and fetches anonymously (the Felhom
# artifacts are world-readable by ruling; the hub-vouched sha256 stays the integrity root).
# A configured credential is still used when present.
#
# v1.11.3 (GL-6 finding F4): --resume correctness — the producer steps (token/enroll/grows) now
# run every pass so their in-memory outputs (pve token, hub host_id/api_key, volume grows) are
# repopulated for the later steps that consume them; a resumed install no longer writes a config
# missing hub.host_id/proxmox.token (daemon crash-loop) or passes an empty -rootfs-grow. golden's
# GOLDEN_VOLID is re-derived from the local archive on resume.
#
# v1.12.0 (GL-8, BYO coexistence hardening — GL-6 findings): F3 — the populated-host leaf guard
# now fires only when a FELHOM guest exists (felhom_guests), so a fresh byo install on a host that
# only runs the OWNER's guests proceeds without --allow-new-leaf (the pin protection for real Felhom
# guests is unchanged). F6 — byo preflight REFUSES (never mutates) when a foreign resolver is bound
# to :53, with remediation (Felhom needs the guest reachable by name; it won't stop the owner's DNS).
# F1 — uninstall now purges the agent config's .bak* siblings (one held a live hub api_key), not just
# agent.json.
#
# Grounding: documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md
#
# Usage:
# sudo ./felhom-host-install.sh --customer-id ID --mode appliance|byo [options]
#
# Required:
# --customer-id ID Customer (must already exist in the hub)
# --mode appliance|byo Install profile (REQUIRED for install/--resume — no default):
# appliance a Felhom-owned/managed box. Full Day-0 incl. the
# break-glass root@pam credential (step 4b).
# byo a host the operator does NOT own (BYO / pilot). NEVER
# touches root@pam (step 4b skipped), REQUIRES explicit
# --cores AND --memory, refuses --enable-oob and
# --rotate-recovery, dies on non-9.x PVE (unless --force)
# and on --acl-storages entries absent from this box
# (except felhom-pbs — the PBS-DR tier creates it later),
# prints the full host-mutation list and requires typing
# the host's short hostname, and asserts the agent config
# keeps lan_resolver/oob OFF (wg_tunnel is ON in BOTH
# modes since v1.15.0 — base infra, disclosed in the ack).
# (dr the 10D DR stub — documented seam, not implemented.)
#
# Options:
# --preflight-only run ONLY the (mode-aware) step-1 checks: prints PRE-FLIGHT PASS/FAIL,
# exits 0/1, writes NO state, executes NO later step. Re-runnable; valid
# with either mode (mode-specific checks need --mode).
# --hub-url URL default https://hub.felhom.eu
# --vmid N guest VMID to provision. Default 9201; if omitted and 9201 is already
# in use, the script auto-picks the next free id (pct+qm) and asks to
# confirm. An EXPLICIT --vmid stays deterministic (dies unless --force).
# --golden VOLID golden archive volid (default: newest vzdump of the
# golden build VMID on the archive storage; else fetched
# from Gitea per the hub artifact manifest)
# --golden-vmid N golden build guest vmid for auto-discovery (default 9100)
# --archive-storage NAME storage holding the golden vzdump (default local)
# --force-gitea-golden ignore any local golden; fetch+verify the golden from
# Gitea (proves the fetch path; used by the live test)
# --node NAME PVE node name (default: pvesh /nodes, else hostname)
# --bridge-ip IP[:PORT] local-api listen addr; an explicit value overrides the island default
# (default: appliance → island 169.254.253.1:8443; byo → vmbr0 IP:8443)
# --no-island appliance only: keep the historical LAN bind instead of the R-50 island
# --rootfs-grow N grow OS rootfs by N GiB (default: auto-compute)
# --datavol-grow N grow Docker-data vol by N GiB (default: auto-compute)
# --sysdata-grow N grow user-data vol by N GiB (default: auto-compute)
#
# Guest cap (appliance: optional — protect a SHARED host's other guests; byo: BOTH REQUIRED —
# the only noisy-neighbor protection on a host you do not own; needs agent >= v0.52.0):
# --cores N cap the guest to N CPU cores (0/unset = golden default)
# --memory M cap the guest RAM to M MiB (0/unset = golden default)
#
# --passphrase-file PATH read the retrieval passphrase from a 0600 file
# (default: secure no-echo prompt)
# --operator-pubkey-file PATH pin the operator signing PUBLIC keys at day-0 (GL-4). One key per
# line: "operational <authorized_keys line>" or "recovery <authorized_keys
# line>" (# comments/blank ok; the authorized_keys comment field is the
# key_id and is required). Overrides the script's OPERATOR_KEY_* constants.
# Written to authz.signers in the agent config; no keys resolved = a
# preserved config's signers are KEPT, else self-update stays dormant.
# --preserve-from PATH merge non-Day-0 sections (privileged/storage/backup/
# local_api/authz/lan_resolver) from an existing config
# --preserve-state-from PATH carry the prior agent leaf+key+token-store (local-api.crt/key,
# local-tokens.log) over so the pinned fingerprint STAYS STABLE across a
# reinstall (no controller re-bootstrap). Use an aside copy of the old
# /var/lib/felhom-agent.
# --allow-new-leaf opt in to REGENERATE the agent leaf on a host that already has guests
# (the populated-host guard otherwise refuses; every guest must then be
# re-bootstrapped — only use intentionally).
# --force allow provisioning over an EXISTING vmid (destructive)
# --skip-provision install + configure + verify the agent, but do NOT
# provision a guest (re-install/upgrade an agent on a host
# that already has live guests; also the agent-only path)
# --dry-run print every mutating command without executing
# --resume skip steps already recorded in the state file
# -h, --help this help
#
# Uninstall (local host teardown — no hub contact, no passphrase):
# --uninstall cleanly revert an install: destroy the Felhom guest, remove the agent
# (unit/sudoers/binary/state/config/user + runtime artifacts: shared-parent
# unit, mkfs + pbs-apply wrappers, hook snippet, dnsmasq snippets), the pveum
# role/user/token/ACL, and the install state file. Refuses a non-Felhom guest (no
# /etc/felhom-bootstrap mount) and skips host-level removal if OTHER Felhom
# guests remain (both overridable with --force). Typed vmid confirmation
# required. Reuses --vmid (else the recorded provisioned_vmid), --force,
# --archive-storage, --golden-vmid, --dry-run.
# --remove-golden with --uninstall, also delete the golden vzdump from the archive storage
#
# Retrofit (local, non-destructive — no hub contact, no passphrase):
# --adopt-pool add an EXISTING Felhom guest to the `felhom` pool (creates the pool if
# needed). Resolves the guest from --vmid else the recorded provisioned_vmid;
# refuses a non-Felhom guest unless --force. Touches ONLY pool membership —
# never reconfigures/restarts the guest. (A fresh provision joins the pool
# automatically; this retrofits already-installed boxes.)
# --rescope-acl migrate an existing install from the pre-3b broad-`/` token grant to the
# pool-scoped ACL (Guest@/pool/felhom + Store@each storage + Sys.Audit/SDN.Use@/).
# ACL-only (no data touched). SUPERVISED: run with felhom-agent STOPPED, then
# deploy agent >= v0.53.0, then start (the scoped ACL + pool-param agent are
# mutually dependent).
# --acl-storages "a b c" override the storages the scoped ACL grants Datastore.* on
# (default: "local local-lvm felhom-pbs"). Used by fresh install + --rescope-acl.
#
# State (idempotent/resumable): /var/lib/felhom-install/state.json
# Agent config written 0600 to the systemd unit's -config path
# (auto-detected; else /etc/felhom-agent/agent.json).
#
# SECURITY: the passphrase is read no-echo or from a 0600 file — never a CLI arg,
# never echoed, never written to the state file or logs. The minted pve-token
# secret + per-host hub api_key live ONLY in the agent config (0600, root).
#===============================================================================
set -euo pipefail
SCRIPT_VERSION="1.22.0" # the SINGLE version source (F-1): -h and the run banners follow it.
# The hub used to carry a copy for its Setup tab; R-94 DELETED it
# (2026-08-02) because the hub cannot know which version a box runs —
# the Setup command fetches this script at run time. scripts/
# hostinstall_gates.py now asserts the hub carries NO version literal.
# 1.16.0: the FELHOM_ESCROW sudoers alias (controller-driven escrow
# ceremony, agent v0.88.0) ships on every install/update — it rides the
# CANONICAL sudoers fetch below (configs/felhom-agent.sudoers from the
# agent repo, visudo-gated), no separate installer step.
# Operator signing keys pinned at day-0 (GL-4; doc 04 §3 two-key model). PINNED below to the real
# ceremony keypairs (the pin ceremony is DONE — felhom-op-1 / felhom-rec-1); --operator-pubkey-file
# at install time overrides them. If these were ever cleared to empty, no authz.signers is written
# and agent self-update stays DORMANT (the safe fallback; the verify step warns). PUBLIC keys only —
# this script never generates, reads, or references private key material.
OPERATOR_KEY_OPERATIONAL_ID="felhom-op-1"
OPERATOR_KEY_OPERATIONAL_LINE="ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIL8z0qCNgA3x2xxAB0Qj5ro8waFjGZ8Ta/sWB63tlLw+ felhom-op-1"
OPERATOR_KEY_RECOVERY_ID="felhom-rec-1"
OPERATOR_KEY_RECOVERY_LINE="ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIC/gZI5mvXmUzImh81iHdjdVYwmQ4fRGKqn+9td0fI5/ felhom-rec-1"
#-------------------------------------------------------------------------------
# Logging (mirrors felhom-controller/scripts/docker-setup.sh)
#-------------------------------------------------------------------------------
RED='\033[0;31m'; GREEN='\033[0;32m'; YELLOW='\033[1;33m'
BLUE='\033[0;34m'; CYAN='\033[0;36m'; NC='\033[0m'
log_info() { echo -e "${GREEN}[INFO]${NC} $1"; }
log_warn() { echo -e "${YELLOW}[WARN]${NC} $1"; }
log_error() { echo -e "${RED}[ERROR]${NC} $1" >&2; }
log_step() { echo -e "${BLUE}[STEP]${NC} $1"; }
log_success() { echo -e "${GREEN}[OK]${NC} $1"; }
log_skip() { echo -e "${CYAN}[SKIP]${NC} $1"; }
log_dry() { echo -e "${CYAN}[DRY-RUN]${NC} $1"; }
die() { log_error "$1"; exit 1; }
#-------------------------------------------------------------------------------
# Defaults
#-------------------------------------------------------------------------------
CUSTOMER_ID=""
MODE="" # --mode appliance|byo (GL-2: REQUIRED for install/--resume, no default; "dr" = the 10D stub)
PREFLIGHT_ONLY=false # --preflight-only: run the mode-aware step-1 checks only; no state writes, no later step
HUB_URL="https://hub.felhom.eu"
VMID="9201"
VMID_EXPLICIT=false # set true when --vmid is given; gates the auto-pick-a-free-vmid behavior
GOLDEN_VOLID=""
GOLDEN_VMID="9100"
ARCHIVE_STORAGE="local"
NODE=""
NODE_EXPLICIT=false # set true when --node is given; gates the multi-node wrong-node guard
BRIDGE_ADDR=""
# R-50 island control plane (SPIKE-island-bridge-2026-07-25, GO). The appliance default binds the
# local API on a host-internal, portless bridge with a fixed private /30 so the controller→agent
# channel survives any LAN/DHCP/site move (the F1 fix). These are the spike's VALIDATED parameters —
# do not change them. Set ISLAND_ON=1 only in appliance mode with no explicit --bridge-ip override
# (byo and an explicit --bridge-ip keep the historical LAN bind). --no-island opts an appliance out.
ISLAND_BRIDGE="vmbr9"
ISLAND_HOST_CIDR="169.254.253.1/30"
ISLAND_HOST_IP="169.254.253.1"
ISLAND_GUEST_CIDR="169.254.253.2/30"
ISLAND_PORT="8443"
ISLAND_ON=0
WANT_ISLAND=true # appliance default; --no-island sets false
LAN_IP="" # vmbr0 IPv4 (computed in step_agent_config; the LAN anchor for lan_resolver)
ROOTFS_GROW=""
DATAVOL_GROW=""
SYSDATA_GROW=""
CPU_CORES="" # --cores: optional appliance CPU-core cap (empty/unset = golden default)
MEM_MIB="" # --memory: optional appliance RAM cap in MiB (empty/unset = golden default)
PASSPHRASE_FILE=""
OPERATOR_PUBKEY_FILE="" # --operator-pubkey-file: "operational|recovery <authorized_keys line>" per line; overrides the OPERATOR_KEY_* constants (GL-4)
PRESERVE_FROM=""
PRESERVE_STATE_FROM="" # dir holding a prior local-api.{crt,key} + local-tokens.log to carry over (keeps the pin stable across a reinstall)
ALLOW_NEW_LEAF=false # opt-in to intentionally regenerate the agent leaf on a populated host (else the guard refuses)
FORCE=false
FORCE_GITEA_GOLDEN=false
SKIP_PROVISION=false
DRY_RUN=false
RESUME=false
UNINSTALL=false # --uninstall: local host teardown (destroy guest + remove agent/pveum/state)
REMOVE_GOLDEN=false # --remove-golden: also delete the golden vzdump during --uninstall
ADOPT_POOL=false # --adopt-pool: retrofit an EXISTING Felhom guest into the felhom pool (non-destructive)
RESCOPE_ACL=false # --rescope-acl: migrate an existing install from the broad-/ token to the scoped ACL
ROTATE_RECOVERY=false # --rotate-recovery: regenerate + re-vault the break-glass root@pam password (TASK G1)
ENABLE_OOB=false # the OOB belt (TASK H1). DEFAULT-ON for appliance mode since v1.25.0 (F9 ruling
# 2026-07-23: an operator-owned box must be reachable OOB) — see the appliance
# default below; --enable-oob forces it, --no-oob opts an appliance box out.
NO_OOB=false # --no-oob: opt an APPLIANCE install out of the default belt (byo is off already)
# --- Gitea (artifact source) + agent install model (BUNDLE slice) ---
GITEA_BASE="https://gitea.dooplex.hu"
GITEA_OWNER="admin"
AGENT_REPO="felhom-agent" # for the raw unit/sudoers fetch (config text, canonical source)
AGENT_USER="felhom-agent" # the non-root service user the unit + sudoers name
AGENT_BIN="/usr/local/bin/felhom-agent"
AGENT_SUDOERS="/etc/sudoers.d/felhom-agent"
AGENT_UNIT="/etc/systemd/system/felhom-agent.service"
AGENT_STATE_DIR="/var/lib/felhom-agent"
PVE_USER="felhom-agent@pve"
PVE_TOKENID="agent"
PVE_POOL="felhom" # dedicated pool every Felhom-managed guest joins (fleet uniformity + ACL scope)
PVE_ROLE="FelhomAgent" # the PRE-3b single broad role (removed on rescope/uninstall if present — legacy)
# Pool-scoped ACL (3b, validated by SPIKE-pool-scoped-acl-2026-07-01): the agent's privileges are split
# across THREE roles applied at scoped paths so the token can only touch Felhom's own guests + storages
# (blast-radius containment). `pveum acl` grants a whole role per path, hence 3 roles not 1. Each role is
# granted to BOTH the user AND the token (privsep intersection). Guest privs (incl. Pool.Allocate so the
# agent restores INTO the pool) live at /pool/felhom; Datastore WRITE privs at each agent-touched storage.
# `Datastore.Audit` is box-wide in Base (3b-fix v1.7.0): the agent must ENUMERATE every storage incl. the
# dynamically-enrolled removable drives (felhom-usb/felhom-flash) it observes but never registers — a
# per-storage Audit grant hid them → false "drive detached" alerts. Audit is read-only, so box-wide Audit
# keeps WRITE containment (Allocate/AllocateSpace stay per-storage). Only Sys.Audit/SDN.Use/Datastore.Audit box-wide.
# `Pool.Audit` (v1.9.0, audit A1): the agent's stale-lock reaper reads GET /pools/felhom as its
# ownership registry (agent v0.62.0+); without it the reaper fail-safes (skips) and reports the
# `pve:pool-read` capability degraded. NOTE: Pool.Allocate does NOT satisfy the read — the spike
# (SPIKE-a1-pool-membership-read-2026-07-03 T2) 403'd with Allocate granted; Audit is required.
PVE_ROLE_GUEST="FelhomAgentGuest"
PVE_ROLE_STORE="FelhomAgentStore"
PVE_ROLE_BASE="FelhomAgentBase"
PVE_PRIVS_GUEST="VM.Allocate VM.Audit VM.Config.Disk VM.Config.CPU VM.Config.Memory VM.Config.Network VM.Config.Options VM.PowerMgmt VM.Snapshot VM.Snapshot.Rollback VM.Backup Pool.Allocate Pool.Audit"
PVE_PRIVS_STORE="Datastore.Allocate Datastore.AllocateSpace"
PVE_PRIVS_BASE="Sys.Audit SDN.Use Datastore.Audit"
# Storages the agent reads/writes (archive+dump=local, restore=local-lvm, offsite DR=felhom-pbs). The
# offsite felhom-pbs MUST be included or the agent's DR backup 403s (SPIKE residual #1). --acl-storages overrides.
PVE_STORAGES=(local local-lvm felhom-pbs)
# E-2: the whole-guest backup target storage id, and what configure_backup_target resolved to.
# BACKUP_TARGET_RESOLVED feeds agent.json's local_backup_target; "local" means DEGRADED (Case B).
BACKUP_TARGET_ID="felhom-backup"
BACKUP_TARGET_RESOLVED="local"
# F-LEAK (Campaign 8, v1.21.0): the restore-test's scratch VMID band, granted the GUEST role at each
# /vms/<id> PATH. WHY THIS EXISTS: the Guest role is granted at /pool/felhom, and a guest only joins
# that pool when its restore COMPLETES. A restore-test whose restore FAILS therefore leaves a scratch
# guest that exists but is in no pool — and the agent's own teardown then 403s on VM.Allocate, so it
# leaks a half-restored guest holding its disks until a human removes it.
# THE ALTERNATIVE THAT DOES NOT WORK: adopting the stranded guest into the pool first. Tried, shipped,
# and REFUTED live on 2026-07-28 — `PUT /pools/felhom` with vms=<id> ALSO requires VM.Allocate on
# /vms/<id>, so Pool.Allocate cannot bootstrap its own membership.
# WHY THIS IS NOT A WIDENING: PVE ACLs are path-scoped, so ten explicit /vms/<id> grants authorise the
# agent on exactly these ten scratch IDs and nowhere else. It is NOT granted at /vms — that would
# authorise destroying every guest on the box, including a co-tenant's. Verify with
# pvesh get /access/permissions --userid 'felhom-agent@pve!agent' --path /vms/<some-real-vmid>
# which must NOT list VM.Allocate. The agent keeps its own in-band check; this is the outer fence.
PVE_SCRATCH_VMID_MIN=990000
PVE_SCRATCH_VMID_MAX=990009
# FELHOM_INSTALL_STATE_DIR: test-harness-only override (hostinstall-mode-harness.sh) so dry/refusal
# cases can never touch a live install's state.json. Production runs never set it.
STATE_DIR="${FELHOM_INSTALL_STATE_DIR:-/var/lib/felhom-install}"
STATE_FILE="${STATE_DIR}/state.json"
AGENT_CONFIG="" # resolved in preflight
HARD_MIN_LVM_GIB=120 # a useful appliance won't fit below this on local-lvm
# Runtime carriers (never logged)
PASSPHRASE=""
PVE_TOKEN="" # felhom-agent@pve!agent=<secret>
HOST_ID=""
HOST_API_KEY=""
GIT_USER="" # from controller.yaml (config-retrieve) — Gitea fetch credential
GIT_TOKEN="" # from controller.yaml — NEVER logged
ART_AGENT_VER="" # hub artifact manifest: agent version + sha256
ART_AGENT_SHA=""
ART_GOLDEN_VER="" # hub artifact manifest: golden version + sha256
ART_GOLDEN_SHA=""
#-------------------------------------------------------------------------------
# Helpers
#-------------------------------------------------------------------------------
# Print the header through the last option line. Markers, not line numbers (v1.9.1's numeric
# '2,95p' had drifted): start at the "# Day-0 host-bootstrap" line, end before "# State (".
# The version line is printed from SCRIPT_VERSION — the ONLY version source (F-1).
usage() {
echo "felhom-host-install.sh v${SCRIPT_VERSION}"
sed -n '/^# Day-0 host-bootstrap/,/^# State (/p' "$0" | sed '$d' | sed 's/^# \{0,1\}//'
exit 0
}
run() { # simple (no pipes/redirects) mutating command
if $DRY_RUN; then log_dry "$*"; else "$@"; fi
}
# used_vmids — every in-use guest id on this host. LXC (pct) and VMs (qm) SHARE the id space,
# so both are consulted; headers (non-numeric first column) are filtered out.
used_vmids() {
{ pct list 2>/dev/null; qm list 2>/dev/null; } | awk '{print $1}' | grep -E '^[0-9]+$'
}
# _vmid_in_use ID — true if ID is present in the pct+qm used-set (more complete than `pct status`,
# which only knows LXC).
_vmid_in_use() {
local target="$1" used
used=" $(used_vmids | tr '\n' ' ') "
[[ "$used" == *" $target "* ]]
}
# next_free_vmid BASE — the first id >= BASE not in the used-set, scanning upward.
next_free_vmid() {
local base="$1" used id
used=" $(used_vmids | tr '\n' ' ') "
id="$base"
while [[ "$used" == *" $id "* ]]; do id=$((id + 1)); done
echo "$id"
}
# State helpers (robust JSON via python3).
_state_has() {
[[ -f "$STATE_FILE" ]] || return 1
STATE_FILE="$STATE_FILE" python3 -c "import json,os,sys;f=os.environ['STATE_FILE'];d=json.load(open(f));sys.exit(0 if sys.argv[1] in d.get('completed',[]) else 1)" "$1" 2>/dev/null
}
_state_mark() {
$DRY_RUN && return 0
mkdir -p "$STATE_DIR"
STATE_FILE="$STATE_FILE" python3 -c "import json,os,sys;f=os.environ['STATE_FILE'];d=json.load(open(f)) if os.path.exists(f) else {'completed':[]};c=d.setdefault('completed',[]);(c.append(sys.argv[1]) if sys.argv[1] not in c else None);json.dump(d,open(f,'w'),indent=2)" "$1"
}
should_skip() { # returns 0 (skip) if --resume AND step already done
if $RESUME && _state_has "$1"; then log_skip "step '$1' already completed"; return 0; fi
return 1
}
# _state_put KEY VALUE — set a top-level string key in state.json (creates the file if absent).
# Mirrors _state_mark: dry-run no-ops (writes nothing), robust JSON via python3.
_state_put() {
$DRY_RUN && return 0
mkdir -p "$STATE_DIR"
STATE_FILE="$STATE_FILE" python3 -c "import json,os,sys;f=os.environ['STATE_FILE'];d=json.load(open(f)) if os.path.exists(f) else {'completed':[]};d[sys.argv[1]]=sys.argv[2];json.dump(d,open(f,'w'),indent=2)" "$1" "$2"
}
# _state_get KEY — print the top-level string value for KEY (empty if the file/key is absent).
_state_get() {
[[ -f "$STATE_FILE" ]] || return 0
STATE_FILE="$STATE_FILE" python3 -c "import json,os,sys;d=json.load(open(os.environ['STATE_FILE']));print(d.get(sys.argv[1],''))" "$1" 2>/dev/null
}
http_code() { # GET, prints status code only (read-only preflight)
curl -fsS -o /dev/null -w '%{http_code}' "$@" 2>/dev/null || curl -sS -o /dev/null -w '%{http_code}' "$@" 2>/dev/null
}
#-------------------------------------------------------------------------------
# Artifact + Gitea helpers (BUNDLE slice)
#-------------------------------------------------------------------------------
# Resolve the hub-vouched artifact manifest (agent + golden version+sha256). Passphrase-authed,
# same trust root as config-retrieve. Sets ART_* globals. Empty fields are valid (caller falls back).
resolve_artifacts() {
local resp code body
resp=$(curl -sS -w $'\n%{http_code}' "$HUB_URL/api/v1/artifacts/$CUSTOMER_ID" \
-H "X-Retrieval-Password: $PASSPHRASE")
code=$(tail -n1 <<<"$resp"); body=$(sed '$d' <<<"$resp")
[[ "$code" == "200" ]] || die "artifact manifest fetch failed: HTTP $code"
ART_AGENT_VER=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['agent']['version'])" "$body" 2>/dev/null || echo "")
ART_AGENT_SHA=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['agent']['sha256'])" "$body" 2>/dev/null || echo "")
ART_GOLDEN_VER=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['golden']['version'])" "$body" 2>/dev/null || echo "")
ART_GOLDEN_SHA=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['golden']['sha256'])" "$body" 2>/dev/null || echo "")
}
# Resolve the Gitea fetch credential (git username + token) from the customer's controller.yaml —
# the SAME secret config-retrieve already hands out (NO new credential). Sets GIT_USER / GIT_TOKEN.
# Parses the git: block without a YAML lib (fresh PVE has no PyYAML).
# v1.11.2 (GL-6 Gate-0 operator ruling, G3): EMPTY credentials are ALLOWED — the fetches then run
# ANONYMOUSLY (the Felhom artifacts are deliberately world-readable; integrity comes from the
# hub-vouched sha256 + visudo/sh -n validation, never from transport auth). Loud warn, never die:
# a customer config without a git token must not block day-0 while the anonymous path works.
resolve_git_creds() {
local yaml
yaml=$(curl -fsS "$HUB_URL/api/v1/config/$CUSTOMER_ID" -H "X-Retrieval-Password: $PASSPHRASE") \
|| die "controller.yaml fetch failed (for the git fetch credential)"
GIT_USER=$(awk '/^[^[:space:]#]/{ingit=($1=="git:")} ingit&&$1=="username:"{print $2}' <<<"$yaml" | head -1)
GIT_TOKEN=$(awk '/^[^[:space:]#]/{ingit=($1=="git:")} ingit&&$1=="token:"{print $2}' <<<"$yaml" | head -1)
# strip any surrounding quotes
GIT_USER="${GIT_USER%\"}"; GIT_USER="${GIT_USER#\"}"
GIT_TOKEN="${GIT_TOKEN%\"}"; GIT_TOKEN="${GIT_TOKEN#\"}"
if [[ -z "$GIT_TOKEN" ]]; then
log_warn " no git credential in controller.yaml — fetching artifacts ANONYMOUSLY (they are world-readable; sha256 verification unchanged)"
fi
}
# _git_auth_args populates the caller's curl arg array with basic-auth ONLY when a credential is
# present — empty creds = anonymous fetch (v1.11.2). `curl -u user:` with an empty token would NOT
# fall back to anonymous (a wrong credential 401s even on world-readable content), hence the array.
_git_auth_args() {
local -n _out=$1
_out=()
if [[ -n "$GIT_TOKEN" ]]; then
_out=(-u "${GIT_USER}:${GIT_TOKEN}")
fi
}
# _dry_auth_hint — the dry-run print's auth stanza (drill F-2): the old hardcoded "curl -u <git>"
# lied on the anonymous-fetch branch. Requires resolve_git_creds to have run (both call sites do).
_dry_auth_hint() {
if [[ -n "$GIT_TOKEN" ]]; then echo "curl -u <git>"; else echo "curl (anonymous — no git credential)"; fi
}
# Fetch a Gitea generic-package URL to a dest (authed when a credential exists, else anonymous),
# then VERIFY its sha256 against the expected (hub-vouched) value. Aborts on any mismatch —
# verify-before-use; the sha is the integrity root either way. $1=url $2=dest $3=expected_sha
fetch_verify() {
local url="$1" dest="$2" want="$3"
[[ -n "$want" ]] || die "refusing to install an artifact with no expected sha256 (manifest incomplete): $url"
local -a _auth; _git_auth_args _auth
curl -fsS "${_auth[@]}" -o "$dest" "$url" || die "fetch failed: $url"
local got; got=$(sha256sum "$dest" | awk '{print $1}')
if [[ "$got" != "$want" ]]; then
rm -f "$dest"
die "sha256 MISMATCH for $url — expected $want got $got. Refusing to install (verify-before-use)."
fi
log_success " verified sha256 ${got:0:16}… matches the hub manifest"
}
# Fetch a raw config file (the canonical unit/sudoers) from the agent repo (authed when a credential
# exists, else anonymous). These are non-executable text (not the integrity-checked binary); the
# sudoers is `visudo -cf`-validated before install, which catches corruption/tampering that would
# matter. $1=repo-path $2=dest
fetch_raw() {
local path="$1" dest="$2"
local -a _auth; _git_auth_args _auth
curl -fsS "${_auth[@]}" -o "$dest" \
"$GITEA_BASE/$GITEA_OWNER/$AGENT_REPO/raw/branch/main/$path" \
|| die "raw fetch failed: $path"
[[ -s "$dest" ]] || die "raw fetch empty: $path"
}
#-------------------------------------------------------------------------------
# Uninstall (local host teardown) — reverse of install; no hub contact, no passphrase
#-------------------------------------------------------------------------------
# felhom_guests — every vmid on this host that carries the /etc/felhom-bootstrap bind mount (the
# read-only bootstrap mount an agent-provisioned guest always has). Matched by the CONSTANT guest
# PATH, not a hardcoded mpN slot (the slot drifts; on the demo host it's mp9).
felhom_guests() {
local id
for id in $(used_vmids); do
pct config "$id" 2>/dev/null | grep -q 'mp=/etc/felhom-bootstrap' && echo "$id"
done
}
#-------------------------------------------------------------------------------
# felhom pool (fleet uniformity) — every managed guest joins the `felhom` pool. All pool ops run as
# root@pam from the installer, so NO agent/token/ACL change is involved (that is the separate 3b spike).
# API shapes confirmed on PVE 9: `pvesh get /pools` → [{poolid,comment}]; `pvesh get /pools/<id>` →
# {poolid,comment,members:[{vmid,...}]}. Pool ops: `pveum pool add|delete <id>`, `pveum pool modify
# <id> --vms <ids>` (additive).
#-------------------------------------------------------------------------------
# pool_exists — true if the felhom pool is present.
pool_exists() {
pvesh get /pools --output-format json 2>/dev/null \
| python3 -c "import json,sys;sys.exit(0 if any(p.get('poolid')=='$PVE_POOL' for p in json.load(sys.stdin)) else 1)" 2>/dev/null
}
# pool_members — space-separated vmids currently in the felhom pool (empty if none / pool absent).
pool_members() {
pvesh get "/pools/$PVE_POOL" --output-format json 2>/dev/null \
| python3 -c "import json,sys
try: d=json.load(sys.stdin)
except Exception: sys.exit(0)
print(' '.join(str(m.get('vmid')) for m in d.get('members',[]) if m.get('vmid') is not None))" 2>/dev/null
}
# ensure_felhom_pool — create the pool if absent (idempotent no-op otherwise). Via run() (dry-run-aware).
ensure_felhom_pool() {
if pool_exists; then
log_skip " pool $PVE_POOL already exists"
else
run pveum pool add "$PVE_POOL" --comment "Felhom-managed guests"
fi
}
# pool_add_guest VMID — add a guest to the felhom pool unless it is already a member (idempotent).
pool_add_guest() {
local vmid="$1" members
members=" $(pool_members) "
if [[ "$members" == *" $vmid "* ]]; then
log_skip " guest $vmid already in pool $PVE_POOL"
else
run pveum pool modify "$PVE_POOL" -vms "$vmid"
log_success " guest $vmid added to pool $PVE_POOL"
fi
}
#-------------------------------------------------------------------------------
# Pool-scoped ACL helpers (3b). All ops run as root@pam (installer) — no privilege change to the agent.
#-------------------------------------------------------------------------------
# _role_exists NAME — true if a pveum role NAME exists.
_role_exists() {
pveum role list --output-format json 2>/dev/null \
| python3 -c "import json,sys;sys.exit(0 if any(r['roleid']==sys.argv[1] for r in json.load(sys.stdin)) else 1)" "$1" 2>/dev/null
}
# _ensure_role NAME "PRIVS" — create the role, or modify it to the exact priv set (idempotent).
_ensure_role() {
local name="$1" privs="$2"
if _role_exists "$name"; then
log_info " role $name exists — ensuring exact privileges"
run pveum role modify "$name" -privs "$privs"
else
run pveum role add "$name" -privs "$privs"
fi
}
# _grant PATH ROLE — grant ROLE at PATH to BOTH the user AND the token (privsep intersection). `acl
# modify` is idempotent so this is safe to repeat / re-apply after a token rotation.
_grant() {
local path="$1" role="$2"
run pveum acl modify "$path" -user "$PVE_USER" -role "$role"
run pveum acl modify "$path" -token "${PVE_USER}!${PVE_TOKENID}" -role "$role"
}
# ---- E-2 Part 2: the whole-guest backup target -------------------------------------------------
#
# CASE A an eligible secondary drive is already mounted at install time -> create `felhom-backup`
# on that drive's OWN mountpoint (F-1), is_mountpoint 1 (F-2), grant the agent (F-3), and
# point the primary tier at it.
# CASE B system drive only -> the target stays on the system drive. This is RECORDED AS DEGRADED,
# not as normal: it protects against corruption and NOT against drive loss. The install
# still succeeds — a single-drive appliance is a valid product, it just cannot survive
# drive loss.
#
# PHASE-0 REALITY CHECK (E-2, 2026-07-29): the installer has NO drive-enrollment step —
# /mnt/felhom-drives is agent-installed at runtime and drives are formatted/mounted by the dashboard
# later. So on a FRESH appliance Case A almost never fires; the common case is Case B, and the drive
# arrives months later through the wizard (E-2 Part 3). Case A exists for the reinstall/recovery
# shape, where an agent-generated .mount unit already brings the drive up by fs-UUID.
#
# eligible_backup_drive — echoes the mountpoint of a usable secondary drive, or nothing.
# Attributes SUGGEST and REFUSE THE ABSURD; they never decide by transport. The reference hardware
# forbids that: demo-felhom's backup drive is an external USB HDD, and BOTH demo boxes' drives
# report removable=0 — a transport rule would disqualify the reference drive, a removable rule
# would find no candidate at all.
eligible_backup_drive() {
local root_dev mp dev sz
root_dev=$(findmnt -no SOURCE / 2>/dev/null || true)
# Only real, already-mounted filesystems under /mnt — never / , never a bare directory.
while read -r mp dev; do
[[ -n "$mp" && -n "$dev" ]] || continue
[[ "$mp" == /mnt/* ]] || continue
[[ "$dev" != "$root_dev" ]] || continue
[[ "$dev" == /dev/* ]] || continue # skip NAS/tmpfs/overlay sources
# Refuse the absurd: a thumb-drive-sized volume is not a whole-guest backup target. A guest
# archive measured 1.6-6.0 GB on the demo fleet, and keep-last=3 multiplies it.
sz=$(findmnt -nbo SIZE "$mp" 2>/dev/null || echo 0)
[[ "$sz" -ge $((64 * 1024 * 1024 * 1024)) ]] || continue
echo "$mp"; return 0
done < <(findmnt -rn -o TARGET,SOURCE -t ext4,xfs,btrfs 2>/dev/null)
return 1
}
# configure_backup_target — Case A/B. Idempotent and SCENARIO-F SAFE.
configure_backup_target() {
# SCENARIO F, the hard rule: never touch a box that already has a target. The agent.json write
# below is already protected structurally (`setdefault` preserves an existing `backup` block
# WHOLE), and this is the second, explicit guard for the STORAGE half — `pvesm add` on an
# existing id is an error, and repointing a live target is exactly the silent retarget this
# whole arc closes.
if pvesm status --storage "$BACKUP_TARGET_ID" >/dev/null 2>&1; then
log_skip " backup target '$BACKUP_TARGET_ID' already exists — leaving it exactly as it is (Scenario F)"
BACKUP_TARGET_RESOLVED="$BACKUP_TARGET_ID"
return 0
fi
local mp
if mp=$(eligible_backup_drive); then
# CASE A. Run through the SAME guarded wrapper the agent uses, so the installer and the
# runtime path cannot drift and the laws are enforced in exactly one place.
if $DRY_RUN; then
log_dry "felhom-backup-target-apply create $BACKUP_TARGET_ID $mp ; grant"
else
/usr/local/sbin/felhom-backup-target-apply create "$BACKUP_TARGET_ID" "$mp" \
|| die "backup target creation failed on $mp"
/usr/local/sbin/felhom-backup-target-apply grant "$BACKUP_TARGET_ID" \
|| die "backup target grant failed — every backup would 403"
fi
BACKUP_TARGET_RESOLVED="$BACKUP_TARGET_ID"
log_success " backup target: $BACKUP_TARGET_ID on $mp (separate device — survives a system-drive failure)"
else
# CASE B — degraded, and SAID SO. Not a failure; a single-drive appliance is valid.
BACKUP_TARGET_RESOLVED="local"
log_warn " backup target: DEGRADED — no eligible second drive, so the whole-system backup stays on the SYSTEM drive."
log_warn " It protects against file corruption but NOT against a disk failure. Attach a second drive and assign it in the dashboard."
fi
}
# apply_scoped_acl — create the 3 scoped roles and grant each at its path(s). Requires the pool to exist.
# ORDER (3b-fix): Base (which holds box-wide Datastore.Audit) is ensured + granted BEFORE Store, so a
# RE-APPLY on a live box adds Audit@/ before Store drops its per-storage Audit → the agent never loses
# storage-enumeration visibility mid-apply (gap-free).
apply_scoped_acl() {
_ensure_role "$PVE_ROLE_BASE" "$PVE_PRIVS_BASE"
_ensure_role "$PVE_ROLE_GUEST" "$PVE_PRIVS_GUEST"
_ensure_role "$PVE_ROLE_STORE" "$PVE_PRIVS_STORE"
_grant / "$PVE_ROLE_BASE"
_grant "/pool/$PVE_POOL" "$PVE_ROLE_GUEST"
local s
for s in "${PVE_STORAGES[@]}"; do
_grant "/storage/$s" "$PVE_ROLE_STORE"
done
# F-LEAK: the scratch band, per-VMID. A failed restore-test's scratch never joins the pool, so the
# pool grant cannot reach it; these path-scoped grants can, and reach nothing else.
local v
for ((v = PVE_SCRATCH_VMID_MIN; v <= PVE_SCRATCH_VMID_MAX; v++)); do
_grant "/vms/$v" "$PVE_ROLE_GUEST"
done
log_success " scoped ACL applied (Base@/, Guest@/pool/$PVE_POOL + /vms/${PVE_SCRATCH_VMID_MIN}..${PVE_SCRATCH_VMID_MAX}, Store@[${PVE_STORAGES[*]}])"
}
# _acl_grant_present PATH TYPE UGID ROLE — true if that exact ACL grant exists.
_acl_grant_present() {
pveum acl list --output-format json 2>/dev/null | python3 -c "import json,sys
p,t,u,r=sys.argv[1:5]
sys.exit(0 if any(e.get('path')==p and e.get('type')==t and e.get('ugid')==u and e.get('roleid')==r for e in json.load(sys.stdin)) else 1)" "$1" "$2" "$3" "$4" 2>/dev/null
}
# remove_scoped_acl — delete the 3-role scoped grants (user+token at each path), then the 3 roles
# (roles last — PVE refuses to delete a referenced role). Presence-checked, tolerate-absent.
remove_scoped_acl() {
local s
if _acl_grant_present "/pool/$PVE_POOL" user "$PVE_USER" "$PVE_ROLE_GUEST"; then run pveum acl delete "/pool/$PVE_POOL" --users "$PVE_USER" --roles "$PVE_ROLE_GUEST"; fi
if _acl_grant_present "/pool/$PVE_POOL" token "${PVE_USER}!${PVE_TOKENID}" "$PVE_ROLE_GUEST"; then run pveum acl delete "/pool/$PVE_POOL" --tokens "${PVE_USER}!${PVE_TOKENID}" --roles "$PVE_ROLE_GUEST"; fi
# F-LEAK scratch-band grants. These MUST be removed before the role delete below — PVE refuses to
# delete a role that is still referenced by any ACL, so leaving them behind breaks the uninstall.
local v
for ((v = PVE_SCRATCH_VMID_MIN; v <= PVE_SCRATCH_VMID_MAX; v++)); do
if _acl_grant_present "/vms/$v" user "$PVE_USER" "$PVE_ROLE_GUEST"; then run pveum acl delete "/vms/$v" --users "$PVE_USER" --roles "$PVE_ROLE_GUEST"; fi
if _acl_grant_present "/vms/$v" token "${PVE_USER}!${PVE_TOKENID}" "$PVE_ROLE_GUEST"; then run pveum acl delete "/vms/$v" --tokens "${PVE_USER}!${PVE_TOKENID}" --roles "$PVE_ROLE_GUEST"; fi
done
for s in "${PVE_STORAGES[@]}"; do
if _acl_grant_present "/storage/$s" user "$PVE_USER" "$PVE_ROLE_STORE"; then run pveum acl delete "/storage/$s" --users "$PVE_USER" --roles "$PVE_ROLE_STORE"; fi
if _acl_grant_present "/storage/$s" token "${PVE_USER}!${PVE_TOKENID}" "$PVE_ROLE_STORE"; then run pveum acl delete "/storage/$s" --tokens "${PVE_USER}!${PVE_TOKENID}" --roles "$PVE_ROLE_STORE"; fi
done
if _acl_grant_present / user "$PVE_USER" "$PVE_ROLE_BASE"; then run pveum acl delete / --users "$PVE_USER" --roles "$PVE_ROLE_BASE"; fi
if _acl_grant_present / token "${PVE_USER}!${PVE_TOKENID}" "$PVE_ROLE_BASE"; then run pveum acl delete / --tokens "${PVE_USER}!${PVE_TOKENID}" --roles "$PVE_ROLE_BASE"; fi
local name
for name in "$PVE_ROLE_GUEST" "$PVE_ROLE_STORE" "$PVE_ROLE_BASE"; do
if _role_exists "$name"; then run pveum role delete "$name"; else log_skip " role $name already absent"; fi
done
}
# remove_old_broad_acl — remove the PRE-3b single FelhomAgent role granted at / (user+token) + the role.
# Tolerate-absent (fresh 3b installs have none). Used by --rescope-acl (migration) and step_token (so a
# re-install can't leave the old broad grant unioned with the new scoped one).
remove_old_broad_acl() {
if _acl_grant_present / user "$PVE_USER" "$PVE_ROLE"; then run pveum acl delete / --users "$PVE_USER" --roles "$PVE_ROLE"; fi
if _acl_grant_present / token "${PVE_USER}!${PVE_TOKENID}" "$PVE_ROLE"; then run pveum acl delete / --tokens "${PVE_USER}!${PVE_TOKENID}" --roles "$PVE_ROLE"; fi
if _role_exists "$PVE_ROLE"; then run pveum role delete "$PVE_ROLE"; else log_skip " old broad role $PVE_ROLE already absent"; fi
}
# _guest_drive_note VMID — GL-4 (guest-only mode): best-effort list of the drives the agent's bind
# store records for THIS vmid (/var/lib/felhom-agent/guest-binds.json, vmid -> durable-ids); store
# unreadable/absent → the generic note. Read-only, never dies.
_guest_drive_note() {
local vmid="$1" binds=""
binds=$(python3 -c "import json
try:
d=json.load(open('$AGENT_STATE_DIR/guest-binds.json'))
print(', '.join(d.get('$vmid',[])))
except Exception:
pass" 2>/dev/null || true)
if [[ -n "$binds" ]]; then
log_info " drives recorded as bound to guest $vmid: $binds"
log_info " they stay mounted (other guests may share the host) — eject via the dashboard BEFORE uninstalling if this guest owned them."
else
log_info " enrolled drives (if any) stay mounted; eject a drive via the dashboard BEFORE uninstall if it belonged to this guest."
fi
}
# _uninstall_statement full|guest-only — GL-4: the explicit end-of-teardown KEPT-vs-WIPED statement.
# WIPED mirrors what THIS mode actually ran; KEPT names everything that deliberately lives on.
# Statement-only: this script NEVER deletes PBS backups, hub records, escrow blobs, or drive data.
# Reads run_uninstall's locals (vmid, pool_removed, _busy_mounts, _had_break_glass) via bash's
# dynamic scoping — call it from run_uninstall only.
_uninstall_statement() {
local scope="$1"
echo ""
log_step "kept vs wiped — read before pulling drives or closing the customer out"
echo " WIPED (this run):"
echo " - guest $vmid (container + its OS/Docker/user-data volumes)"
if [[ "$scope" == "full" ]]; then
echo " - the felhom-agent: binary, unit, sudoers, config (+ its .bak backups), state dir, service user"
echo " - self-update artifacts: guarded wrapper, A/B slots (.prev/.new.*), rollback unit, start-limit drop-in"
echo " - break-glass watchdog + OOB artifacts (where present); guest-hook snippet; dnsmasq snippets; the mkfs + pbs-apply wrappers"
echo " - pveum: the Felhom roles/user/token/scoped ACL$( $pool_removed && printf '; the emptied %s pool' "$PVE_POOL")"
echo " - the install state file"
if $REMOVE_GOLDEN; then echo " - the golden vzdump (--remove-golden)"; fi
else
echo " - NOTHING host-level (other Felhom guests remain: agent, token/ACL, pool, state all stay)"
fi
echo " KEPT (lives on deliberately — remove/rotate these out-of-band if the customer is leaving):"
if [[ "$scope" == "full" ]]; then
echo " - the enrolled drives + ALL data under /mnt/felhom-drives — unmounted only, NEVER wiped;"
if [[ ${#_busy_mounts[@]} -gt 0 ]]; then
echo " physically removable now, EXCEPT still mounted (busy — stop the apps and retry): ${_busy_mounts[*]}"
else
echo " the drives are physically removable now."
fi
else
echo " - the enrolled drives + ALL data under /mnt/felhom-drives — left MOUNTED (remaining guests may use them)"
fi
if pvesm status 2>/dev/null | awk '$2=="pbs"{found=1} END{exit !found}'; then
echo " - the PBS backups + this customer's namespace on the PBS side — delete there if wanted"
fi
echo " - the hub host/customer record + report history (operator UI / DB)"
echo " - the escrow blob in the hub, if one exists (operator UI)"
if $_had_break_glass; then
echo " - the hub-vaulted root@pam recovery credential — the box KEEPS the password step 4b set; rotate it if the box leaves Felhom management"
fi
echo ""
}
# run_uninstall — the full guarded teardown. Every mutation goes through run() so --dry-run prints it
# and executes nothing. Ordering is the reverse of install: guest -> agent -> pveum(ACL,token,user,
# role) -> golden(opt-in) -> state file. See the TASK spec §7/§8.
run_uninstall() {
log_step "UNINSTALL — local host teardown"
# GL-4: snapshot state facts BEFORE any removal (the closing statement needs them; the state
# file itself is deleted in step 7).
local _had_break_glass=false _busy_mounts=()
_state_has break_glass && _had_break_glass=true
# 1. Resolve the target vmid: --vmid, else the recorded provisioned_vmid, else die.
local state_vmid vmid pool_removed=false
state_vmid=$(_state_get provisioned_vmid)
if $VMID_EXPLICIT; then
vmid="$VMID"
elif [[ -n "$state_vmid" ]]; then
vmid="$state_vmid"
log_info " no --vmid given; using recorded provisioned_vmid=$vmid from $STATE_FILE"
else
die "pass --vmid N (state has no recorded vmid)"
fi
# state-mismatch: an explicit --vmid that disagrees with the recorded one needs --force.
if $VMID_EXPLICIT && [[ -n "$state_vmid" && "$state_vmid" != "$vmid" ]]; then
if $FORCE; then
log_warn " --vmid $vmid differs from the recorded provisioned_vmid=$state_vmid — --force given, proceeding"
else
die "--vmid $vmid differs from the recorded provisioned_vmid=$state_vmid. Pass --force to override."
fi
fi
# 2. Guest teardown (guarded: ours-check + typed confirm).
if _vmid_in_use "$vmid"; then
# ours-check: a Felhom guest carries the /etc/felhom-bootstrap bind mount (constant guest path).
if pct config "$vmid" 2>/dev/null | grep -q 'mp=/etc/felhom-bootstrap'; then
log_info " vmid $vmid looks like a Felhom guest (has the /etc/felhom-bootstrap mount)"
elif $FORCE; then
log_warn " vmid $vmid has NO /etc/felhom-bootstrap mount — --force given, destroying anyway"
else
die "vmid $vmid does not look like a Felhom-provisioned guest (no /etc/felhom-bootstrap mount).
Refusing to destroy. Pass --force to override."
fi
# show the config so the operator can eyeball what is about to be destroyed
log_info " pct config $vmid:"
pct config "$vmid" 2>/dev/null | sed 's/^/ /'
# typed confirmation — mandatory, never skipped except in --dry-run (nothing is destroyed there).
if $DRY_RUN; then
log_dry "would prompt: Type the vmid ($vmid) to confirm PERMANENT destruction"
else
local ans
read -rp "Type the vmid ($vmid) to confirm PERMANENT destruction: " ans < /dev/tty
[[ "$ans" == "$vmid" ]] || die "confirmation mismatch (got '$ans', expected '$vmid') — aborting, nothing destroyed"
fi
# stop (tolerate already-stopped) then destroy
local gstat; gstat=$(pct status "$vmid" 2>/dev/null | awk '{print $2}')
if [[ "$gstat" == "running" ]]; then
run pct stop "$vmid"
else
log_skip " guest $vmid not running (status: ${gstat:-unknown}) — skip stop"
fi
run pct destroy "$vmid"
log_success " guest $vmid destroyed"
else
log_skip " guest $vmid already absent — skipping guest teardown"
# host-level removal is still ours-gated: allowed if state's provisioned_vmid matches; else --force.
if [[ -n "$state_vmid" && "$state_vmid" == "$vmid" ]]; then
log_info " recorded provisioned_vmid matches $vmid — host-level removal permitted"
elif ! $FORCE; then
die "guest $vmid is absent and is not the recorded provisioned_vmid ('${state_vmid:-none}') — refusing host-level removal without --force."
fi
fi
# 3. Other-Felhom-guests detector — the safe default. If any OTHER Felhom guest remains and no
# --force, stop after the guest teardown and leave every host-level component in place.
local others others_csv
others=$(felhom_guests | grep -vx "$vmid" || true)
if [[ -n "$others" ]] && ! $FORCE; then
others_csv=$(echo "$others" | tr '\n' ' ' | sed 's/ */ /g;s/^ //;s/ $//;s/ /, /g')
echo ""
log_warn "Other Felhom guests remain (${others_csv}); leaving the agent + PVE token + state in place."
log_warn "Re-run --uninstall --force to remove host-level components anyway (this orphans ${others_csv})."
# GL-4 (Scenario B): NO umounts in guest-only mode — drives may serve the remaining guests.
_guest_drive_note "$vmid"
_uninstall_statement guest-only
log_success "UNINSTALL (guest-only) complete — removed guest $vmid; host-level components preserved."
log_info " NOTE: the host record still exists in the hub — remove it there if desired."
$DRY_RUN && log_warn " DRY-RUN: nothing above was actually executed."
return 0
fi
# ── host-level removal (reverse of install) ──────────────────────────────────────────────────
log_step "host-level removal"
# 4. Agent removal — service, unit(+.bak), sudoers, binary(+.bak), state dir, config, user.
# NEVER `sudo`. Resolve the agent config path BEFORE the unit is removed (mirrors preflight —
# the unit's -config arg is the truth, else the default); the config holds the per-host hub
# api_key and must not survive an uninstall (drill finding R1).
local agent_cfg=""
if systemctl cat felhom-agent >/dev/null 2>&1; then
agent_cfg=$(systemctl cat felhom-agent 2>/dev/null | grep -oP '(?<=-config )\S+' | head -1)
fi
[[ -n "$agent_cfg" ]] || agent_cfg="/etc/felhom-agent/agent.json"
if systemctl list-unit-files felhom-agent.service >/dev/null 2>&1; then
systemctl is-active --quiet felhom-agent 2>/dev/null && run systemctl stop felhom-agent
systemctl is-enabled --quiet felhom-agent 2>/dev/null && run systemctl disable felhom-agent
else
log_skip " felhom-agent unit not loaded — skip stop/disable"
fi
if [[ -f "$AGENT_UNIT" ]]; then run rm -f "$AGENT_UNIT"; else log_skip " $AGENT_UNIT already absent"; fi
local bak
for bak in "${AGENT_UNIT}".bak-*; do [[ -e "$bak" ]] && run rm -f "$bak"; done
run systemctl daemon-reload
if [[ -f "$AGENT_SUDOERS" ]]; then run rm -f "$AGENT_SUDOERS"; else log_skip " $AGENT_SUDOERS already absent"; fi
if [[ -f "$AGENT_BIN" ]]; then run rm -f "$AGENT_BIN"; else log_skip " $AGENT_BIN already absent"; fi
for bak in "${AGENT_BIN}".bak-*; do [[ -e "$bak" ]] && run rm -f "$bak"; done
if [[ -d "$AGENT_STATE_DIR" ]]; then run rm -rf "$AGENT_STATE_DIR"; else log_skip " $AGENT_STATE_DIR already absent"; fi
if id "$AGENT_USER" >/dev/null 2>&1; then run userdel "$AGENT_USER"; else log_skip " service user $AGENT_USER already absent"; fi
# 4b. Agent config + its backups (pve token + per-host hub api_key — secrets must not survive;
# drill R1 / GL-6 F1). The config write leaves `${agent_cfg}.bak*` siblings (e.g. .bak-<ver>,
# .bak-ceremony-*, .bak-pre064) — one GL-6 residue still held a LIVE hub api_key. Remove the
# config AND every `.bak*` sibling, then the (now-empty) dir. Paths logged, contents never.
if [[ -f "$agent_cfg" ]]; then run rm -f "$agent_cfg"; else log_skip " $agent_cfg already absent"; fi
local _cfgbak
for _cfgbak in "${agent_cfg}".bak*; do [[ -e "$_cfgbak" ]] && run rm -f "$_cfgbak"; done
run rmdir "$(dirname "$agent_cfg")" 2>/dev/null || true
# 4b2. Management-plane break-glass (TASK G1): timer+oneshot+script+tmpfiles. Stop/disable the
# timer, remove all four artifacts + the runtime heal-marker. We do NOT `rmdir /run/sshd` —
# the stock sshd needs it; leaving the (now unit-less) dir in place is correct (a bare kernel
# /run tmpfs recreates it empty on next boot anyway). Tolerate-absent throughout.
if systemctl list-unit-files felhom-mgmt-watchdog.timer >/dev/null 2>&1; then
systemctl is-active --quiet felhom-mgmt-watchdog.timer 2>/dev/null && run systemctl stop felhom-mgmt-watchdog.timer
systemctl is-enabled --quiet felhom-mgmt-watchdog.timer 2>/dev/null && run systemctl disable felhom-mgmt-watchdog.timer
else
log_skip " felhom-mgmt-watchdog.timer not loaded — skip stop/disable"
fi
run systemctl reset-failed felhom-mgmt-watchdog.service 2>/dev/null || true
local wda
for wda in /etc/systemd/system/felhom-mgmt-watchdog.service /etc/systemd/system/felhom-mgmt-watchdog.timer \
/usr/local/sbin/felhom-mgmt-watchdog /etc/tmpfiles.d/felhom-privsep.conf /run/felhom-mgmt-watchdog.healed; do
if [[ -e "$wda" ]]; then run rm -f "$wda"; fi
done
# 4b3. OOB felhom-sshd instance + belt (TASK H1). Stop/disable felhom-sshd + the belt loader,
# remove units + config tree + belt table + felhom-op user/sudoers. NEVER touch the stock
# sshd/:22/etc/ssh. Tolerate-absent throughout.
if systemctl list-unit-files felhom-sshd.service >/dev/null 2>&1; then
systemctl is-active --quiet felhom-sshd 2>/dev/null && run systemctl stop felhom-sshd
systemctl is-enabled --quiet felhom-sshd 2>/dev/null && run systemctl disable felhom-sshd
fi
if systemctl list-unit-files felhom-oob-nft.service >/dev/null 2>&1; then
systemctl is-active --quiet felhom-oob-nft 2>/dev/null && run systemctl stop felhom-oob-nft
systemctl is-enabled --quiet felhom-oob-nft 2>/dev/null && run systemctl disable felhom-oob-nft
fi
run systemctl reset-failed felhom-sshd felhom-oob-nft 2>/dev/null || true
nft list table inet felhom_oob >/dev/null 2>&1 && run nft delete table inet felhom_oob
local oa
for oa in /etc/systemd/system/felhom-sshd.service /etc/systemd/system/felhom-oob-nft.service \
/etc/felhom-oob.nft /etc/sudoers.d/felhom-op /run/felhom-sshd.pid /run/felhom-sshd.healed; do
if [[ -e "$oa" ]]; then run rm -f "$oa"; fi
done
if [[ -d /etc/felhom-sshd ]]; then run rm -rf /etc/felhom-sshd; fi
if id felhom-op >/dev/null 2>&1; then run userdel -r felhom-op 2>/dev/null || run userdel felhom-op; fi
# 4b4. Self-update artifacts (TASK D1; GL-4 gap-closure). Paths derived from the AUTHORITATIVE
# list in felhom-agent configs/felhom-selfupdate-guarded: the wrapper itself, the A/B slot
# files it creates next to the live binary (.prev snapshot + orphaned .new.* temps), plus
# the rollback unit + start-limit drop-in step 5 installs alongside it. pending.json lives
# under $AGENT_STATE_DIR (already removed in 4). Tolerate-absent throughout.
if systemctl list-unit-files felhom-agent-rollback.service >/dev/null 2>&1; then
systemctl is-active --quiet felhom-agent-rollback 2>/dev/null && run systemctl stop felhom-agent-rollback
systemctl is-enabled --quiet felhom-agent-rollback 2>/dev/null && run systemctl disable felhom-agent-rollback
fi
run systemctl reset-failed felhom-agent-rollback.service 2>/dev/null || true
local sua
for sua in /usr/local/sbin/felhom-selfupdate-guarded /etc/systemd/system/felhom-agent-rollback.service \
"${AGENT_BIN}.prev"; do
if [[ -e "$sua" ]]; then run rm -f "$sua"; fi
done
for sua in "${AGENT_BIN}".new.*; do [[ -e "$sua" ]] && run rm -f "$sua"; done
if [[ -d "${AGENT_UNIT}.d" ]]; then
if [[ -f "${AGENT_UNIT}.d/felhom-agent-limits.conf" ]]; then run rm -f "${AGENT_UNIT}.d/felhom-agent-limits.conf"; fi
run rmdir "${AGENT_UNIT}.d" 2>/dev/null || true
fi
# 4c. Shared-parent unit + wrapper + /mnt/felhom-drives (agent-installed at runtime; drill R2).
# Stop/disable, remove unit + script, unbind + remove the (empty) parent dir. Tolerate-absent.
if systemctl list-unit-files felhom-shared-parent.service 2>/dev/null | grep -q felhom-shared-parent; then
systemctl is-active --quiet felhom-shared-parent 2>/dev/null && run systemctl stop felhom-shared-parent
systemctl is-enabled --quiet felhom-shared-parent 2>/dev/null && run systemctl disable felhom-shared-parent
else
log_skip " felhom-shared-parent unit not loaded — skip stop/disable"
fi
if [[ -f /etc/systemd/system/felhom-shared-parent.service ]]; then run rm -f /etc/systemd/system/felhom-shared-parent.service; else log_skip " felhom-shared-parent.service already absent"; fi
if [[ -f /usr/local/sbin/felhom-shared-parent.sh ]]; then run rm -f /usr/local/sbin/felhom-shared-parent.sh; fi
run systemctl daemon-reload
# GL-4: unmount every enrolled/network drive mounted UNDER /mnt/felhom-drives (deepest first)
# BEFORE the root self-bind. Plain umount ONLY — NEVER -l/-f: a lazy/forced unmount on a busy
# data mount risks the customer's data; a busy mount gets a warning + guidance instead. The
# DATA STAYS ON THE DRIVE — nothing here (or anywhere in this script) wipes or formats it.
local dmnt
while IFS= read -r dmnt; do
[[ -n "$dmnt" ]] || continue
if $DRY_RUN; then
log_dry "umount $dmnt # data stays on the drive"
elif umount "$dmnt" 2>/dev/null; then
log_success " unmounted $dmnt (data stays on the drive)"
else
log_warn " $dmnt is busy — NOT forcing (never umount -l/-f). Eject the drive via the dashboard, or stop the apps using it and re-run."
_busy_mounts+=("$dmnt")
continue
fi
run rmdir "$dmnt" 2>/dev/null || true
done < <(findmnt -rn -o TARGET 2>/dev/null | grep '^/mnt/felhom-drives/' | sort -r || true)
if mountpoint -q /mnt/felhom-drives 2>/dev/null; then
if [[ ${#_busy_mounts[@]} -gt 0 ]]; then
log_warn " /mnt/felhom-drives root bind left mounted (busy drive mounts above must go first)"
else
run umount /mnt/felhom-drives
fi
fi
if [[ -d /mnt/felhom-drives ]]; then run rmdir /mnt/felhom-drives 2>/dev/null || true; fi
# 4d. Guarded-mkfs + pbs-apply wrappers, guest-hook snippet, lan-resolver dnsmasq snippets (drill R3-R5).
if [[ -f /usr/local/sbin/felhom-mkfs-guarded ]]; then run rm -f /usr/local/sbin/felhom-mkfs-guarded; else log_skip " felhom-mkfs-guarded already absent"; fi
if [[ -f /usr/local/sbin/felhom-pbs-apply ]]; then run rm -f /usr/local/sbin/felhom-pbs-apply; else log_skip " felhom-pbs-apply already absent"; fi
if [[ -f /usr/local/sbin/felhom-backup-target-apply ]]; then run rm -f /usr/local/sbin/felhom-backup-target-apply; else log_skip " felhom-backup-target-apply already absent"; fi
if [[ -f /var/lib/vz/snippets/felhom-guest-hook.sh ]]; then run rm -f /var/lib/vz/snippets/felhom-guest-hook.sh; fi
local dconf _dnsmasq_touched=false
for dconf in /etc/dnsmasq.d/felhom-*.conf; do
[[ -e "$dconf" ]] || continue
run rm -f "$dconf"; _dnsmasq_touched=true
done
if $_dnsmasq_touched && systemctl is-active --quiet dnsmasq 2>/dev/null; then
run systemctl restart dnsmasq || true
fi
# 5. pveum removal (presence-checked; tolerate-absent; roles deleted only after their grants).
# Remove the 3-role scoped grants+roles (3b) AND the pre-3b single-role broad grant if present —
# both tolerate-absent so --uninstall works on a box of either shape.
remove_scoped_acl
remove_old_broad_acl
# token then user (token-remove purges its ACL; user-delete purges anything else).
if pveum user token list "$PVE_USER" --output-format json 2>/dev/null | python3 -c "import json,sys;sys.exit(0 if any(t['tokenid']=='$PVE_TOKENID' for t in json.load(sys.stdin)) else 1)" 2>/dev/null; then
run pveum user token remove "$PVE_USER" "$PVE_TOKENID"
else
log_skip " token ${PVE_USER}!${PVE_TOKENID} already absent"
fi
if pveum user list --output-format json 2>/dev/null | python3 -c "import json,sys;sys.exit(0 if any(u['userid']=='$PVE_USER' for u in json.load(sys.stdin)) else 1)" 2>/dev/null; then
run pveum user delete "$PVE_USER"
else
log_skip " user $PVE_USER already absent"
fi
# 5b. felhom pool — delete ONLY if empty (a destroyed guest is auto-removed from its pool). Never
# delete a pool that still holds members (someone else's guests, or another Felhom guest kept
# under --force).
if pool_exists; then
local pool_left; pool_left=$(pool_members)
if [[ -z "$pool_left" ]]; then
run pveum pool delete "$PVE_POOL"
pool_removed=true
else
log_skip " pool $PVE_POOL not empty (members: $pool_left) — leaving it"
fi
else
log_skip " pool $PVE_POOL already absent"
fi
# 6. Golden vzdump (opt-in via --remove-golden; else left in place).
if $REMOVE_GOLDEN; then
local gvols gv
gvols=$(pvesm list "$ARCHIVE_STORAGE" --content backup 2>/dev/null | awk -v v="$GOLDEN_VMID" '$0 ~ ("vzdump-lxc-" v "-"){print $1}')
if [[ -n "$gvols" ]]; then
while IFS= read -r gv; do
[[ -n "$gv" ]] || continue
run pvesm free "$gv"
done <<<"$gvols"
log_success " removed golden vzdump(s) from $ARCHIVE_STORAGE"
else
log_skip " no golden vzdump (vzdump-lxc-${GOLDEN_VMID}-*) on $ARCHIVE_STORAGE"
fi
else
log_skip " golden vzdump left in place (pass --remove-golden to remove)"
fi
# 7. Install state file (only reached when host-level removal ran — safe-skip returned earlier).
if [[ -f "$STATE_FILE" ]]; then run rm -f "$STATE_FILE"; else log_skip " $STATE_FILE already absent"; fi
run rmdir "$STATE_DIR" 2>/dev/null || true
# 8. Summary + the GL-4 kept-vs-wiped statement.
_uninstall_statement full
log_success "UNINSTALL complete — removed: guest $vmid, the felhom-agent (unit/sudoers/binary/state/config+baks/user + selfupdate-artifacts/shared-parent/mkfs-wrapper/pbs-apply-wrapper/hook-snippet/dnsmasq-snippets), the pveum role/user/token/ACL,$( $pool_removed && printf ' the %s pool,' "$PVE_POOL") and $STATE_FILE."
if $REMOVE_GOLDEN; then log_info " golden vzdump: removed."; else log_info " golden vzdump: left in place (--remove-golden to remove)."; fi
log_info " NOTE: the 'sudo' and 'dnsmasq' packages were left installed (system packages); the host record still exists in the hub — remove it there if desired."
$DRY_RUN && log_warn " DRY-RUN: nothing above was actually executed."
return 0
}
# run_adopt_pool — retrofit an EXISTING Felhom guest into the felhom pool. Non-destructive: creates the
# pool if absent + adds the guest; never reconfigures/restarts the guest, never contacts the hub. Guest
# resolves from --vmid else the recorded provisioned_vmid (mirrors run_uninstall).
run_adopt_pool() {
log_step "ADOPT-POOL — add an existing Felhom guest to the $PVE_POOL pool"
local state_vmid vmid
state_vmid=$(_state_get provisioned_vmid)
if $VMID_EXPLICIT; then
vmid="$VMID"
elif [[ -n "$state_vmid" ]]; then
vmid="$state_vmid"
log_info " no --vmid given; using recorded provisioned_vmid=$vmid from $STATE_FILE"
else
die "pass --vmid N (state has no recorded vmid)"
fi
_vmid_in_use "$vmid" || die "guest $vmid not found on this host (nothing to adopt)"
# ours-check: only adopt a Felhom guest (has the /etc/felhom-bootstrap mount) unless --force.
if pct config "$vmid" 2>/dev/null | grep -q 'mp=/etc/felhom-bootstrap'; then
log_info " vmid $vmid looks like a Felhom guest (has the /etc/felhom-bootstrap mount)"
elif $FORCE; then
log_warn " vmid $vmid has NO /etc/felhom-bootstrap mount — --force given, adopting anyway"
else
die "vmid $vmid does not look like a Felhom-provisioned guest (no /etc/felhom-bootstrap mount).
Refusing to adopt. Pass --force to override."
fi
ensure_felhom_pool
pool_add_guest "$vmid"
echo ""
log_success "ADOPT-POOL complete — guest $vmid is in pool $PVE_POOL (guest not otherwise modified)."
$DRY_RUN && log_warn " DRY-RUN: nothing above was actually executed."
return 0
}
# run_rescope_acl — migrate an EXISTING install from the pre-3b broad-/ token to the pool-scoped ACL.
# Non-destructive to data (ACL-only): ensure the pool + the guest is a member, apply the 3-role scoped
# grants, THEN remove the old broad grant. Idempotent + dry-run-aware. Does NOT touch the guest or hub.
# ORDERING (see §13): run this with the agent STOPPED, then deploy agent >= v0.53.0, then start — the
# scoped ACL and the pool-param agent are mutually dependent.
run_rescope_acl() {
log_step "RESCOPE-ACL — migrate to the pool-scoped token ACL"
local state_vmid vmid
state_vmid=$(_state_get provisioned_vmid)
if $VMID_EXPLICIT; then
vmid="$VMID"
elif [[ -n "$state_vmid" ]]; then
vmid="$state_vmid"
log_info " no --vmid given; using recorded provisioned_vmid=$vmid from $STATE_FILE"
else
die "pass --vmid N (state has no recorded vmid)"
fi
_vmid_in_use "$vmid" || die "guest $vmid not found on this host"
if pct config "$vmid" 2>/dev/null | grep -q 'mp=/etc/felhom-bootstrap'; then
log_info " vmid $vmid looks like a Felhom guest (has the /etc/felhom-bootstrap mount)"
elif $FORCE; then
log_warn " vmid $vmid has NO /etc/felhom-bootstrap mount — --force given, rescoping anyway"
else
die "vmid $vmid does not look like a Felhom-provisioned guest (no /etc/felhom-bootstrap mount).
Refusing to rescope. Pass --force to override."
fi
# The guest MUST be a pool member before the scoped token can touch it — ensure it first.
ensure_felhom_pool
pool_add_guest "$vmid"
# Apply the scoped grants, THEN remove the old broad grant (add-before-remove: never leave the token
# with NO grant mid-migration).
apply_scoped_acl
remove_old_broad_acl
echo ""
log_success "RESCOPE-ACL complete — token scoped to /pool/$PVE_POOL + /storage/[${PVE_STORAGES[*]}] + Sys.Audit/SDN.Use@/."
log_warn " NOW deploy agent >= v0.53.0 (restore-into-pool) and (re)start felhom-agent — the scoped ACL needs it."
$DRY_RUN && log_warn " DRY-RUN: nothing above was actually executed."
return 0
}
#-------------------------------------------------------------------------------
# Arg parse
#-------------------------------------------------------------------------------
while [[ $# -gt 0 ]]; do
case "$1" in
--customer-id) CUSTOMER_ID="$2"; shift 2 ;;
--mode) MODE="$2"; shift 2 ;;
--hub-url) HUB_URL="$2"; shift 2 ;;
--vmid) VMID="$2"; VMID_EXPLICIT=true; shift 2 ;;
--golden) GOLDEN_VOLID="$2"; shift 2 ;;
--golden-vmid) GOLDEN_VMID="$2"; shift 2 ;;
--archive-storage) ARCHIVE_STORAGE="$2"; shift 2 ;;
--node) NODE="$2"; NODE_EXPLICIT=true; shift 2 ;;
--bridge-ip) BRIDGE_ADDR="$2"; shift 2 ;;
--no-island) WANT_ISLAND=false; shift ;; # R-50: keep the historical LAN bind on an appliance
--rootfs-grow) ROOTFS_GROW="$2"; shift 2 ;;
--datavol-grow) DATAVOL_GROW="$2"; shift 2 ;;
--sysdata-grow) SYSDATA_GROW="$2"; shift 2 ;;
--cores) CPU_CORES="$2"; shift 2 ;;
--memory) MEM_MIB="$2"; shift 2 ;;
--passphrase-file) PASSPHRASE_FILE="$2"; shift 2 ;;
--operator-pubkey-file) OPERATOR_PUBKEY_FILE="$2"; shift 2 ;;
--preserve-from) PRESERVE_FROM="$2"; shift 2 ;;
--preserve-state-from) PRESERVE_STATE_FROM="$2"; shift 2 ;;
--allow-new-leaf) ALLOW_NEW_LEAF=true; shift ;;
--force) FORCE=true; shift ;;
--force-gitea-golden) FORCE_GITEA_GOLDEN=true; shift ;;
--skip-provision) SKIP_PROVISION=true; shift ;;
--uninstall) UNINSTALL=true; shift ;;
--remove-golden) REMOVE_GOLDEN=true; shift ;;
--adopt-pool) ADOPT_POOL=true; shift ;;
--rescope-acl) RESCOPE_ACL=true; shift ;;
--rotate-recovery) ROTATE_RECOVERY=true; shift ;;
--enable-oob) ENABLE_OOB=true; shift ;;
--no-oob) NO_OOB=true; shift ;;
--acl-storages) read -ra PVE_STORAGES <<< "$2"; shift 2 ;;
--preflight-only) PREFLIGHT_ONLY=true; shift ;;
--dry-run) DRY_RUN=true; shift ;;
--resume) RESUME=true; shift ;;
-h|--help) usage ;;
*) die "Unknown option: $1 (use -h)" ;;
esac
done
#===============================================================================
# UNINSTALL MODE — local host teardown (no hub contact, no passphrase). Dispatched early,
# before any provision/DR logic, and does not require --customer-id.
#===============================================================================
if $UNINSTALL; then
[[ $EUID -eq 0 ]] || die "must run as root"
echo ""
log_info "felhom-host-install v${SCRIPT_VERSION} — mode=uninstall"
$DRY_RUN && log_warn "DRY-RUN: no mutations will be performed"
echo ""
run_uninstall
exit 0
fi
#===============================================================================
# ADOPT-POOL MODE — retrofit an EXISTING Felhom guest into the felhom pool (non-destructive; no hub
# contact, no passphrase, no guest reconfigure beyond pool membership). Dispatched early.
#===============================================================================
if $ADOPT_POOL; then
[[ $EUID -eq 0 ]] || die "must run as root"
echo ""
log_info "felhom-host-install v${SCRIPT_VERSION} — mode=adopt-pool"
$DRY_RUN && log_warn "DRY-RUN: no mutations will be performed"
echo ""
run_adopt_pool
exit 0
fi
#===============================================================================
# RESCOPE-ACL MODE — migrate an existing install to the pool-scoped token ACL (ACL-only, no hub, no
# passphrase). Supervised: run with the agent stopped, then deploy agent >= v0.53.0 (see §13).
#===============================================================================
if $RESCOPE_ACL; then
[[ $EUID -eq 0 ]] || die "must run as root"
echo ""
log_info "felhom-host-install v${SCRIPT_VERSION} — mode=rescope-acl"
$DRY_RUN && log_warn "DRY-RUN: no mutations will be performed"
echo ""
run_rescope_acl
exit 0
fi
#===============================================================================
# DR MODE — documented seam only (10D). NOT implemented.
#===============================================================================
if [[ "$MODE" == "dr" ]]; then
log_error "DR mode not yet implemented (10D)."
cat >&2 <<'EOF'
The DR step skeleton (for the future implementer) mirrors provision EXCEPT
the restore source:
1. pre-flight (root, PVE, hub reachable, customer+passphrase valid)
2. pveum token (identical to provision)
3. host-enroll (mint-once-reuse — the lost host re-binds to its customer)
4. agent config write (identical)
5. RESTORE: instead of the golden, restore the customer's OWN whole-CT PBS
snapshot (continuity preserved) — agent --selftest=bring-up -mode dr
-archive <pbs-snapshot>. Identity/keys come from escrow + the hub recipe.
6. verify (identical)
EOF
exit 2
fi
#===============================================================================
# GL-2 install-profile validation — everything below runs a real install (or its --preflight-only /
# --resume variant), so the profile is decided HERE, before the passphrase prompt and any step. The
# refusals are deliberately argv-time (the harness relies on them firing on a non-PVE machine too).
#===============================================================================
# --- F5 (VALIDATION-n100): appliance guest auto-sizing --------------------------
# The golden default (2 GB RAM) was too small on a 16 GB host. In APPLIANCE mode, when the operator
# passes no explicit cap, size the guest generously from the host — LXC limits are cheap, so err
# generous. An explicit --cores/--memory ALWAYS wins (never overwritten). byo mode requires explicit
# caps (never auto-sized). The host reads are overridable for the mode harness (no real /proc dep).
host_total_mib() {
if [[ -n "${FELHOM_FAKE_MEMTOTAL_MIB:-}" ]]; then echo "$FELHOM_FAKE_MEMTOTAL_MIB"; return; fi
awk '/^MemTotal:/{print int($2/1024)}' /proc/meminfo 2>/dev/null || echo 0
}
host_core_count() {
if [[ -n "${FELHOM_FAKE_NPROC:-}" ]]; then echo "$FELHOM_FAKE_NPROC"; return; fi
nproc 2>/dev/null || echo 0
}
autosize_guest_caps() {
[[ "$MODE" == "appliance" ]] || return 0
if [[ -z "$MEM_MIB" ]]; then
local total; total=$(host_total_mib)
if [[ "${total:-0}" -gt 0 ]]; then
# clamp(host-4096, min 4096, max host-2048), then the hard ceiling host-1024 (never
# over-commit). Order matters: the min floor is applied AFTER the generous max so that on
# a host too small for both the MIN wins (edge rule), bounded only by the host-1024 ceiling.
local mem=$(( total - 4096 ))
local max2=$(( total - 2048 ))
(( mem > max2 )) && mem=$max2
(( mem < 4096 )) && mem=4096
local ceil=$(( total - 1024 ))
(( mem > ceil )) && mem=$ceil
MEM_MIB=$mem
log_info " auto-sized guest RAM: ${MEM_MIB} MiB (host ${total} MiB; clamp(host-4096, min 4096, max host-2048), ceiling host-1024)"
fi
fi
if [[ -z "$CPU_CORES" ]]; then
local cores; cores=$(host_core_count)
if [[ "${cores:-0}" -gt 0 ]]; then
local c=$(( cores - 1 ))
(( c < 2 )) && c=2
CPU_CORES=$c
log_info " auto-sized guest cores: ${CPU_CORES} (host ${cores} cores; host-1, min 2)"
fi
fi
}
case "$MODE" in
appliance|byo) ;;
"")
if $PREFLIGHT_ONLY; then
log_warn "--preflight-only without --mode: running the mode-agnostic base checks only"
else
die "--mode is required: pass --mode appliance (a Felhom-owned box) or --mode byo (a host you do not own). See -h."
fi
;;
provision) die "--mode provision was retired in v1.10.0 — pass --mode appliance (Felhom-owned box) or --mode byo (a host you do not own)." ;;
*) die "Unknown --mode: $MODE (appliance|byo)" ;;
esac
# F5: fill the appliance guest caps from the host when the operator gave none (explicit flags win).
autosize_guest_caps
# BYO argument refusals (C1/C2) — before the passphrase prompt, before any step.
if [[ "$MODE" == "byo" ]]; then
if [[ -z "$CPU_CORES" || -z "$MEM_MIB" ]]; then
die "byo mode requires explicit --cores and --memory (noisy-neighbor caps on a host you do not own)"
fi
if $ENABLE_OOB; then
die "--enable-oob is not allowed in byo mode (OOB access on a host the operator does not own is the owner's call; provision it as a separate, owner-consented step)"
fi
if $ROTATE_RECOVERY; then
die "--rotate-recovery is not allowed in byo mode (byo never touches root@pam — there is no recovery credential to rotate)"
fi
fi
# H1 / F9 (ruling 2026-07-23 "install everywhere"): the OOB belt is now a DEFAULT appliance leg — a box
# the operator owns must have an out-of-band way in (the 07-22 power-outage audit found the belt on
# NEITHER fleet box because the universal ISO never passed --enable-oob). Appliance installs it unless
# --no-oob; byo stays refuse-by-default above (owner-consented, never automatic). The belt is
# lockout-SAFE if the fleet has no operator peer yet: empty sets + no felhom-sshd key = no access
# granted and nothing on :22 touched, until the operator OOB peer is registered hub-side.
if [[ "$MODE" == "appliance" ]] && ! $NO_OOB; then
ENABLE_OOB=true
fi
# GL-4: resolve the operator signing keys — script constants by default; --operator-pubkey-file
# OVERRIDES them (C5). Validated HERE (before the passphrase prompt, before any step) so a bad key
# file dies fast and the harness can exercise it on a non-PVE machine. File format: one key per
# line, "operational <authorized_keys line>" or "recovery <authorized_keys line>"; '#' comments and
# blank lines allowed. The key_id is the authorized_keys COMMENT field — required (the agent's
# signed-op verifier addresses keys by key_id).
RESOLVED_OP_ID="$OPERATOR_KEY_OPERATIONAL_ID"; RESOLVED_OP_LINE="$OPERATOR_KEY_OPERATIONAL_LINE"
RESOLVED_REC_ID="$OPERATOR_KEY_RECOVERY_ID"; RESOLVED_REC_LINE="$OPERATOR_KEY_RECOVERY_LINE"
if [[ -n "$OPERATOR_PUBKEY_FILE" ]]; then
[[ -f "$OPERATOR_PUBKEY_FILE" ]] || die "--operator-pubkey-file not found: $OPERATOR_PUBKEY_FILE"
if [[ -n "$RESOLVED_OP_LINE$RESOLVED_REC_LINE" ]]; then
log_info "--operator-pubkey-file overrides the script's built-in operator key constants"
fi
RESOLVED_OP_ID=""; RESOLVED_OP_LINE=""; RESOLVED_REC_ID=""; RESOLVED_REC_LINE=""
_kseen=false
while IFS= read -r _kline || [[ -n "$_kline" ]]; do
_kline="${_kline%$'\r'}"
[[ -z "$_kline" || "$_kline" == \#* ]] && continue
_krole="${_kline%% *}"; _krest="${_kline#* }"
case "$_krole" in
operational|recovery) ;;
*) die "--operator-pubkey-file: unknown role '$_krole' (want operational|recovery) in line: $_kline" ;;
esac
read -r _ktype _kb64 _kid _ <<<"$_krest"
[[ "$_ktype" == ssh-* || "$_ktype" == sk-ssh-* || "$_ktype" == ecdsa-* ]] \
|| die "--operator-pubkey-file: not an authorized_keys line (bad key type '${_ktype:-<empty>}') in line: $_kline"
[[ -n "$_kb64" && "$_kb64" =~ ^[A-Za-z0-9+/=]+$ ]] \
|| die "--operator-pubkey-file: not an authorized_keys line (missing/invalid key material) in line: $_kline"
[[ -n "$_kid" ]] \
|| die "--operator-pubkey-file: key line has no comment field — the comment IS the key_id; add one (e.g. felhom-op-1): $_kline"
if [[ "$_krole" == operational ]]; then
RESOLVED_OP_ID="$_kid"; RESOLVED_OP_LINE="$_krest"
else
RESOLVED_REC_ID="$_kid"; RESOLVED_REC_LINE="$_krest"
fi
_kseen=true
done < "$OPERATOR_PUBKEY_FILE"
$_kseen || die "--operator-pubkey-file has no key lines (empty/comments only): $OPERATOR_PUBKEY_FILE"
fi
# Resume mode-mismatch (C4): an install keeps the mode it started with — mode-flipping a half-done
# install would skip already-completed steps under the WRONG profile's guarantees.
if $RESUME; then
_recorded_mode=$(_state_get mode || true)
if [[ -n "$_recorded_mode" && "$_recorded_mode" != "$MODE" ]]; then
die "install started as $_recorded_mode; resume with --mode $_recorded_mode or start over"
fi
fi
#===============================================================================
# PROVISION MODE
#===============================================================================
[[ -n "$CUSTOMER_ID" ]] || die "--customer-id is required (use -h)"
echo ""
log_info "felhom-host-install v${SCRIPT_VERSION} — mode=${MODE:-<preflight-only, no mode>} customer=${CUSTOMER_ID} vmid=${VMID}"
$DRY_RUN && log_warn "DRY-RUN: no mutations will be performed"
echo ""
#-------------------------------------------------------------------------------
# Read passphrase (no-echo prompt or 0600 file) — never on argv/logs
#-------------------------------------------------------------------------------
read_passphrase() {
if [[ -n "$PASSPHRASE_FILE" ]]; then
[[ -f "$PASSPHRASE_FILE" ]] || die "--passphrase-file not found: $PASSPHRASE_FILE"
local perm; perm=$(stat -c '%a' "$PASSPHRASE_FILE")
[[ "$perm" == "600" || "$perm" == "400" ]] || log_warn "passphrase file $PASSPHRASE_FILE is mode $perm (want 600)"
PASSPHRASE="$(< "$PASSPHRASE_FILE")"; PASSPHRASE="${PASSPHRASE%$'\n'}"
else
# Read from the terminal explicitly (not stdin), so the no-echo prompt works whether the
# script is run from a file OR piped to bash (curl … | sudo bash) — where stdin is the pipe.
read -rsp "Retrieval passphrase for customer '${CUSTOMER_ID}': " PASSPHRASE < /dev/tty; echo ""
fi
[[ -n "$PASSPHRASE" ]] || die "empty passphrase"
}
#-------------------------------------------------------------------------------
# BYO host-mutation disclosure + typed acknowledgement (GL-2). Printed at the END of preflight so
# every value shown (vmid, storages, config path) is final. The list follows the CODE (steps 2-8 +
# the agent's runtime installs), not memory — keep it in sync when a step grows a new host artifact.
# The ack reads /dev/tty (the script is often `curl | bash`) and mirrors the uninstall confirm's
# dry-run branch. NO credential is ever printed here — artifact PATHS only.
#-------------------------------------------------------------------------------
_byo_disclosure_ack() {
echo ""
log_step "BYO disclosure — installing Felhom on a host the operator does NOT own will create:"
cat <<EOF
pveum: roles ${PVE_ROLE_BASE}/${PVE_ROLE_GUEST}/${PVE_ROLE_STORE} + user ${PVE_USER} + token ${PVE_USER}!${PVE_TOKENID}
+ scoped ACL grants (Base@/, Guest@/pool/${PVE_POOL}, Store@[${PVE_STORAGES[*]}]) + the '${PVE_POOL}' pool
agent: service user ${AGENT_USER} + ${AGENT_BIN} (+ .bak-* backups; A/B slots via the guarded self-update)
+ ${AGENT_UNIT} (+ felhom-agent-rollback.service + start-limit drop-in)
+ ${AGENT_SUDOERS} + state dir ${AGENT_STATE_DIR} + config ${AGENT_CONFIG} (0600)
host: /usr/local/sbin/felhom-mkfs-guarded + /usr/local/sbin/felhom-selfupdate-guarded
+ /usr/local/sbin/felhom-pbs-apply (PBS-DR apply wrapper — DR capability is baked
on every install; ACTIVATION stays a hub flag, off = zero effect on this host)
+ felhom-mgmt-watchdog service+timer+script + /etc/tmpfiles.d/felhom-privsep.conf
+ guest-hook snippet under /var/lib/vz/snippets/ (agent-installed at runtime)
+ the 'sudo' and 'age' packages if absent + install state dir ${STATE_DIR}
wg: an OUTBOUND WireGuard tunnel to the Felhom hub (wg_tunnel.enabled=true — base
infrastructure like the cloudflared tunnel; hands-free peer registration; the
operator can disable it per-customer hub-side on request)
EOF
if $SKIP_PROVISION; then
echo " guest: none (--skip-provision)"
else
echo " guest: the provisioned Felhom LXC (vmid ${VMID}, capped ${CPU_CORES} cores / ${MEM_MIB} MiB) + its volumes"
echo " + the golden vzdump imported onto storage '${ARCHIVE_STORAGE}'"
fi
echo " update: operator-signed self-update authority: ${RESOLVED_OP_ID:-NONE (self-update stays dormant)}"
echo " NOT touched in byo mode: root@pam (no break-glass), host DNS (:53), OOB sshd."
echo ""
if $DRY_RUN; then
log_dry "would prompt: Type this host's short hostname ($(hostname -s)) to acknowledge the byo install"
elif $PREFLIGHT_ONLY; then
log_info " (--preflight-only: acknowledgement prompt skipped — nothing will be mutated)"
else
local _hn _ans; _hn=$(hostname -s)
read -rp "Type this host's short hostname (${_hn}) to acknowledge the byo install: " _ans < /dev/tty
[[ "$_ans" == "$_hn" ]] || die "acknowledgement mismatch (got '${_ans}', expected '${_hn}') — aborting, nothing installed"
log_success " byo install acknowledged for host ${_hn}"
fi
}
#-------------------------------------------------------------------------------
# STEP 1 — pre-flight (fail fast before any mutation)
#-------------------------------------------------------------------------------
step_preflight() {
log_step "1/8 pre-flight"
[[ $EUID -eq 0 ]] || die "must run as root"
command -v pveum >/dev/null || die "pveum not found — is this a Proxmox VE host?"
command -v pct >/dev/null || die "pct not found — is this a Proxmox VE host?"
command -v pvesh >/dev/null || die "pvesh not found"
command -v curl >/dev/null || die "curl not found"
command -v python3>/dev/null || die "python3 not found"
local pvever; pvever=$(pveversion | head -1)
# GL-2 PVE-major gate: byo installs are validated on PVE 9.x ONLY — on a host we don't own, an
# unvalidated major is a refusal, not a warning. Appliance keeps the historical warn-only.
if [[ "$pvever" != *"/9."* && "$MODE" == "byo" ]] && ! $FORCE; then
die "byo mode is validated on PVE 9.x only (got: $pvever). Pass --force to override deliberately."
fi
[[ "$pvever" == *"/9."* ]] || log_warn "expected PVE 9.x, got: $pvever"
log_info " $pvever"
# node — on a MULTI-NODE cluster, auto-selecting nodes[0] is a wrong-node footgun. Require an
# explicit --node unless there is exactly one node (or no guest will be provisioned).
local nodes_json node_count node_names
nodes_json=$(pvesh get /nodes --output-format json 2>/dev/null || echo "[]")
node_count=$(python3 -c "import json,sys;print(len(json.loads(sys.argv[1])))" "$nodes_json" 2>/dev/null || echo 0)
if [[ "${node_count:-0}" -gt 1 ]] && ! $NODE_EXPLICIT && ! $SKIP_PROVISION; then
node_names=$(python3 -c "import json,sys;print(', '.join(n['node'] for n in json.loads(sys.argv[1])))" "$nodes_json" 2>/dev/null || echo "?")
die "this is a ${node_count}-node cluster (${node_names}); pass --node explicitly — auto-selecting nodes[0] risks provisioning on the wrong node."
fi
if [[ -z "$NODE" ]]; then
NODE=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])[0]['node'])" "$nodes_json" 2>/dev/null || hostname)
fi
if $NODE_EXPLICIT; then log_info " node: $NODE (explicit)"; else log_info " node: $NODE (auto)"; fi
# agent config path: honor the existing systemd unit's -config, else default
if systemctl cat felhom-agent >/dev/null 2>&1; then
AGENT_CONFIG=$(systemctl cat felhom-agent 2>/dev/null | grep -oP '(?<=-config )\S+' | head -1)
fi
[[ -n "$AGENT_CONFIG" ]] || AGENT_CONFIG="/etc/felhom-agent/agent.json"
log_info " agent config: $AGENT_CONFIG"
# v1.1.0: the agent binary is no longer a prerequisite — the agent-install step (5/8) fetches it
# from Gitea + verifies it. Just report what's present (if anything).
if command -v felhom-agent >/dev/null 2>&1; then
log_info " agent (existing): $(felhom-agent --version 2>&1 | head -1)"
else
log_info " agent: not installed yet — will be fetched + installed in step 5/8"
fi
# local-lvm free space
local free_gib
free_gib=$(lvs --noheadings --units g -o lv_size,data_percent /dev/pve/data 2>/dev/null | awk '{gsub(/[^0-9.]/,"",$1); used=$2; print int($1*(100-used)/100)}' 2>/dev/null || echo 0)
if [[ "${free_gib:-0}" -gt 0 ]]; then
log_info " local-lvm free: ~${free_gib} GiB"
[[ "$free_gib" -ge "$HARD_MIN_LVM_GIB" ]] || log_warn "local-lvm free ~${free_gib} GiB < hard min ${HARD_MIN_LVM_GIB} GiB"
else
log_warn " could not read local-lvm free space (continuing)"
fi
# RAM floor (soft): a big appliance guest can pressure existing guests on a small box. WARN only.
local mem_avail_mib
mem_avail_mib=$(awk '/^MemAvailable:/{print int($2/1024)}' /proc/meminfo 2>/dev/null || echo 0)
if [[ "${mem_avail_mib:-0}" -gt 0 ]]; then
if [[ "$mem_avail_mib" -lt 2048 ]]; then
log_warn " low free RAM (~${mem_avail_mib} MiB); the appliance guest may pressure existing guests — consider the CPU/mem cap."
else
log_info " free RAM: ~${mem_avail_mib} MiB"
fi
fi
# Appliance-cap sanity (soft): a cap that EXCEEDS host resources won't protect other guests. WARN,
# never die — the operator may know better (e.g. capping below a future hardware upgrade).
if ! $SKIP_PROVISION; then
if [[ -n "$CPU_CORES" ]]; then
local host_cores; host_cores=$(nproc 2>/dev/null || echo 0)
if [[ "${host_cores:-0}" -gt 0 && "$CPU_CORES" -gt "$host_cores" ]]; then
log_warn " requested cap (${CPU_CORES} cores) exceeds host cores (${host_cores}); the cap won't protect other guests."
fi
fi
if [[ -n "$MEM_MIB" ]]; then
local host_mem_mib; host_mem_mib=$(awk '/^MemTotal:/{print int($2/1024)}' /proc/meminfo 2>/dev/null || echo 0)
if [[ "${host_mem_mib:-0}" -gt 0 && "$MEM_MIB" -gt "$host_mem_mib" ]]; then
log_warn " requested cap (${MEM_MIB} MiB) exceeds host RAM (~${host_mem_mib} MiB); the cap won't protect other guests."
fi
fi
fi
# Existing-guest context (GL-2, both modes): cheap situational awareness — on a byo/shared box
# this is the population the caps + scoped ACL protect.
local _guest_count; _guest_count=$( { used_vmids || true; } | wc -l)
log_info " existing guests on this host: ${_guest_count} (pct+qm)"
# archive-storage-exists guard (provision only — the golden lives there + the restore reads it).
if ! $SKIP_PROVISION; then
if pvesm status --storage "$ARCHIVE_STORAGE" >/dev/null 2>&1; then
log_info " archive storage '$ARCHIVE_STORAGE' present"
else
die "archive storage '$ARCHIVE_STORAGE' not found (pvesm status). Pass --archive-storage NAME."
fi
fi
# --acl-storages existence (GL-2): the scoped ACL grants Datastore.* at each of these paths — a
# grant on a storage this box doesn't have silently protects nothing (campaign-2 R2 class). byo
# refuses; appliance warns. EXCEPTION (v1.15.0, drill retrofit-ACL lesson): `felhom-pbs` is
# EXPECTED absent on a fresh box — the PBS-DR tier apply creates it later, and the grant MUST
# be pre-positioned or the apply-bridge 403s (Datastore.Allocate on /storage/felhom-pbs — the
# drill hit exactly this after narrowing --acl-storages). Do NOT narrow the set to silence
# the info line.
local _missing_storages=() _acls
for _acls in "${PVE_STORAGES[@]}"; do
if ! pvesm status --storage "$_acls" >/dev/null 2>&1; then
if [[ "$_acls" == "felhom-pbs" ]]; then
log_info " acl storage 'felhom-pbs' not present yet — expected: the PBS-DR tier creates it; the grant is pre-positioned deliberately"
else
_missing_storages+=("$_acls")
fi
fi
done
if [[ ${#_missing_storages[@]} -gt 0 ]]; then
if [[ "$MODE" == "byo" ]]; then
die "acl storage(s) not found on this box: ${_missing_storages[*]} — pass --acl-storages with this box's real storages"
fi
for _acls in "${_missing_storages[@]}"; do
log_warn " acl storage '$_acls' not found (pvesm status) — its grant will point at nothing"
done
elif [[ "$MODE" == "byo" ]]; then
log_info " acl storages all present or pre-positioned: ${PVE_STORAGES[*]}"
fi
# Host DNS :53 gate (byo only) — GL-8/F6. In byo the agent's own lan-resolver stays OFF (a config
# assert enforces it), so ANY process bound to :53 is the OWNER's — and Felhom needs the guest
# reachable by name on the LAN. Felhom must NOT stop/mask/kill a service on a host it does not own
# (the dnsmasq analog of the break-glass rule), so this REFUSES with the remediation instead of
# mutating anything. `ss` failure/absence degrades to a warn (never block on a missing tool).
# Appliance mode is untouched — there the agent's lan_resolver owns :53 by design.
if [[ "$MODE" == "byo" ]]; then
local _dns53
if _dns53=$(ss -H -ltnup 'sport = :53' 2>/dev/null); then
if [[ -n "$_dns53" ]]; then
log_error " a resolver is already bound to :53 on this host:"
echo "$_dns53" | tr -s '[:space:]' ' ' | cut -c1-200 | sed 's/^/ /' >&2
die "a resolver is already bound to :53 on this host — Felhom needs the guest reachable by name on your LAN.
Stop or reconfigure that resolver, OR point your LAN DNS at the guest's address, then re-run.
(Felhom does NOT touch DNS services on a host it does not own — this is a refusal, not a change.)"
fi
log_info " host DNS (:53): free"
else
log_warn " host DNS (:53): could not query (ss failed/absent) — skipping the :53 gate"
fi
fi
# hub reachable
local hc; hc=$(http_code "$HUB_URL/api/v1/config/$CUSTOMER_ID" -H "X-Retrieval-Password: preflight-no-op" || echo 000)
[[ "$hc" != "000" ]] || die "hub unreachable at $HUB_URL"
log_info " hub reachable ($HUB_URL)"
# customer + passphrase valid (read-only GET /config/{id})
read_passphrase
local code; code=$(http_code "$HUB_URL/api/v1/config/$CUSTOMER_ID" -H "X-Retrieval-Password: $PASSPHRASE")
case "$code" in
200) log_success " customer '$CUSTOMER_ID' exists + passphrase valid" ;;
401) die "passphrase REJECTED (401) for customer '$CUSTOMER_ID'" ;;
404) die "customer '$CUSTOMER_ID' not found in hub (404) — create it in the hub first" ;;
*) die "unexpected hub status $code on config preflight" ;;
esac
# golden archive — auto-discover a LOCAL one for info; the golden step (7/8) ensures one exists
# (local else Gitea-fetched + verified), so a missing local golden is no longer fatal here.
if [[ -z "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then
GOLDEN_VOLID=$(pvesm list "$ARCHIVE_STORAGE" --content backup 2>/dev/null | awk -v v="$GOLDEN_VMID" '$0 ~ ("vzdump-lxc-" v "-"){print $1}' | sort | tail -1)
fi
if [[ -n "$GOLDEN_VOLID" ]]; then
pvesm list "$ARCHIVE_STORAGE" --content backup 2>/dev/null | grep -q "$(basename "$GOLDEN_VOLID")" || die "golden volid not resolvable: $GOLDEN_VOLID"
log_info " golden (local): $GOLDEN_VOLID"
else
log_info " golden: none local — will fetch + verify from Gitea in step 7/8"
fi
# vmid guard (irrelevant when --skip-provision: we never touch a guest). "In use" is checked
# against the pct+qm id-set (LXC and VMs share the space), not just `pct status`.
if $SKIP_PROVISION; then
log_info " --skip-provision: agent install/config only, no guest will be provisioned"
elif _vmid_in_use "$VMID"; then
if $VMID_EXPLICIT; then
# Explicit --vmid stays deterministic: die unless --force (which over-provisions, destructive).
if $FORCE; then
log_warn " vmid $VMID already exists — --force given, it WILL be destroyed by provision"
else
die "vmid $VMID already exists. Refusing to clobber a live guest. Pass --force to provision over it."
fi
elif $FORCE; then
# Default vmid + --force: honor the destructive over-provision without prompting.
log_warn " vmid $VMID already exists — --force given, it WILL be destroyed by provision"
else
# Default vmid in use, no --force: auto-pick the next free id and CONFIRM (never silent).
local free_vmid; free_vmid=$(next_free_vmid "$VMID")
log_info " vmid $VMID is in use; next free vmid is $free_vmid"
local ans; read -rp "VMID $VMID is in use. Use next free VMID $free_vmid? [y/N] " ans < /dev/tty
[[ "$ans" == "y" || "$ans" == "Y" ]] || die "no free vmid confirmed"
VMID="$free_vmid"
log_success " using auto-selected vmid $VMID"
fi
fi
# BYO disclosure + typed ack — last, so every disclosed value (vmid incl. auto-pick) is final.
if [[ "$MODE" == "byo" ]]; then
_byo_disclosure_ack
fi
# Record the customer + install profile into the state (foundation for a later automatic
# --uninstall; the mode key backs the C4 resume-mismatch refusal). --preflight-only writes NO
# state — it must stay re-runnable and must never let a later --resume skip the real preflight.
if ! $PREFLIGHT_ONLY; then
_state_put customer_id "$CUSTOMER_ID"
_state_put mode "$MODE"
_state_mark preflight
fi
log_success "pre-flight passed"
}
#-------------------------------------------------------------------------------
# STEP 2 — Proxmox API token (idempotent pveum; reuse-if-working else rotate)
#-------------------------------------------------------------------------------
step_token() {
log_step "2/8 Proxmox API token"
# Fast-path only when the token is ALREADY in memory this process (avoids a double reuse-check if
# called twice). On --resume PVE_TOKEN starts empty, so this correctly re-derives it from the
# on-disk config (reuse) or a rotation — GL6-F4. Uses _state_has (no misleading SKIP log here).
if _state_has token && [[ -n "$PVE_TOKEN" ]]; then return 0; fi
# Pool BEFORE the ACL: /pool/felhom must exist before apply_scoped_acl grants on it (3b). Always —
# even under --skip-provision (the token exists now; a later provision-into-pool needs pool + grant).
# The 3 scoped roles + grants are created by apply_scoped_acl below (AFTER the token exists).
ensure_felhom_pool
# user: tolerate-exists
if pveum user list --output-format json 2>/dev/null | python3 -c "import json,sys;sys.exit(0 if any(u['userid']=='$PVE_USER' for u in json.load(sys.stdin)) else 1)"; then
log_info " user $PVE_USER exists"
else
run pveum user add "$PVE_USER"
fi
# token: reuse if the existing agent config token still authenticates, else rotate
local reused=false
if [[ -f "$AGENT_CONFIG" ]] && python3 -c "import json,sys;d=json.load(open('$AGENT_CONFIG'));sys.exit(0 if d.get('proxmox',{}).get('token') else 1)" 2>/dev/null; then
log_info " existing agent config has a token — testing it (read-only --selftest)"
if felhom-agent --config "$AGENT_CONFIG" --selftest >/dev/null 2>&1; then
log_success " existing token authenticates — REUSING (no rotation)"
PVE_TOKEN=$(python3 -c "import json;print(json.load(open('$AGENT_CONFIG'))['proxmox']['token'])")
reused=true
else
log_warn " existing token failed selftest — will rotate"
fi
fi
if ! $reused; then
if $DRY_RUN; then
log_dry "pveum user token remove $PVE_USER $PVE_TOKENID # if present"
log_dry "pveum user token add $PVE_USER $PVE_TOKENID --privsep 1 --output-format json # capture .value"
PVE_TOKEN="${PVE_USER}!${PVE_TOKENID}=<DRY-RUN-SECRET>"
else
if pveum user token list "$PVE_USER" --output-format json 2>/dev/null | python3 -c "import json,sys;sys.exit(0 if any(t['tokenid']=='$PVE_TOKENID' for t in json.load(sys.stdin)) else 1)"; then
log_info " removing stale token $PVE_TOKENID (secret unrecoverable — rotating)"
pveum user token remove "$PVE_USER" "$PVE_TOKENID"
fi
local secret
secret=$(pveum user token add "$PVE_USER" "$PVE_TOKENID" --privsep 1 --output-format json | python3 -c "import json,sys;print(json.load(sys.stdin)['value'])")
[[ -n "$secret" ]] || die "failed to capture new token secret"
PVE_TOKEN="${PVE_USER}!${PVE_TOKENID}=${secret}"
log_success " token minted (secret captured, not logged)"
fi
fi
# Scoped ACL grants — AFTER the token exists (`pveum user token remove` purges the token's ACL, so
# re-applying post-rotate is mandatory; `acl modify` is idempotent so this is safe on the reuse path).
apply_scoped_acl
# If this box previously ran the pre-3b broad grant (re-install/upgrade), remove it — else the old
# FelhomAgent role at / would UNION with the scoped grant and defeat containment. Tolerate-absent.
remove_old_broad_acl
_state_mark token
}
#-------------------------------------------------------------------------------
# STEP 3 — compute grows (floors) if not passed
#-------------------------------------------------------------------------------
step_grows() {
log_step "3/8 compute volume grows"
# Golden base: rootfs 32G + Docker-data 16G + user-data 8G (build-golden.sh).
if [[ -z "$ROOTFS_GROW$DATAVOL_GROW$SYSDATA_GROW" ]]; then
local free_gib
free_gib=$(lvs --noheadings --units g -o lv_size,data_percent /dev/pve/data 2>/dev/null | awk '{gsub(/[^0-9.]/,"",$1); used=$2; print int($1*(100-used)/100)}' 2>/dev/null || echo 0)
# Reserve headroom; split the rest ~ docker 80% / sysdata 20%; rootfs stays golden.
ROOTFS_GROW=0
if [[ "${free_gib:-0}" -ge 300 ]]; then
DATAVOL_GROW=184; SYSDATA_GROW=42 # reproduces the standard 200G/50G appliance
elif [[ "${free_gib:-0}" -ge 150 ]]; then
DATAVOL_GROW=84; SYSDATA_GROW=22
else
DATAVOL_GROW=34; SYSDATA_GROW=12 # minimal floors
fi
log_info " auto-computed from ~${free_gib} GiB free"
fi
ROOTFS_GROW="${ROOTFS_GROW:-0}"; DATAVOL_GROW="${DATAVOL_GROW:-0}"; SYSDATA_GROW="${SYSDATA_GROW:-0}"
log_info " grows: rootfs +${ROOTFS_GROW}G (->$((32+ROOTFS_GROW))G), docker +${DATAVOL_GROW}G (->$((16+DATAVOL_GROW))G), sys_drive +${SYSDATA_GROW}G (->$((8+SYSDATA_GROW))G)"
_state_mark grows
}
#-------------------------------------------------------------------------------
# STEP 4 — host enroll (option C; single secret, no global key)
#-------------------------------------------------------------------------------
step_enroll() {
log_step "4/8 host enrollment (POST /host-enroll)"
if $DRY_RUN; then
log_dry "curl -fsS -X POST $HUB_URL/api/v1/host-enroll -H 'X-Retrieval-Password: <pass>' -d '{\"customer_id\":\"$CUSTOMER_ID\"}'"
HOST_ID="<dry-run-host-id>"; HOST_API_KEY="<dry-run-key>"; _state_mark enroll; return 0
fi
local resp code body
resp=$(curl -sS -w $'\n%{http_code}' -X POST "$HUB_URL/api/v1/host-enroll" \
-H "X-Retrieval-Password: $PASSPHRASE" -H 'Content-Type: application/json' \
-d "{\"customer_id\":\"$CUSTOMER_ID\"}")
code=$(tail -n1 <<<"$resp"); body=$(sed '$d' <<<"$resp")
case "$code" in
201) log_success " host MINTED (first enroll)" ;;
200) log_success " host REUSED (idempotent — existing credential)" ;;
401) die "host-enroll 401 (passphrase) — should have been caught in preflight" ;;
404) die "host-enroll 404 (unknown customer)" ;;
*) die "host-enroll unexpected $code: $body" ;;
esac
HOST_ID=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['host_id'])" "$body")
HOST_API_KEY=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['api_key'])" "$body")
[[ -n "$HOST_ID" && -n "$HOST_API_KEY" ]] || die "host-enroll: malformed response"
log_info " host_id: $HOST_ID (api_key captured, not logged)"
_state_mark enroll
}
#-------------------------------------------------------------------------------
# STEP 4b — break-glass credential (TASK G1): generate + set + vault the root@pam console password
#-------------------------------------------------------------------------------
# The human fallback for when BOTH the sshd path AND the agent-independent auto-heal (layers 1+2) have
# failed: a strong root@pam password lets the operator reach the PVE WEB CONSOLE (pveproxy :8006 — a
# failure domain distinct from sshd) and run the one-line /run/sshd fix. Generated with strong entropy,
# set via chpasswd, and vaulted to the hub over the enroll-authenticated channel (host api_key). The
# password is NEVER logged, printed, or written to any file — it goes stdin→chpasswd and stdin→curl
# only (SPIKE-felhom-sshd finding #9 / TASK G1 trap 3). Idempotent: skipped if already vaulted unless
# --rotate-recovery (a re-set would strand the operator's saved copy).
step_break_glass() {
log_step "4b/8 break-glass credential (root@pam console password → hub vault)"
if [[ -z "${HOST_ID:-}" || -z "${HOST_API_KEY:-}" ]]; then
log_warn " no host_id/api_key (enroll skipped?) — cannot vault a recovery credential; skipping"
return 0
fi
if $DRY_RUN; then
log_dry "openssl rand → strong root@pam password (never logged) ; chpasswd ; PUT $HUB_URL/api/v1/hosts/$HOST_ID/recovery-credential (Bearer host key)"
_state_mark break_glass; return 0
fi
if _state_has break_glass && ! $ROTATE_RECOVERY; then
log_skip " recovery credential already vaulted (use --rotate-recovery to regenerate)"
return 0
fi
# Strong password: 24 url-safe bytes (~144 bits). Kept ONLY in a local shell var, never on disk.
local newpw
newpw=$(openssl rand -base64 24 2>/dev/null | tr -d '\n' | tr '+/' '-_')
[[ ${#newpw} -ge 24 ]] || die "failed to generate a strong recovery password"
# Set root@pam (= the Linux root user on PVE) via chpasswd on STDIN — no argv, no log.
if ! printf 'root:%s\n' "$newpw" | chpasswd 2>/dev/null; then
newpw="" # scrub
die "chpasswd failed to set the root@pam recovery password"
fi
# Vault to the hub over the host-key-authenticated channel; password only on stdin (-d @-).
local code
code=$(printf '{"username":"root@pam","password":"%s"}' "$newpw" \
| curl -sS -o /dev/null -w '%{http_code}' -X PUT \
"$HUB_URL/api/v1/hosts/$HOST_ID/recovery-credential" \
-H "Authorization: Bearer $HOST_API_KEY" -H 'Content-Type: application/json' -d @- 2>/dev/null)
newpw="" # scrub the plaintext from the shell var the moment it is vaulted
case "$code" in
200) log_success " root@pam password set + vaulted to the hub (retrieve via the operator /admin path; never logged here)"
# Drill F-8: say this OUT LOUD — the operator was live-surprised by a 401 at the PVE
# GUI. Pointer only, never the credential.
log_warn " NOTE: the root@pam password just CHANGED — the old one now fails at the PVE web GUI (:8006)."
log_warn " Retrieve the new one at hub → host page (vaulted recovery credential)." ;;
401|403) die "recovery-credential vault rejected ($code) — host key/authorization problem" ;;
*) die "recovery-credential vault failed (HTTP $code)" ;;
esac
_state_mark break_glass
}
#-------------------------------------------------------------------------------
# STEP 5 — agent install: fetch+verify the binary, ensure the service user, sudoers, unit
#-------------------------------------------------------------------------------
# Closes the old prerequisite "install the agent binary + unit manually". Fetches the binary from
# Gitea (git token from controller.yaml), VERIFIES its sha256 against the hub manifest, then installs
# the non-root felhom-agent user + binary + sudoers + unit. The SERVICE is started in step 6 (after the
# config is written) — here we only install + daemon-reload + enable.
# Part 2 (v1.25.0): a fresh appliance install boots on the stock PVE ENTERPRISE apt repo with no
# subscription key, so every `apt-get update` 401s and the box can never pull security updates. Switch
# the Proxmox-managed repos to no-subscription — idempotent, APPLIANCE-ONLY, and scoped to the repos
# Proxmox itself ships (pve-enterprise + ceph enterprise, deb822 `.sources` on trixie/PVE9, or the
# legacy `.list`). The customer's own sources are never touched. No upgrades are run — repo alignment
# only. Runs BEFORE the first apt operation (the sudo/age installs below) so they don't hit the 401.
align_apt_repos() {
[[ "$MODE" == "appliance" ]] || { log_skip " apt repo alignment skipped (byo — the owner manages repos)"; return 0; }
if $DRY_RUN; then log_dry "disable pve/ceph enterprise .sources (Enabled: no) + legacy .list; add pve-no-subscription.sources; apt-get update"; return 0; fi
local d=/etc/apt/sources.list.d changed=0 suite="" keyring="" f
if [[ -f "$d/pve-enterprise.sources" ]]; then
suite=$(awk -F': *' '/^Suites:/{print $2; exit}' "$d/pve-enterprise.sources")
keyring=$(awk -F': *' '/^Signed-By:/{print $2; exit}' "$d/pve-enterprise.sources")
fi
[[ -n "$suite" ]] || suite="trixie"
[[ -n "$keyring" ]] || keyring="/usr/share/keyrings/proxmox-archive-keyring.gpg"
# 1a. disable enterprise deb822 .sources (append `Enabled: no`, idempotent)
for f in "$d/pve-enterprise.sources" "$d/ceph.sources"; do
[[ -f "$f" ]] || continue
grep -qiE '^Enabled:[[:space:]]*no' "$f" && continue
sed -i '/^[Ee]nabled:/d' "$f"; printf 'Enabled: no\n' >> "$f"; changed=1
log_info " disabled $(basename "$f") (Enabled: no)"
done
# 1b. legacy .list enterprise line → commented
for f in "$d/pve-enterprise.list"; do
[[ -f "$f" ]] || continue
grep -qE '^[[:space:]]*deb[[:space:]].*enterprise\.proxmox\.com' "$f" || continue
sed -i 's|^[[:space:]]*\(deb[[:space:]].*enterprise\.proxmox\.com.*\)$|# \1 # felhom: no-subscription|' "$f"; changed=1
log_info " commented enterprise line in $(basename "$f")"
done
# 2. add pve-no-subscription if it is not already present in any list/sources
if ! grep -rqsE 'download\.proxmox\.com/debian/pve([[:space:]/]|$)' "$d"/*.sources "$d"/*.list 2>/dev/null; then
cat > "$d/pve-no-subscription.sources" <<SRC
Types: deb
URIs: http://download.proxmox.com/debian/pve
Suites: $suite
Components: pve-no-subscription
Signed-By: $keyring
SRC
changed=1; log_info " added pve-no-subscription.sources (suite=$suite)"
fi
# 3. refresh — must be clean now (a stale enterprise line would 401)
if apt-get update -q >/dev/null 2>&1; then
log_success " apt repos aligned to no-subscription ($([[ $changed == 1 ]] && echo changed || echo 'already aligned'); apt-get update OK)"
else
log_warn " apt-get update non-zero after repo alignment — inspect $d (non-fatal to the install)"
fi
}
step_agent_install() {
log_step "5/8 agent install (fetch + verify + install)"
align_apt_repos # Part 2: switch enterprise→no-subscription before any apt-get (appliance only)
# Manifest + git fetch credential (both passphrase / config-retrieve — NO new credential).
resolve_artifacts
resolve_git_creds
[[ -n "$ART_AGENT_VER" ]] || die "hub artifact manifest has no agent version — set it in the operator UI (Configs → Day-0 artifacts)"
log_info " manifest: agent v$ART_AGENT_VER (sha ${ART_AGENT_SHA:0:16}…), golden v${ART_GOLDEN_VER:-<unset>}"
# Idempotent skip: same version already installed AND the service is healthy.
local cur=""
[[ -x "$AGENT_BIN" ]] && cur=$("$AGENT_BIN" --version 2>/dev/null | awk '{print $2}')
if [[ "$cur" == "$ART_AGENT_VER" ]] && systemctl is-active --quiet felhom-agent 2>/dev/null; then
log_skip " agent v$cur already installed + service active — skipping binary install"
else
local url="$GITEA_BASE/api/packages/$GITEA_OWNER/generic/felhom-agent/$ART_AGENT_VER/felhom-agent"
log_info " fetching agent binary v$ART_AGENT_VER from Gitea …"
if $DRY_RUN; then
log_dry "$(_dry_auth_hint) -o /tmp/felhom-agent.new $url ; verify sha256=$ART_AGENT_SHA ; install -m0755 -> $AGENT_BIN"
else
local tmp; tmp=$(mktemp -t felhom-agent.XXXXXX)
fetch_verify "$url" "$tmp" "$ART_AGENT_SHA"
# back up any existing binary before replacing
if [[ -f "$AGENT_BIN" ]]; then
cp -a "$AGENT_BIN" "${AGENT_BIN}.bak-$(date +%s)" 2>/dev/null || true
fi
install -m 0755 -o root -g root "$tmp" "$AGENT_BIN"
rm -f "$tmp"
log_success " installed $AGENT_BIN ($("$AGENT_BIN" --version 2>&1 | head -1))"
fi
fi
# The non-root model REQUIRES the `sudo` package (provides both `sudo` and `visudo`). A host that
# previously ran the agent as root+`direct` won't have it installed. Install it idempotently before
# the sudoers (visudo validates it) and before the daemon starts (it shells out via `sudo -n`).
if ! command -v sudo >/dev/null 2>&1 || ! command -v visudo >/dev/null 2>&1; then
if $DRY_RUN; then
log_dry "apt-get install -y sudo # required for the non-root agent (provides sudo + visudo)"
else
log_info " installing the 'sudo' package (required for the non-root agent model) …"
DEBIAN_FRONTEND=noninteractive apt-get install -y -q sudo >/dev/null 2>&1 \
|| { apt-get update -q >/dev/null 2>&1; DEBIAN_FRONTEND=noninteractive apt-get install -y -q sudo >/dev/null 2>&1; } \
|| die "failed to install the 'sudo' package (needed for the non-root agent)"
log_success " sudo installed ($(sudo --version 2>/dev/null | head -1))"
fi
fi
# `age` (drill F-10) — the escrow ceremony's identity-wrap dependency. NOT installed by PVE by
# default; a fresh-box ceremony died on it live (the demo host only had it from its spike-era
# install). DR-tier-by-default: baked on every install, both modes. Idempotent.
if ! command -v age >/dev/null 2>&1; then
if $DRY_RUN; then
log_dry "apt-get install -y age # escrow ceremony identity-wrap dependency (F-10)"
else
log_info " installing the 'age' package (escrow ceremony identity-wrap dependency) …"
DEBIAN_FRONTEND=noninteractive apt-get install -y -q age >/dev/null 2>&1 \
|| { apt-get update -q >/dev/null 2>&1; DEBIAN_FRONTEND=noninteractive apt-get install -y -q age >/dev/null 2>&1; } \
|| die "failed to install the 'age' package (the escrow ceremony needs it)"
log_success " age installed ($(age --version 2>/dev/null | head -1))"
fi
fi
# Resolve visudo by absolute path too (non-login SSH PATH can miss /usr/sbin).
local VISUDO; VISUDO=$(command -v visudo 2>/dev/null || echo /usr/sbin/visudo)
# Service user (system, no login, no home dir creation needed beyond state).
if $DRY_RUN; then
log_dry "useradd --system --no-create-home --shell /usr/sbin/nologin $AGENT_USER # if absent"
elif id "$AGENT_USER" >/dev/null 2>&1; then
log_info " service user $AGENT_USER exists"
else
useradd --system --no-create-home --shell /usr/sbin/nologin "$AGENT_USER"
log_success " created service user $AGENT_USER"
fi
# systemd-journal group: the NAS verify pipeline (agent v0.81.0) classifies mount failures from
# an UNPRIVILEGED `journalctl -u <unit>` read — group membership instead of a sudoers grant
# (rule: no journalctl via sudo). Idempotent. Existing installs: run the usermod by hand, then
# `systemctl restart felhom-agent`.
if $DRY_RUN; then
log_dry "usermod -aG systemd-journal $AGENT_USER"
elif id -nG "$AGENT_USER" | tr ' ' '\n' | grep -qx systemd-journal; then
log_info " $AGENT_USER already in systemd-journal"
else
usermod -aG systemd-journal "$AGENT_USER"
log_success " added $AGENT_USER to systemd-journal (unprivileged journal read for NAS verify)"
fi
# State dir (the old root deployment may have created it root-owned; StateDirectory= also adjusts
# on start, but chown here so the very first start has a writable dir).
run mkdir -p "$AGENT_STATE_DIR"
run chown -R "${AGENT_USER}:${AGENT_USER}" "$AGENT_STATE_DIR"
run chmod 0750 "$AGENT_STATE_DIR"
# ── Agent local-API leaf lifecycle (B.2) ──────────────────────────────────────────────────────
# The leaf's SHA-256 is pinned into EVERY guest's bootstrap. A reinstall that REGENERATES the leaf
# invalidates every controller's pin (the 2026-06-28 root→non-root incident → controller↔agent dead
# for days). Two protections:
# (a) --preserve-state-from DIR: carry the prior leaf+key+token-store over → the fp stays STABLE,
# no re-bootstrap needed. (Distinct from --preserve-from, which merges config sections only.)
# (b) populated-host guard: REFUSE to proceed leaf-less on a host that already has guests, unless
# --preserve-state-from or an explicit --allow-new-leaf is given. Converts the silent footgun
# into a hard stop.
local _have_leaf=false
[[ -f "$AGENT_STATE_DIR/local-api.crt" && -f "$AGENT_STATE_DIR/local-api.key" ]] && _have_leaf=true
if [[ -n "$PRESERVE_STATE_FROM" ]]; then
[[ -f "$PRESERVE_STATE_FROM/local-api.crt" && -f "$PRESERVE_STATE_FROM/local-api.key" ]] \
|| die "--preserve-state-from $PRESERVE_STATE_FROM: local-api.crt/key not found there"
openssl x509 -in "$PRESERVE_STATE_FROM/local-api.crt" -noout >/dev/null 2>&1 \
|| die "--preserve-state-from: $PRESERVE_STATE_FROM/local-api.crt does not parse as an X.509 cert — refusing"
if $DRY_RUN; then
log_dry "preserve agent state: copy local-api.{crt,key}+local-tokens.log from $PRESERVE_STATE_FROM -> $AGENT_STATE_DIR (chown $AGENT_USER; 644/600/600)"
else
install -o "$AGENT_USER" -g "$AGENT_USER" -m 0644 "$PRESERVE_STATE_FROM/local-api.crt" "$AGENT_STATE_DIR/local-api.crt"
install -o "$AGENT_USER" -g "$AGENT_USER" -m 0600 "$PRESERVE_STATE_FROM/local-api.key" "$AGENT_STATE_DIR/local-api.key"
[[ -f "$PRESERVE_STATE_FROM/local-tokens.log" ]] && \
install -o "$AGENT_USER" -g "$AGENT_USER" -m 0600 "$PRESERVE_STATE_FROM/local-tokens.log" "$AGENT_STATE_DIR/local-tokens.log"
log_success " preserved agent leaf+token store from $PRESERVE_STATE_FROM (pin stays stable — no re-bootstrap)"
fi
_have_leaf=true
fi
# GL-8/F3: the guard fires only when a FELHOM guest is present — regenerating the leaf orphans
# THAT guest's pinned fingerprint (the 2026-06-28 incident). A BYO host's own non-Felhom guests
# have no Felhom pin to break, so `felhom_guests` (the /etc/felhom-bootstrap detector) — NOT
# `pct list` (any guest) — is the correct populated-host test: a fresh byo install on a host that
# only runs the owner's guests now proceeds without `--allow-new-leaf`.
if ! $_have_leaf && ! $ALLOW_NEW_LEAF; then
if [[ -n "$(felhom_guests)" ]]; then
die "this host already has a Felhom guest but $AGENT_STATE_DIR has no agent leaf to preserve.
Re-running here will REGENERATE the leaf and invalidate every controller's pin (the 2026-06-28 incident).
Pass --preserve-state-from <old state dir> to keep the pin stable, or --allow-new-leaf to regenerate
intentionally (every guest must then be re-bootstrapped)."
fi
fi
# Guarded-mkfs wrapper (Impl-1 Part B) — the ONLY mkfs path the sudoers permits. Install it BEFORE
# the sudoers (which allowlists it), 0755 root:root under /usr/local/sbin. bash -n before install.
if $DRY_RUN; then
log_dry "fetch configs/felhom-mkfs-guarded.sh ; bash -n ; install 0755 -> /usr/local/sbin/felhom-mkfs-guarded"
else
local wtmp; wtmp=$(mktemp -t felhom-mkfs.XXXXXX)
fetch_raw "configs/felhom-mkfs-guarded.sh" "$wtmp"
bash -n "$wtmp" || { rm -f "$wtmp"; die "fetched felhom-mkfs-guarded.sh failed bash -n — refusing to install"; }
install -m 0755 -o root -g root "$wtmp" /usr/local/sbin/felhom-mkfs-guarded
rm -f "$wtmp"
log_success " installed /usr/local/sbin/felhom-mkfs-guarded (0755, the guarded mkfs path)"
fi
# Self-update guarded wrapper (TASK D1) — the ONLY binary-swap path the sudoers permits. Install
# it BEFORE the sudoers (which allowlists it), 0755 root:root under /usr/local/sbin. sh -n first.
if $DRY_RUN; then
log_dry "fetch configs/felhom-selfupdate-guarded ; sh -n ; install 0755 -> /usr/local/sbin/felhom-selfupdate-guarded"
else
local sutmp; sutmp=$(mktemp -t felhom-selfupd.XXXXXX)
fetch_raw "configs/felhom-selfupdate-guarded" "$sutmp"
sh -n "$sutmp" || { rm -f "$sutmp"; die "fetched felhom-selfupdate-guarded failed sh -n — refusing to install"; }
install -m 0755 -o root -g root "$sutmp" /usr/local/sbin/felhom-selfupdate-guarded
rm -f "$sutmp"
log_success " installed /usr/local/sbin/felhom-selfupdate-guarded (0755, the guarded A/B binary-swap path)"
fi
# PBS-DR apply wrapper (drill F-7) — the ONLY pbs-apply path the FELHOM_PBSDR sudoers alias
# permits. The sudoers has shipped the alias since the PBS-DR slice; the binary itself did not
# ship, so the 3 pbsdr-* capabilities were born DEGRADED ("binary not found") on every fresh
# box. DR-tier-by-default (2026-07-12): capability is baked on EVERY install (activation is a
# hub flag). Install BEFORE the sudoers, 0755 root:root, bash -n first — the mkfs pattern.
if $DRY_RUN; then
log_dry "fetch configs/felhom-pbs-apply ; bash -n ; install 0755 -> /usr/local/sbin/felhom-pbs-apply"
else
local patmp; patmp=$(mktemp -t felhom-pbsapply.XXXXXX)
fetch_raw "configs/felhom-pbs-apply" "$patmp"
bash -n "$patmp" || { rm -f "$patmp"; die "fetched felhom-pbs-apply failed bash -n — refusing to install"; }
install -m 0755 -o root -g root "$patmp" /usr/local/sbin/felhom-pbs-apply
rm -f "$patmp"
log_success " installed /usr/local/sbin/felhom-pbs-apply (0755, the guarded PBS-DR apply path)"
fi
# E-2a: backup-target apply wrapper — the ONLY path the FELHOM_BACKUPTARGET sudoers alias permits.
# The agent cannot create a PVE storage (Datastore.Allocate at /storage) or grant an ACL
# (Permissions.Modify); both are deliberately outside its role, so the move runs through this
# fenced shim. Same pattern as pbs-apply: install BEFORE the sudoers, 0755 root:root, bash -n first.
if $DRY_RUN; then
log_dry "fetch configs/felhom-backup-target-apply ; bash -n ; install 0755 -> /usr/local/sbin/felhom-backup-target-apply"
else
local bttmp; bttmp=$(mktemp -t felhom-bt.XXXXXX)
fetch_raw "configs/felhom-backup-target-apply" "$bttmp"
bash -n "$bttmp" || { rm -f "$bttmp"; die "fetched felhom-backup-target-apply failed bash -n — refusing to install"; }
install -m 0755 -o root -g root "$bttmp" /usr/local/sbin/felhom-backup-target-apply
rm -f "$bttmp"
log_success " installed /usr/local/sbin/felhom-backup-target-apply (0755, the guarded backup-target path)"
fi
# Sudoers — fetch the canonical file, validate with visudo -cf BEFORE installing (0440 root:root).
if $DRY_RUN; then
log_dry "fetch configs/felhom-agent.sudoers ; visudo -cf ; install 0440 -> $AGENT_SUDOERS"
else
local sdtmp; sdtmp=$(mktemp -t felhom-sudoers.XXXXXX)
fetch_raw "configs/felhom-agent.sudoers" "$sdtmp"
"$VISUDO" -cf "$sdtmp" >/dev/null || { rm -f "$sdtmp"; die "fetched sudoers failed visudo -cf — refusing to install"; }
install -m 0440 -o root -g root "$sdtmp" "$AGENT_SUDOERS"
rm -f "$sdtmp"
# re-validate the live drop-in in the full sudoers context
"$VISUDO" -cf /etc/sudoers >/dev/null || die "sudoers invalid after installing $AGENT_SUDOERS"
log_success " installed $AGENT_SUDOERS (0440, visudo-validated)"
fi
# systemd unit — fetch the canonical unit, install, daemon-reload, enable (NOT start — no config yet).
if $DRY_RUN; then
log_dry "fetch configs/felhom-agent.service -> $AGENT_UNIT ; systemctl daemon-reload ; systemctl enable felhom-agent"
else
local untmp; untmp=$(mktemp -t felhom-unit.XXXXXX)
fetch_raw "configs/felhom-agent.service" "$untmp"
grep -q "User=$AGENT_USER" "$untmp" || { rm -f "$untmp"; die "fetched unit does not run as $AGENT_USER — refusing"; }
if [[ -f "$AGENT_UNIT" ]]; then cp -a "$AGENT_UNIT" "${AGENT_UNIT}.bak-$(date +%s)" 2>/dev/null || true; fi
install -m 0644 -o root -g root "$untmp" "$AGENT_UNIT"
rm -f "$untmp"
systemctl daemon-reload
systemctl enable felhom-agent >/dev/null 2>&1 || true
log_success " installed $AGENT_UNIT + enabled (started in step 6 after config)"
fi
# Self-update rollback unit + start-limit drop-in (TASK D1). The drop-in gives felhom-agent the
# tuned start-limit (SPIKE-agent-selfupdate: [Unit] StartLimitIntervalSec=120 + Burst=4 → a
# crash-looping update reaches terminal `failed` in ~20s instead of looping forever) and the
# OnFailure= that auto-reverts. Both are idempotent installs (+ daemon-reload). Non-fatal on a
# box whose agent repo predates them (the raw fetch would 404) — self-update just stays manual.
if $DRY_RUN; then
log_dry "fetch configs/felhom-agent-rollback.service -> /etc/systemd/system/ ; fetch configs/felhom-agent-limits.conf -> $AGENT_UNIT.d/ ; daemon-reload"
else
local rbtmp; rbtmp=$(mktemp -t felhom-rollback.XXXXXX)
if fetch_raw "configs/felhom-agent-rollback.service" "$rbtmp" 2>/dev/null; then
install -m 0644 -o root -g root "$rbtmp" /etc/systemd/system/felhom-agent-rollback.service
local ddir="${AGENT_UNIT}.d"; mkdir -p "$ddir"
local lmtmp; lmtmp=$(mktemp -t felhom-limits.XXXXXX)
fetch_raw "configs/felhom-agent-limits.conf" "$lmtmp"
# [SF-3] the keys MUST be in [Unit] — refuse a drop-in that would half-apply in [Service].
grep -q '^\[Unit\]' "$lmtmp" || { rm -f "$lmtmp" "$rbtmp"; die "self-update limits drop-in missing [Unit] section — refusing"; }
install -m 0644 -o root -g root "$lmtmp" "$ddir/felhom-agent-limits.conf"
rm -f "$lmtmp"
systemctl daemon-reload
log_success " installed self-update rollback unit + start-limit drop-in (auto-rollback armed)"
else
log_skip " self-update rollback artifacts not in the agent repo yet — self-update stays manual"
fi
rm -f "$rbtmp"
fi
# Management-plane break-glass layers 1+2 (TASK G1). Three artifacts that keep the host reachable
# even if a second sshd (H1) removes the SHARED /run/sshd privsep dir (SPIKE-felhom-sshd §8):
# • felhom-privsep.conf (tmpfiles) — layer 1: /run/sshd is boot-persistent, owned by no unit.
# • felhom-mgmt-watchdog (script) — layer 2 heal action (recreate dir + reset-failed sshd).
# • .service + .timer — run it every ~60s, AGENT-INDEPENDENTLY (heals with the
# agent down — the whole point; trap 1).
# Non-fatal if the agent repo predates them (raw fetch 404s → break-glass just stays manual).
# HARD GUARD: refuse ANY fetched unit that declares RuntimeDirectory= — that directive is the very
# incident G1 closes (a second sshd's `RuntimeDirectory=sshd` removed the shared /run/sshd).
install_mgmt_watchdog
# H1: dedicated felhom-sshd OOB instance + static belt (appliance default since v1.25.0; --no-oob opts out).
install_oob
_state_mark agent_install
}
# install_mgmt_watchdog fetches + installs the G1 break-glass host artifacts (idempotent; enables the
# timer). Split out for readability; called from step_agent_install. Every unit is RuntimeDirectory-
# guarded (trap 2). Non-fatal on a repo that predates the artifacts.
install_mgmt_watchdog() {
if $DRY_RUN; then
log_dry "fetch configs/felhom-privsep.tmpfiles -> /etc/tmpfiles.d/felhom-privsep.conf ; systemd-tmpfiles --create"
log_dry "fetch configs/felhom-mgmt-watchdog.sh -> /usr/local/sbin/felhom-mgmt-watchdog (0755)"
log_dry "fetch configs/felhom-mgmt-watchdog.{service,timer} -> /etc/systemd/system/ ; enable --now felhom-mgmt-watchdog.timer"
return 0
fi
local wdtmp; wdtmp=$(mktemp -t felhom-mgmt-wd.XXXXXX)
if ! fetch_raw "configs/felhom-mgmt-watchdog.sh" "$wdtmp" 2>/dev/null; then
log_skip " mgmt-watchdog artifacts not in the agent repo yet — break-glass auto-heal stays manual"
rm -f "$wdtmp"; return 0
fi
sh -n "$wdtmp" || { rm -f "$wdtmp"; die "fetched felhom-mgmt-watchdog.sh failed sh -n — refusing to install"; }
install -m 0755 -o root -g root "$wdtmp" /usr/local/sbin/felhom-mgmt-watchdog
rm -f "$wdtmp"
# tmpfiles (layer 1) — install + create now so /run/sshd is guaranteed present immediately.
local tftmp; tftmp=$(mktemp -t felhom-privsep.XXXXXX)
fetch_raw "configs/felhom-privsep.tmpfiles" "$tftmp"
install -m 0644 -o root -g root "$tftmp" /etc/tmpfiles.d/felhom-privsep.conf
rm -f "$tftmp"
systemd-tmpfiles --create /etc/tmpfiles.d/felhom-privsep.conf 2>/dev/null || true
# units (layer 2) — RuntimeDirectory guard on BOTH before install (trap 2 / the incident cause).
local svctmp tmrtmp; svctmp=$(mktemp -t felhom-wd-svc.XXXXXX); tmrtmp=$(mktemp -t felhom-wd-tmr.XXXXXX)
fetch_raw "configs/felhom-mgmt-watchdog.service" "$svctmp"
fetch_raw "configs/felhom-mgmt-watchdog.timer" "$tmrtmp"
if grep -qiE '^[[:space:]]*RuntimeDirectory[[:space:]]*=' "$svctmp" "$tmrtmp"; then
rm -f "$svctmp" "$tmrtmp"
die "mgmt-watchdog unit declares RuntimeDirectory= — that is the incident G1 fixes; refusing to install"
fi
install -m 0644 -o root -g root "$svctmp" /etc/systemd/system/felhom-mgmt-watchdog.service
install -m 0644 -o root -g root "$tmrtmp" /etc/systemd/system/felhom-mgmt-watchdog.timer
rm -f "$svctmp" "$tmrtmp"
systemctl daemon-reload
systemctl enable --now felhom-mgmt-watchdog.timer >/dev/null 2>&1 || true
log_success " installed break-glass layers 1+2 (tmpfiles /run/sshd + agent-independent watchdog timer)"
}
# install_oob installs the dedicated felhom-sshd OOB instance + the static nft belt (TASK H1). The
# STATIC parts land here (unit, host key, felhom-op user + sudoers, belt table + loader); the agent
# reconciles the DYNAMIC parts (port claim, config render, authorized_keys, belt set elements) once
# oob.enabled=true. Gated by --enable-oob. Non-fatal if the agent repo predates the artifacts.
# The felhom-sshd unit is RuntimeDirectory-guarded (the G1 incident cause) before install.
install_oob() {
$ENABLE_OOB || { log_skip " OOB (felhom-sshd) off (byo, or appliance --no-oob) — skipping"; return 0; }
if $DRY_RUN; then
log_dry "mkdir /etc/felhom-sshd{,/authorized_keys} ; ssh-keygen host key ; useradd felhom-op ; install felhom-op sudoers"
log_dry "fetch felhom-sshd.service (RuntimeDirectory-guarded) + felhom-oob.nft + felhom-oob-nft.service ; enable"
return 0
fi
install -d -o root -g root -m 0755 /etc/felhom-sshd /etc/felhom-sshd/authorized_keys
# dedicated host key (stable across reloads) — generate once.
if [[ ! -f /etc/felhom-sshd/ssh_host_ed25519_key ]]; then
ssh-keygen -t ed25519 -N "" -f /etc/felhom-sshd/ssh_host_ed25519_key -C felhom-sshd-hostkey -q
chmod 600 /etc/felhom-sshd/ssh_host_ed25519_key
fi
# operator login user (no ambient privilege; escalates only via the felhom-op sudoers verbs).
id felhom-op >/dev/null 2>&1 || useradd --create-home --shell /bin/bash felhom-op
# felhom-op scoped sudoers (visudo-validated before install).
local optmp; optmp=$(mktemp -t felhom-op-sudoers.XXXXXX)
if fetch_raw "configs/felhom-op.sudoers" "$optmp" 2>/dev/null; then
"$VISUDO" -cf "$optmp" >/dev/null || { rm -f "$optmp"; die "fetched felhom-op sudoers failed visudo -cf"; }
install -m 0440 -o root -g root "$optmp" /etc/sudoers.d/felhom-op
fi
rm -f "$optmp"
# felhom-sshd unit — RuntimeDirectory guard (the G1 incident cause) BEFORE install.
local sutmp; sutmp=$(mktemp -t felhom-sshd-unit.XXXXXX)
if ! fetch_raw "configs/felhom-sshd.service" "$sutmp" 2>/dev/null; then
log_skip " felhom-sshd.service not in the agent repo yet — OOB stays unconfigured"
rm -f "$sutmp"; return 0
fi
if grep -qiE '^[[:space:]]*RuntimeDirectory[[:space:]]*=' "$sutmp"; then
rm -f "$sutmp"; die "felhom-sshd unit declares RuntimeDirectory= — the SPIKE-felhom-sshd §8 incident cause; refusing"
fi
install -m 0644 -o root -g root "$sutmp" /etc/systemd/system/felhom-sshd.service
rm -f "$sutmp"
# static belt table + boot loader.
local ntmp utmp; ntmp=$(mktemp -t felhom-oob-nft.XXXXXX); utmp=$(mktemp -t felhom-oob-unit.XXXXXX)
fetch_raw "configs/felhom-oob.nft" "$ntmp"
fetch_raw "configs/felhom-oob-nft.service" "$utmp"
nft -c -f "$ntmp" >/dev/null 2>&1 || { rm -f "$ntmp" "$utmp"; die "felhom-oob.nft failed nft -c — refusing"; }
install -m 0644 -o root -g root "$ntmp" /etc/felhom-oob.nft
install -m 0644 -o root -g root "$utmp" /etc/systemd/system/felhom-oob-nft.service
rm -f "$ntmp" "$utmp"
systemctl daemon-reload
systemctl enable --now felhom-oob-nft.service >/dev/null 2>&1 || true # load the static belt now
systemctl enable felhom-sshd >/dev/null 2>&1 || true # NOT start — the agent renders the config first
log_success " installed OOB felhom-sshd instance + static belt (agent renders config + fills sets once oob.enabled)"
}
#-------------------------------------------------------------------------------
# ensure_island_bridge — create the R-50 host-internal island bridge (vmbr9, portless, fixed /30) if
# absent, idempotently: append an ifupdown2 stanza to /etc/network/interfaces + `ifreload -a`.
# Portless (bridge-ports none) = no physical NIC, so the /30 cannot collide with any customer LAN and
# survives any LAN renumber; vmbr0 is a separate stanza and is never touched. Safe to re-run (no-ops
# when present). Spike-proven non-disruptive (SPIKE-island-bridge-2026-07-25, probe P2).
#-------------------------------------------------------------------------------
ensure_island_bridge() {
local ifaces=/etc/network/interfaces
if grep -qE "^[[:space:]]*iface[[:space:]]+${ISLAND_BRIDGE}[[:space:]]" "$ifaces" 2>/dev/null; then
log_info " island bridge ${ISLAND_BRIDGE} already present — leaving it"
return 0
fi
if $DRY_RUN; then
log_dry "append ${ISLAND_BRIDGE} stanza (portless, ${ISLAND_HOST_CIDR}) to $ifaces + ifreload -a"
return 0
fi
log_info " creating island bridge ${ISLAND_BRIDGE} (portless, ${ISLAND_HOST_CIDR})"
cp -a "$ifaces" "${ifaces}.pre-island.bak" 2>/dev/null || true
cat >> "$ifaces" <<EOF
# R-50 island control plane (felhom-host-install) — host-internal, no physical port. The agent binds
# ${ISLAND_HOST_IP}:${ISLAND_PORT} here; the guest gets ${ISLAND_GUEST_CIDR} on net1. Never add a
# bridge-port — the portlessness is what makes the /30 uncollidable with any customer LAN.
auto ${ISLAND_BRIDGE}
iface ${ISLAND_BRIDGE} inet static
address ${ISLAND_HOST_CIDR}
bridge-ports none
bridge-stp off
bridge-fd 0
EOF
if command -v ifreload >/dev/null 2>&1; then
run ifreload -a || log_warn " ifreload -a returned non-zero — verify ${ISLAND_BRIDGE} is up"
else
run ifup "${ISLAND_BRIDGE}" || log_warn " ifup ${ISLAND_BRIDGE} returned non-zero"
fi
ip -4 -o addr show "${ISLAND_BRIDGE}" 2>/dev/null | grep -q "${ISLAND_HOST_IP}" \
&& log_success " ${ISLAND_BRIDGE} up: ${ISLAND_HOST_CIDR}" \
|| log_warn " ${ISLAND_BRIDGE} did not come up with ${ISLAND_HOST_IP} — check $ifaces"
}
#-------------------------------------------------------------------------------
# STEP 6 — write agent config + ensure service healthy
#-------------------------------------------------------------------------------
step_agent_config() {
log_step "6/8 agent config + service"
# E-2 Part 2: resolve the whole-guest backup target BEFORE the config is written — the resolved
# value feeds local_backup_target below. Scenario F safe: it returns early, changing nothing, when
# a target already exists, and `setdefault` preserves an existing box's whole backup block anyway.
configure_backup_target
# TLS pin: the SERVED leaf cert fingerprint (not pvesh node info — may differ)
local fp
fp=$(echo | openssl s_client -connect 127.0.0.1:8006 2>/dev/null | openssl x509 -noout -fingerprint -sha256 2>/dev/null | sed 's/.*=//')
[[ -n "$fp" ]] || log_warn " could not compute TLS fingerprint (leaving empty — agent will use system trust)"
# bridge / local-api addr. LAN_IP (vmbr0) is always resolved — it is the lan_resolver anchor even
# on an island install (Finding-1), and the historical bind on a non-island install.
LAN_IP=$(ip -4 -o addr show vmbr0 2>/dev/null | awk '{print $4}' | cut -d/ -f1 | head -1)
if [[ "$MODE" == "appliance" && "$WANT_ISLAND" == true && -z "$BRIDGE_ADDR" ]]; then
# R-50 appliance default: bind the local API on the host-internal island bridge (LAN-move
# immune). Create the bridge first so the agent can bind it; point listen_addr at it; the guest
# gets net1 from the agent (island_bridge/island_guest_addr, written below). lan_resolver.host_ip
# is pinned to the LAN IP in the config write — never let it default off the island listen_addr.
ensure_island_bridge
BRIDGE_ADDR="${ISLAND_HOST_IP}:${ISLAND_PORT}"
ISLAND_ON=1
log_info " R-50 island ON: local_api=${BRIDGE_ADDR} (${ISLAND_BRIDGE}); guest net1=${ISLAND_GUEST_CIDR}; lan_resolver.host_ip=${LAN_IP:-<none>}"
elif [[ -z "$BRIDGE_ADDR" ]]; then
BRIDGE_ADDR="${LAN_IP:-127.0.0.1}:8443" # byo / --no-island / explicit-LAN: historical bind
elif [[ "$BRIDGE_ADDR" != *:* ]]; then
BRIDGE_ADDR="${BRIDGE_ADDR}:8443"
fi
log_info " node=$NODE local_api=$BRIDGE_ADDR tls_fp=${fp:0:17}"
if $DRY_RUN; then
log_dry "write $AGENT_CONFIG (0600): proxmox{endpoint,node=$NODE,token=<secret>,tls.fingerprint=$fp} hub{url=$HUB_URL,host_id=$HOST_ID,api_key=<secret>} local_api{$BRIDGE_ADDR}"
if [[ -n "$RESOLVED_OP_LINE$RESOLVED_REC_LINE" ]]; then
log_dry "write authz.signers: operational=${RESOLVED_OP_ID:-<none>} recovery=${RESOLVED_REC_ID:-<none>} (operator-signed self-update authority)"
else
log_dry "authz.signers: none resolved — self-update stays dormant (preserved signers, if any, are kept)"
fi
if [[ "$MODE" == "byo" ]]; then
log_dry "assert (byo) written config: lan_resolver.enable is false/absent"
log_dry "assert (byo) written config: oob.enabled is false/absent"
log_dry "(byo NOTE: wg_tunnel.enabled=true is now the DEFAULT both modes — WG is base infrastructure; consent = the typed disclosure ack)"
fi
log_dry "systemctl restart felhom-agent && felhom-agent --config $AGENT_CONFIG --selftest"
_state_mark agent_config; return 0
fi
mkdir -p "$(dirname "$AGENT_CONFIG")"
# Build config: optional preserve base + fresh-host defaults + Day-0 overrides.
# Secrets passed via env (NOT argv) to avoid ps exposure.
PVE_TOKEN="$PVE_TOKEN" HOST_API_KEY="$HOST_API_KEY" \
NODE="$NODE" FP="$fp" HUB_URL="$HUB_URL" HOST_ID="$HOST_ID" BRIDGE_ADDR="$BRIDGE_ADDR" \
ISLAND_ON="$ISLAND_ON" ISLAND_BRIDGE="$ISLAND_BRIDGE" ISLAND_GUEST_CIDR="$ISLAND_GUEST_CIDR" LAN_IP="$LAN_IP" \
OP_KEY_ID="$RESOLVED_OP_ID" OP_KEY_LINE="$RESOLVED_OP_LINE" \
REC_KEY_ID="$RESOLVED_REC_ID" REC_KEY_LINE="$RESOLVED_REC_LINE" \
PRESERVE_FROM="$PRESERVE_FROM" INSTALL_MODE="$MODE" ENABLE_OOB="$ENABLE_OOB" BACKUP_TARGET_RESOLVED="$BACKUP_TARGET_RESOLVED" OUT="$AGENT_CONFIG" python3 <<'PY'
import json, os, sys
out = os.environ['OUT']
base = {}
pf = os.environ.get('PRESERVE_FROM','')
if pf and os.path.exists(pf):
try: base = json.load(open(pf))
except Exception: base = {}
# fresh-host defaults for any section not preserved
base.setdefault('log_level','info')
# CAMPAIGN-3 Part 6: deployment_mode gates the agent's node self-heal (host networking recovery at
# boot, F12-class). Authoritative from --mode: "appliance" unlocks the remedy; "byo" (or anything
# else, including a stale preserved value) leaves it CHECK-only. Written explicitly so a byo box can
# never inherit an appliance self-heal from a --preserve-from config.
base['deployment_mode'] = 'appliance' if os.environ.get('INSTALL_MODE') == 'appliance' else 'byo'
# H1 / F9 (v1.25.0): when the belt is installed (appliance default, or explicit --enable-oob), turn the
# agent's OOB loop on so it claims a port, renders felhom-sshd's config, and fills the belt sets each
# tick. Without this the static units sit inert (the pre-v1.25.0 state that left the belt unusable). byo
# never reaches here with ENABLE_OOB true (refused above) and the byo asserts below re-check oob is off.
if os.environ.get('ENABLE_OOB') == 'true':
base.setdefault('oob', {})['enabled'] = True
# privileged.mode = "sudo": the canonical unit runs the agent as the NON-root felhom-agent user, so
# every host-root op goes through `sudo -n` against /etc/sudoers.d/felhom-agent. ("direct" was the old
# dev/CI shortcut for a root agent.) Force the mode authoritative (a stale preserved "direct" config
# would otherwise break the non-root daemon); the binary paths MUST match the sudoers allowlist.
base.setdefault('privileged', {})
base['privileged']['mode'] = 'sudo'
base['privileged'].setdefault('sudo_path','sudo')
for _k,_v in {"unit_dir":"/etc/systemd/system","stage_dir":"/var/lib/felhom-agent/units","systemctl":"/usr/bin/systemctl","install":"/usr/bin/install","smartctl":"/usr/sbin/smartctl","lvs":"/usr/sbin/lvs"}.items():
base['privileged'].setdefault(_k,_v)
base.setdefault('storage', {"watchdog_interval_seconds":5,"watchdog_debounce_seconds":15,"known_refresh_seconds":20})
# R-82: local DAILY + offsite WEEKLY. The two tiers carry SEPARATE cadences and retentions —
# keep_last=3 is three DAYS on the daily tier and three WEEKS on a weekly one, so one shared knob
# would guarantee that one of them is wrong. keep_last=2 on the offsite tier = two weeks (operator
# ruling 2026-07-26).
#
# The offsite tier is written even though `felhom-pbs` does not exist yet: that storage appears only
# when the hub provisions the DR tier. The agent DEFERS a tier whose target storage is absent
# (localapi targetStoragePresent), so the tier stays silent until it is real and goes live with no
# restart the moment it is provisioned. Writing it here means a box is never left with a DR tier that
# is provisioned but unscheduled — the exact "applied and empty" state R-82 exists to end.
#
# setdefault: an EXISTING box's backup block is preserved WHOLE. Upgrades never gain the tier here —
# they are migrated explicitly (R-82 Slice D.2), so an in-place upgrade can never silently start
# writing to an offsite datastore.
base.setdefault('backup', {"local_backup_target":os.environ.get('BACKUP_TARGET_RESOLVED','local'),"local_backup_retention":3,"restore_storage":"local-lvm","restore_test_cadence_seconds":0,"scratch_vmid_min":990000,"scratch_vmid_max":990009,"pbs_secret_dir":"/etc/pve/priv/storage","backup_cadence_seconds":0,"backup_targets":[{"target_id":"felhom-pbs","cadence_seconds":604800,"keep_last":2}]})
base.setdefault('local_api', {})
base['local_api'].setdefault('enable', True)
base['local_api']['listen_addr'] = os.environ['BRIDGE_ADDR']
# R-50: on an island install, record the island bridge + the guest net1 CIDR so the provisioner
# (bringup) attaches the guest's static island NIC. All-or-nothing — the agent's config.Validate
# rejects a half-set island. On a non-island install these keys are ABSENT (pre-R-50, LAN-only).
if os.environ.get('ISLAND_ON') == '1':
base['local_api']['island_bridge'] = os.environ['ISLAND_BRIDGE']
base['local_api']['island_guest_addr'] = os.environ['ISLAND_GUEST_CIDR']
base['local_api'].setdefault('cert_file','/var/lib/felhom-agent/local-api.crt')
base['local_api'].setdefault('key_file','/var/lib/felhom-agent/local-api.key')
base['local_api'].setdefault('token_store','/var/lib/felhom-agent/local-tokens.log')
# GL-2: byo defaults the lan-resolver OFF (the installer must never take over the owner's host DNS;
# the post-write byo asserts refuse a true value however it got in). Appliance keeps the historical
# default-on. setdefault either way — a --preserve-from section wins and the asserts judge it.
if os.environ.get('INSTALL_MODE') == 'byo':
base.setdefault('lan_resolver', {"enable": False})
else:
base.setdefault('lan_resolver', {"enable": True})
# R-50 Finding-1 (the dnsmasq trap): pin the split-horizon DNS to the LAN IP EXPLICITLY. Without
# this, config.go's LANResolverConfig.WithDefaults derives host_ip from listen_addr — which on an
# island install is the island IP (169.254.253.1) — silently moving LAN DNS onto the island bridge
# and killing it (spike-confirmed live). Set whenever the LAN IP is known (island or LAN bind).
_lan = os.environ.get('LAN_IP','').strip()
if _lan:
base['lan_resolver']['host_ip'] = _lan
# WG tunnel (drill F-9; operator decision 2026-07-12 #5): WG is BASE infrastructure like
# cloudflared — always-on for every install, both modes (it is also the future OOB-management
# transport; the PBS-DR tier merely RIDES it). Registration is hands-free (the drill proved it:
# keygen -> hub peer 10.77.0.x/32 -> handshake, no vouch). setdefault, not override: a
# --preserve-from config that deliberately disabled it is respected (the per-customer opt-out
# lives hub-side; byo consent = the disclosure ack + the tester agreement).
base.setdefault('wg_tunnel', {"enabled": True})
# Day-0 overrides (always authoritative)
base['proxmox'] = {
"endpoint":"https://127.0.0.1:8006",
"node": os.environ['NODE'],
"token": os.environ['PVE_TOKEN'],
"tls": {"fingerprint": os.environ['FP'], "insecure_skip_verify": False},
}
base['hub'] = {
"url": os.environ['HUB_URL'],
"host_id": os.environ['HOST_ID'],
"api_key": os.environ['HOST_API_KEY'],
"poll_seconds": base.get('hub',{}).get('poll_seconds',900),
"timeout_seconds": base.get('hub',{}).get('timeout_seconds',30),
}
# GL-4: operator-signed self-update authority (authz.signers, schema = agent config.go SignerKey).
# Script/file-provided keys are authoritative ONLY when non-empty — a reinstall with NO keys keeps
# a manually-pinned box's preserved signers (never silently clobber; C4). nonce_store_path is never
# invented here: the agent defaults it, and a preserved value rides along untouched in base['authz'].
signers = []
if os.environ.get('OP_KEY_LINE',''):
signers.append({"key_id": os.environ['OP_KEY_ID'], "role": "operational", "public_key": os.environ['OP_KEY_LINE']})
if os.environ.get('REC_KEY_LINE',''):
signers.append({"key_id": os.environ['REC_KEY_ID'], "role": "recovery", "public_key": os.environ['REC_KEY_LINE']})
if signers:
authz = base.setdefault('authz', {})
prev = authz.get('signers') or []
if prev:
print("[config] replacing %d preserved authz signer(s) with the script/file-provided operator key(s) (pin rotation)" % len(prev), file=sys.stderr)
authz['signers'] = signers
fd = os.open(out, os.O_WRONLY|os.O_CREAT|os.O_TRUNC, 0o600)
with os.fdopen(fd,'w') as f:
json.dump(base, f, indent=2); f.write('\n')
PY
# The non-root felhom-agent daemon must READ this config (token + hub api_key live here). Own it by
# the service user, 0600 (root still reads it for the provision one-shot).
chown "${AGENT_USER}:${AGENT_USER}" "$AGENT_CONFIG" 2>/dev/null || chmod 600 "$AGENT_CONFIG"
chmod 600 "$AGENT_CONFIG"
log_success " wrote $AGENT_CONFIG (0600 ${AGENT_USER})"
# GL-2 byo config asserts: these three features take over host DNS / open tunnels — forbidden on
# a customer/BYO box until their production endpoints exist. Parse the JUST-WRITTEN file (same
# interpreter pattern as the write above) and refuse to start the daemon on any true value —
# a --preserve-from carrying lan_resolver.enable=true is exactly what this catches.
# NOTE (v1.15.0 / decision 5): wg_tunnel.enabled was REMOVED from the byo-forbidden set — WG is
# base infrastructure (always-on both modes, like cloudflared); byo consent moved to the typed
# disclosure ack (the WG tunnel is now on that list) + the tester-agreement disclosure line.
if [[ "$MODE" == "byo" ]]; then
AGENT_CONFIG="$AGENT_CONFIG" python3 <<'PY' || die "byo config assert FAILED — see the keys above; a byo box must keep lan_resolver/oob off and deployment_mode=byo (fix the preserved config and re-run with --resume)"
import json, os, sys
d = json.load(open(os.environ['AGENT_CONFIG']))
bad = []
if d.get('lan_resolver', {}).get('enable'): bad.append('lan_resolver.enable (takes over host DNS on :53)')
if d.get('oob', {}).get('enabled'): bad.append('oob.enabled (starts a second operator sshd)')
if d.get('deployment_mode') == 'appliance': bad.append('deployment_mode=appliance (unlocks host-service self-heal on a host we do not own)')
if bad:
print('byo-forbidden config keys are TRUE: ' + '; '.join(bad), file=sys.stderr)
sys.exit(1)
PY
log_success " byo config asserts passed (lan_resolver.enable / oob.enabled off; deployment_mode=byo; wg_tunnel intentionally ON — base infrastructure)"
fi
# health: read-only selftest (proxmox) must pass before provisioning
if ! felhom-agent --config "$AGENT_CONFIG" --selftest >/dev/null 2>&1; then
felhom-agent --config "$AGENT_CONFIG" --selftest 2>&1 | tail -20 >&2
die "agent --selftest FAILED with the new config (token/ACL/TLS problem) — fix before provisioning"
fi
log_success " agent --selftest (read-only) passed"
# start the daemon (host-report loop) as the felhom-agent user and confirm it stays up. is-active is
# the real proof the NON-root user can read the 0600 config (the root selftest above can't show that).
if systemctl list-unit-files felhom-agent.service >/dev/null 2>&1; then
run systemctl enable felhom-agent >/dev/null 2>&1 || true
run systemctl restart felhom-agent
if ! $DRY_RUN; then
sleep 3
if systemctl is-active --quiet felhom-agent; then
log_success " felhom-agent service active (non-root $AGENT_USER reads the config OK)"
else
systemctl status felhom-agent --no-pager -l 2>&1 | tail -20 >&2
journalctl -u felhom-agent -n 20 --no-pager 2>&1 | tail -20 >&2
die "felhom-agent did not stay active after restart — see status/journal above"
fi
fi
else
log_warn " no felhom-agent systemd unit — daemon host-report loop not started (provision one-shot still works)"
fi
_state_mark agent_config
}
#-------------------------------------------------------------------------------
# STEP 7 — golden: ensure a restorable golden archive (local else Gitea-fetched + verified)
#-------------------------------------------------------------------------------
# Local auto-discovery is the default + fallback. When no local golden exists (or --force-gitea-golden),
# fetch the golden from Gitea (git token), VERIFY its sha256 against the hub manifest, and import it
# into the archive storage's dump dir under a valid vzdump name so the provision restore can use it.
step_golden() {
log_step "7/8 golden archive"
if [[ -n "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then
log_skip " using local golden: $GOLDEN_VOLID"
_state_mark golden; return 0
fi
# Need the manifest + git creds (already resolved in step 5, but re-resolve on a fresh --resume run).
[[ -n "$ART_GOLDEN_VER" ]] || resolve_artifacts
[[ -n "$GIT_TOKEN" ]] || resolve_git_creds
[[ -n "$ART_GOLDEN_VER" && -n "$ART_GOLDEN_SHA" ]] || die "hub manifest has no golden version/sha256 — set it in the operator UI, or pass --golden VOLID"
local url="$GITEA_BASE/api/packages/$GITEA_OWNER/generic/felhom-golden/$ART_GOLDEN_VER/golden.tar.zst"
if $DRY_RUN; then
log_dry "$(_dry_auth_hint) -o <dump>/vzdump-lxc-${GOLDEN_VMID}-<ts>.tar.zst $url ; verify sha256=$ART_GOLDEN_SHA ; set GOLDEN_VOLID"
GOLDEN_VOLID="${ARCHIVE_STORAGE}:backup/vzdump-lxc-${GOLDEN_VMID}-<dry-run>.tar.zst"
_state_mark golden; return 0
fi
# Resolve the archive storage's dump dir (pvesm path maps a volid → fs path without needing it to exist).
local dump_dir fname dest
dump_dir=$(dirname "$(pvesm path "${ARCHIVE_STORAGE}:backup/vzdump-lxc-${GOLDEN_VMID}-2000_01_01-00_00_00.tar.zst" 2>/dev/null)")
[[ -d "$dump_dir" ]] || die "could not resolve dump dir for storage $ARCHIVE_STORAGE (got '$dump_dir')"
fname="vzdump-lxc-${GOLDEN_VMID}-$(date +%Y_%m_%d-%H_%M_%S).tar.zst"
dest="${dump_dir}/${fname}"
log_info " fetching golden v$ART_GOLDEN_VER from Gitea → $dest"
fetch_verify "$url" "$dest" "$ART_GOLDEN_SHA"
GOLDEN_VOLID="${ARCHIVE_STORAGE}:backup/${fname}"
pvesm list "$ARCHIVE_STORAGE" --content backup 2>/dev/null | grep -q "$fname" \
|| log_warn " imported golden not yet listed by pvesm (continuing — restore uses the volid directly)"
log_success " golden imported + verified: $GOLDEN_VOLID"
_state_mark golden
}
#-------------------------------------------------------------------------------
# STEP 8 — provision (golden restore -> resize -> bootstrap.json -> onboot:1)
#-------------------------------------------------------------------------------
step_provision() {
log_step "8/8 provision guest $VMID"
# NOTE: -hub-password is passed on argv (the agent's only input for it) — briefly
# visible in ps. Tracked as an Observation (candidate: env/stdin in the agent).
# Optional operator CPU/RAM cap — passed to the agent ONLY when set (an agent < v0.52.0 would
# reject the unknown flag and die; opt-in means no one hits that until they intentionally cap).
local -a cap_args=()
[[ -n "$CPU_CORES" ]] && cap_args+=(-cores "$CPU_CORES")
[[ -n "$MEM_MIB" ]] && cap_args+=(-memory "$MEM_MIB")
# felhom pool: ensure it exists before the restore (step_token already created it; this is a
# belt-and-suspenders no-op that also covers a --resume path). The AGENT adds the guest to the pool
# atomically via restore --pool (v0.53.0) — no separate script-side pool_add_guest.
ensure_felhom_pool
if $DRY_RUN; then
log_dry "felhom-agent --config $AGENT_CONFIG --selftest=provision -archive $GOLDEN_VOLID -vmid $VMID -customer-id $CUSTOMER_ID -hub-password <pass> -rootfs-grow $ROOTFS_GROW -datavol-grow $DATAVOL_GROW -sysdata-grow $SYSDATA_GROW ${cap_args[*]} # agent restores INTO pool $PVE_POOL"
log_dry "record provisioned_vmid=$VMID in $STATE_FILE (for a later automatic --uninstall)"
_state_mark provision; return 0
fi
if ! felhom-agent --config "$AGENT_CONFIG" --selftest=provision \
-archive "$GOLDEN_VOLID" -vmid "$VMID" \
-customer-id "$CUSTOMER_ID" -hub-password "$PASSPHRASE" \
-rootfs-grow "$ROOTFS_GROW" -datavol-grow "$DATAVOL_GROW" -sysdata-grow "$SYSDATA_GROW" \
"${cap_args[@]}"; then
die "provision FAILED — see the agent error above. Fix and re-run with --resume."
fi
log_success " provision completed"
_state_mark provision
# Record the provisioned vmid so a later --uninstall resolves the target automatically + safely.
_state_put provisioned_vmid "$VMID"
# (No pool_add_guest here — the agent's restore --pool already made the guest a member.)
# Reboot the guest ONCE: the golden's controller-bootstrap unit evaluates its
# ConditionPathExists=/etc/felhom-bootstrap/bootstrap.json at BOOT, and the back-half attaches
# the bootstrap mount to the ALREADY-RUNNING guest — without a reboot the unit stays skipped and
# the controller never deploys (drill finding R6; the agent's own provision output says
# "next: reboot the guest"). On fast hosts the first boot sometimes wins the race — the reboot
# is idempotent either way (the unit no-ops when the controller already runs).
log_info " rebooting guest $VMID so the baked controller-bootstrap unit picks up the mount"
run pct reboot "$VMID"
}
#-------------------------------------------------------------------------------
# STEP 7 — verify
#-------------------------------------------------------------------------------
step_verify() {
log_step "verify"
if $DRY_RUN; then log_dry "pct status/config $VMID; docker ps in-guest; host-report includes $VMID; assert guest $VMID in pool $PVE_POOL + scoped ACL grants present (user+token: Base@/, Guest@/pool/$PVE_POOL, Store@[${PVE_STORAGES[*]}])"; return 0; fi
local ok=true
local st; st=$(pct status "$VMID" 2>/dev/null | awk '{print $2}')
[[ "$st" == "running" ]] && log_success " pct status: running" || { log_error " pct status: $st"; ok=false; }
if pct config "$VMID" 2>/dev/null | grep -q '^onboot: 1'; then log_success " onboot: 1"; else log_error " onboot NOT 1"; ok=false; fi
pct config "$VMID" 2>/dev/null | grep -E '^(rootfs|mp0|mp1|mp8):' | sed 's/^/ /'
# GL-2 pool + ACL asserts (BOTH modes — campaign-2 R2 proved membership can silently drop, and a
# guest outside the pool is a guest the scoped token cannot manage).
local members; members=" $(pool_members || true) "
if [[ "$members" == *" $VMID "* ]]; then
log_success " pool: guest $VMID is a member of $PVE_POOL"
else
log_error " pool: guest $VMID NOT in pool $PVE_POOL (the scoped token cannot manage it — adopt with --adopt-pool)"; ok=false
fi
local _pairs=("/ $PVE_ROLE_BASE" "/pool/$PVE_POOL $PVE_ROLE_GUEST") _pair _apath _arole _acls2
for _acls2 in "${PVE_STORAGES[@]}"; do _pairs+=("/storage/$_acls2 $PVE_ROLE_STORE"); done
# F-LEAK: assert the scratch band too — a missing grant here is invisible until a restore-test
# FAILS, which is exactly the case that leaked a guest before v1.21.0.
local _sv; for ((_sv = PVE_SCRATCH_VMID_MIN; _sv <= PVE_SCRATCH_VMID_MAX; _sv++)); do _pairs+=("/vms/$_sv $PVE_ROLE_GUEST"); done
for _pair in "${_pairs[@]}"; do
_apath="${_pair% *}"; _arole="${_pair#* }"
if _acl_grant_present "$_apath" user "$PVE_USER" "$_arole" \
&& _acl_grant_present "$_apath" token "${PVE_USER}!${PVE_TOKENID}" "$_arole"; then
log_success " acl: $_arole@$_apath present (user+token)"
else
log_error " acl: $_arole@$_apath MISSING for the user and/or the token (re-apply with --rescope-acl)"; ok=false
fi
done
# GL-4: operator-key pin state. Dormant is the SAFE default, not an error — WARN only.
local _signers
_signers=$(python3 -c "import json;print(len(json.load(open('$AGENT_CONFIG')).get('authz',{}).get('signers') or []))" 2>/dev/null || echo 0)
if [[ "${_signers:-0}" -ge 1 ]]; then
log_success " authz signers: $_signers (operator-signed self-update armed)"
else
log_warn " no operator key pinned — agent self-update stays dormant (pin via --operator-pubkey-file or the script's OPERATOR_KEY_* constants; the key ceremony is an operator step)"
fi
# controller container healthy in-guest — bounded wait (the post-provision reboot + docker start
# take a while, especially on modest hardware; drill R6 re-verify)
local cstat="" _waited=0
while [[ -z "$cstat" && $_waited -lt 180 ]]; do
cstat=$(pct exec "$VMID" -- docker ps --filter name=felhom-controller --format '{{.Status}}' 2>/dev/null | head -1)
[[ -n "$cstat" ]] || { sleep 5; _waited=$((_waited+5)); }
done
if [[ -n "$cstat" ]]; then log_success " controller: $cstat (after ~${_waited}s)"; else log_warn " controller container not visible after ${_waited}s — check 'pct exec $VMID -- journalctl -u felhom-controller-bootstrap'"; fi
local cver; cver=$(pct exec "$VMID" -- docker ps --filter name=felhom-controller --format '{{.Image}}' 2>/dev/null | head -1)
[[ -n "$cver" ]] && log_info " controller image: $cver"
# tunnel
local tun; tun=$(pct exec "$VMID" -- docker ps --filter name=cloudflared --format '{{.Status}}' 2>/dev/null | head -1)
[[ -n "$tun" ]] && log_info " cloudflared: $tun" || log_warn " cloudflared not visible yet"
# host-report includes the guest (best-effort via the agent's hub selftest)
log_info " (confirm in the hub UI that host $HOST_ID reports guest $VMID)"
_state_mark verify
echo ""
if $ok; then
log_success "Day-0 provision SUCCESS — vmid=$VMID host_id=$HOST_ID customer=$CUSTOMER_ID golden=$GOLDEN_VOLID"
else
log_warn "Day-0 provision completed WITH WARNINGS — review the checks above"
fi
# Drill F-8: repeat the break-glass pointer in the FINAL summary (the 4b line scrolls away).
if _state_has break_glass; then
log_info "root@pam was rotated + vaulted at step 4b — retrieve at hub → host page (the old GUI password no longer works)."
fi
}
#-------------------------------------------------------------------------------
# verify (agent-only, for --skip-provision): the agent is installed, runs non-root, and reports.
#-------------------------------------------------------------------------------
step_verify_agent() {
log_step "verify (agent only)"
if $DRY_RUN; then log_dry "felhom-agent --version; systemctl is-active felhom-agent; --selftest=hub (one collect+report)"; return 0; fi
local ok=true
log_info " binary: $("$AGENT_BIN" --version 2>&1 | head -1)"
log_info " runs as: $(systemctl show felhom-agent -p User --value 2>/dev/null) (want $AGENT_USER)"
if systemctl is-active --quiet felhom-agent; then log_success " service active"; else log_error " service NOT active"; ok=false; fi
# one explicit collect+report to prove the hub link end-to-end (host-report lands).
if felhom-agent --config "$AGENT_CONFIG" --selftest=hub >/dev/null 2>&1; then
log_success " --selftest=hub OK (a host-report reached the hub)"
else
log_warn " --selftest=hub did not confirm (the daemon loop still reports every poll_seconds)"
fi
_state_mark verify
echo ""
if $ok; then
log_success "Agent install SUCCESS — $("$AGENT_BIN" --version 2>&1 | head -1) as $AGENT_USER, host_id=$HOST_ID customer=$CUSTOMER_ID"
else
log_warn "Agent install completed WITH WARNINGS — review the checks above"
fi
}
#-------------------------------------------------------------------------------
# Main
#-------------------------------------------------------------------------------
# Scrub secret carriers on any exit; under --preflight-only also print the explicit FAIL verdict
# (the PASS verdict is printed by the preflight-only branch below).
# shellcheck disable=SC2154 # _rc IS assigned first inside the trap string itself
trap '_rc=$?; PASSPHRASE=""; PVE_TOKEN=""; HOST_API_KEY=""; GIT_TOKEN=""; if [[ $_rc -ne 0 ]] && $PREFLIGHT_ONLY; then log_error "PRE-FLIGHT FAIL (exit $_rc) — fix the finding above and re-run"; fi' EXIT
# --preflight-only: the full (mode-aware) step-1 checks, an explicit verdict, exit — no state marks,
# no later step (GL-2).
if $PREFLIGHT_ONLY; then
step_preflight
echo ""
log_success "PRE-FLIGHT PASS (mode=${MODE:-unset}) — no state written, no install step executed"
exit 0
fi
if $RESUME && _state_has preflight; then
# still need the passphrase for enroll/provision even on resume
read_passphrase
# re-resolve cheap derived values skipped steps would have set
[[ -n "$NODE" ]] || NODE=$(pvesh get /nodes --output-format json 2>/dev/null | python3 -c "import json,sys;print(json.load(sys.stdin)[0]['node'])" 2>/dev/null || hostname)
if [[ -z "$AGENT_CONFIG" ]] && systemctl cat felhom-agent >/dev/null 2>&1; then
AGENT_CONFIG=$(systemctl cat felhom-agent 2>/dev/null | grep -oP '(?<=-config )\S+' | head -1)
fi
[[ -n "$AGENT_CONFIG" ]] || AGENT_CONFIG="/etc/felhom-agent/agent.json"
# Backfill display values from the already-written config so the summary is complete.
[[ -f "$AGENT_CONFIG" ]] && HOST_ID=$(python3 -c "import json;print(json.load(open('$AGENT_CONFIG')).get('hub',{}).get('host_id',''))" 2>/dev/null || true)
# GL6-F4: golden's GOLDEN_VOLID feeds provision, but the resume path skips preflight (where local
# auto-discovery sets it). If the golden step already completed, the archive is on the local
# storage from run 1 — re-derive the volid so provision doesn't get an empty -archive. (When
# golden hasn't completed yet, step_golden runs and sets it.)
if [[ -z "$GOLDEN_VOLID" ]] && _state_has golden; then
GOLDEN_VOLID=$(pvesm list "$ARCHIVE_STORAGE" --content backup 2>/dev/null | awk -v v="$GOLDEN_VMID" '$0 ~ ("vzdump-lxc-" v "-"){print $1}' | sort | tail -1)
[[ -n "$GOLDEN_VOLID" ]] && log_info " golden (resumed from local): $GOLDEN_VOLID"
fi
log_skip "pre-flight (resumed)"
else
step_preflight
fi
# GL6-F4: token/enroll/grows are PRODUCERS whose IN-MEMORY outputs (PVE_TOKEN; HOST_ID/HOST_API_KEY;
# ROOTFS_GROW/DATAVOL_GROW/SYSDATA_GROW) are consumed by later steps (agent_config writes the token +
# hub.host_id into the config; provision passes the grows as flags). They MUST run every pass — even
# on --resume — or a resumed install writes a config missing hub.host_id/proxmox.token (daemon
# crash-loops) and provision gets `-rootfs-grow ""` (flag parse error). All three are idempotent:
# token reuses the on-disk token if it still authenticates (else rotates), enroll is a mint-once-
# reuse POST (200 REUSED), grows is a pure recompute — so unconditional re-run is cheap + correct.
# (This is why they are NOT behind `should_skip … ||`; step_token's own `_state_has` fast-path still
# short-circuits the reuse check when the token is already in memory within one process.)
step_token
step_grows
step_enroll
# GL-2: break-glass is gated at the CALL SITE (not inside the step) so the byo skip is auditable in
# one place. byo = a host the operator does not own: root@pam is the OWNER's credential — never
# reset, never vaulted.
if [[ "$MODE" == "appliance" ]]; then
should_skip break_glass || step_break_glass
else
log_skip "break-glass (step 4b) SKIPPED (byo: the host owner manages break-glass; root@pam untouched, nothing vaulted)"
fi
should_skip agent_install || step_agent_install
should_skip agent_config || step_agent_config
should_skip golden || step_golden
if $SKIP_PROVISION; then
log_skip "provision (--skip-provision) — agent install/config verified only"
step_verify_agent
else
should_skip provision || step_provision
step_verify
fi