bee6848458
gates / gates (push) Successful in 8s
Two channels moved off main in the same change, because either one left behind makes the other cosmetic. Channel 1 — the served script. webpage.yaml git-synced /scripts/ from --branch=main every 30s and nginx served that tree, so pushing this file WAS publishing it: within half a minute it was what every new machine downloaded and ran as root, with no staging and no rollback but another push. The sync is now SPLIT: the website keeps tracking main at the same cadence (a copy edit must never need a release) and /scripts/ tracks the tag installer-v<SCRIPT_VERSION>. PROVEN before the manifest was touched: git-sync v4.4.0 follows a tag AND notices a MOVED one — measured on a throwaway sync against this repo, "update required ... local:<old> remote:<new>" -> "updated successfully", within one period. The moved-tag half is what the publish model rests on. Channel 2 — the sixteen files fetched at run time. fetch_raw pulled from $AGENT_REPO/raw/branch/main; it now pulls raw/tag/v$ART_AGENT_VER. That is a correctness fix, not only a channel one (R-183): a fresh install fetched the vouched agent BINARY while taking its unit file, sudoers and guarded wrappers from whatever main held. Two refs, one install, nothing compared them. Their correct ref was never SCRIPT_VERSION — they do not live in this repo. No fallback to a branch: a vouched version whose tag is missing fails loudly rather than quietly serving main. Channel 3 — the URL — needed no change, recorded rather than left silent: https://felhom.eu/scripts/felhom-host-install.sh never carried a ref, so both producers follow the tag with no edit. No hub change, no hub version bump. Gate 6 in hostinstall_gates.py pins all three structurally with no network, so it stays in --fast and runs in CI. It deliberately does NOT assert "a tag exists for the current SCRIPT_VERSION": that would go red on the very push that bumps the version, before publishing — and publishing being separate is the ruling.
2835 lines
172 KiB
Bash
2835 lines
172 KiB
Bash
#!/bin/bash
|
|
#===============================================================================
|
|
# felhom-host-install.sh — version: the SCRIPT_VERSION variable below (single source;
|
|
# `-h` and the run banner both read it — the v1.14.0/v1.13.0 header/banner drift was drill F-1).
|
|
# Day-0 host-bootstrap for a Felhom Proxmox host (operator-deploy model).
|
|
#
|
|
# Run by the operator on a FRESHLY-PVE-INSTALLED box (after a manual PVE install
|
|
# + SSH in). Given a customer-id + retrieval passphrase, it fully automates
|
|
# Day-0: Proxmox API token -> hub host enrollment -> AGENT INSTALL (fetch from
|
|
# Gitea + verify sha256 + install) -> agent config -> golden -> guest provision
|
|
# -> verify. It composes already-proven mechanisms (the pveum role/token
|
|
# sequence, hub POST /host-enroll [option C], felhom-agent --selftest=provision).
|
|
# The agent renders bootstrap.json and the controller pulls its own
|
|
# controller.yaml in-guest; this script does NOT fetch that.
|
|
#
|
|
# v1.1.0 (BUNDLE slice): the agent binary + golden are now fetched from Gitea
|
|
# generic packages and VERIFIED against the hub-vouched artifact manifest
|
|
# (GET /api/v1/artifacts/{id}) before install/use. The fetch credential is the
|
|
# git token already inside the customer's controller.yaml (config-retrieve) — NO
|
|
# new credential. The checksum trust root is the HUB, not Gitea. This removes the
|
|
# old prerequisite "install the agent binary + unit manually".
|
|
#
|
|
# v1.15.0 (DR-tier-by-default, DRILL-day0-vm-2026-07-12 sweep): every install now deploys the FULL
|
|
# DR plumbing uniformly — capability is baked, activation is a hub flag. F-7: ships the
|
|
# `felhom-pbs-apply` wrapper (like mkfs/selfupdate — pbsdr capabilities no longer born DEGRADED).
|
|
# F-10: installs `age` (escrow-ceremony identity-wrap dependency). F-9: renders
|
|
# `wg_tunnel.enabled: true` (WG is BASE infrastructure like cloudflared, not a DR component —
|
|
# operator decision 2026-07-12; byo consent moves to the disclosure ack + tester agreement, the
|
|
# byo wg-off config assert is retired). The default --acl-storages set keeps `felhom-pbs`
|
|
# (drill-proven: narrowing it 403s the PBS-DR apply-bridge later); a missing felhom-pbs storage
|
|
# is expected pre-tier and no longer a byo refusal. F-1: single VERSION source (SCRIPT_VERSION).
|
|
# F-2: dry-run fetch lines show the real auth mode (anonymous vs -u). F-8: step 4b + the final
|
|
# summary now say OUT LOUD that root@pam was rotated + where to retrieve it.
|
|
# v1.14.0 (CAMPAIGN-3 Part 6): --mode now also templates the agent's top-level `deployment_mode`
|
|
# ("appliance" | "byo") — authoritative from --mode, written even over a --preserve-from config, so a
|
|
# byo box can never inherit an appliance host-service self-heal. Gates the agent's node self-heal
|
|
# (host networking recovery at boot, F12-class; agent v0.85.0). The byo config assert now also refuses
|
|
# deployment_mode=appliance. absent/typo → byo (fail-safe).
|
|
# v1.10.0 (GL-2, go-live G2/G4/G5): explicit --mode appliance|byo install profile — the flag is now
|
|
# REQUIRED for a fresh install (no default; build-golden v2.0.0 precedent: defaults rot). byo =
|
|
# BYO-host hardening for a Proxmox host the operator does NOT own: break-glass (step 4b) gated OFF
|
|
# at its call site (root@pam is never touched), mandatory --cores/--memory noisy-neighbor caps,
|
|
# strict preflight (PVE-major gate, --acl-storages existence, host-DNS :53 info line, host-mutation
|
|
# disclosure + typed-hostname acknowledgement), byo config asserts (lan_resolver/wg_tunnel/oob must
|
|
# stay off), pool+ACL verify asserts (BOTH modes — campaign-2 R2 lesson), and --preflight-only.
|
|
# Test harness: scripts/hostinstall-mode-harness.sh (static tier runs anywhere; PVE tier dry-only).
|
|
#
|
|
# v1.11.0 (GL-4, go-live G6 + the G1 key-pin fold): (A) uninstall gap-closure — the teardown now
|
|
# also removes the self-update artifacts (guarded wrapper, .prev/.new.* A/B slots, rollback unit,
|
|
# start-limit drop-in), unmounts every enrolled drive under /mnt/felhom-drives (plain umount only,
|
|
# NEVER -l/-f; busy = warn + guidance, data always stays on the drive), and ends with an explicit
|
|
# KEPT-vs-WIPED statement (PBS backups, hub record, escrow, vaulted recovery credential live on).
|
|
# (B) operator-key day-0 fold — OPERATOR_KEY_* constants (empty until the pin ceremony) and
|
|
# --operator-pubkey-file write authz.signers into the agent config at step 6; empty keys keep a
|
|
# preserved config's signers (never clobber a manually-pinned box); verify reports armed/dormant.
|
|
#
|
|
# v1.11.2 (GL-6 Gate-0 operator ruling, G3): ANONYMOUS artifact fetch supported — empty
|
|
# git.username/git.token in the customer config now WARNS and fetches anonymously (the Felhom
|
|
# artifacts are world-readable by ruling; the hub-vouched sha256 stays the integrity root).
|
|
# A configured credential is still used when present.
|
|
#
|
|
# v1.11.3 (GL-6 finding F4): --resume correctness — the producer steps (token/enroll/grows) now
|
|
# run every pass so their in-memory outputs (pve token, hub host_id/api_key, volume grows) are
|
|
# repopulated for the later steps that consume them; a resumed install no longer writes a config
|
|
# missing hub.host_id/proxmox.token (daemon crash-loop) or passes an empty -rootfs-grow. golden's
|
|
# GOLDEN_VOLID is re-derived from the local archive on resume.
|
|
#
|
|
# v1.12.0 (GL-8, BYO coexistence hardening — GL-6 findings): F3 — the populated-host leaf guard
|
|
# now fires only when a FELHOM guest exists (felhom_guests), so a fresh byo install on a host that
|
|
# only runs the OWNER's guests proceeds without --allow-new-leaf (the pin protection for real Felhom
|
|
# guests is unchanged). F6 — byo preflight REFUSES (never mutates) when a foreign resolver is bound
|
|
# to :53, with remediation (Felhom needs the guest reachable by name; it won't stop the owner's DNS).
|
|
# F1 — uninstall now purges the agent config's .bak* siblings (one held a live hub api_key), not just
|
|
# agent.json.
|
|
#
|
|
# Grounding: documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md
|
|
#
|
|
# Usage:
|
|
# sudo ./felhom-host-install.sh --customer-id ID --mode appliance|byo [options]
|
|
#
|
|
# Required:
|
|
# --customer-id ID Customer (must already exist in the hub)
|
|
# --mode appliance|byo Install profile (REQUIRED for install/--resume — no default):
|
|
# appliance a Felhom-owned/managed box. Full Day-0 incl. the
|
|
# break-glass root@pam credential (step 4b).
|
|
# byo a host the operator does NOT own (BYO / pilot). NEVER
|
|
# touches root@pam (step 4b skipped), REQUIRES explicit
|
|
# --cores AND --memory, refuses --enable-oob and
|
|
# --rotate-recovery, dies on non-9.x PVE (unless --force)
|
|
# and on --acl-storages entries absent from this box
|
|
# (except felhom-pbs — the PBS-DR tier creates it later),
|
|
# prints the full host-mutation list and requires typing
|
|
# the host's short hostname, and asserts the agent config
|
|
# keeps lan_resolver/oob OFF (wg_tunnel is ON in BOTH
|
|
# modes since v1.15.0 — base infra, disclosed in the ack).
|
|
# (dr the 10D DR stub — documented seam, not implemented.)
|
|
#
|
|
# Options:
|
|
# --preflight-only run ONLY the (mode-aware) step-1 checks: prints PRE-FLIGHT PASS/FAIL,
|
|
# exits 0/1, writes NO state, executes NO later step. Re-runnable; valid
|
|
# with either mode (mode-specific checks need --mode).
|
|
# --hub-url URL default https://hub.felhom.eu
|
|
# --vmid N guest VMID to provision. Default 9201; if omitted and 9201 is already
|
|
# in use, the script auto-picks the next free id (pct+qm) and asks to
|
|
# confirm. An EXPLICIT --vmid stays deterministic (dies unless --force).
|
|
# --golden VOLID golden archive volid (default: newest vzdump of the
|
|
# golden build VMID on the archive storage; else fetched
|
|
# from Gitea per the hub artifact manifest)
|
|
# --golden-vmid N golden build guest vmid for auto-discovery (default 9100)
|
|
# --archive-storage NAME storage holding the golden vzdump (default local)
|
|
# --force-gitea-golden ignore any local golden; fetch+verify the golden from
|
|
# Gitea (proves the fetch path; used by the live test)
|
|
# --node NAME PVE node name (default: pvesh /nodes, else hostname)
|
|
# --bridge-ip IP[:PORT] local-api listen addr; an explicit value overrides the island default
|
|
# (default: appliance → island 169.254.253.1:8443; byo → vmbr0 IP:8443)
|
|
# --no-island appliance only: keep the historical LAN bind instead of the R-50 island
|
|
# --rootfs-grow N grow OS rootfs by N GiB (default: auto-compute)
|
|
# --datavol-grow N grow the single data volume by N GiB (default: auto-compute from the pool)
|
|
# --sysdata-grow N DEPRECATED (R-165): added to --datavol-grow; there is one volume now
|
|
#
|
|
# Guest cap (appliance: optional — protect a SHARED host's other guests; byo: BOTH REQUIRED —
|
|
# the only noisy-neighbor protection on a host you do not own; needs agent >= v0.52.0):
|
|
# --cores N cap the guest to N CPU cores (0/unset = golden default)
|
|
# --memory M cap the guest RAM to M MiB (0/unset = golden default)
|
|
#
|
|
# --passphrase-file PATH read the retrieval passphrase from a 0600 file
|
|
# (default: secure no-echo prompt)
|
|
# --operator-pubkey-file PATH pin the operator signing PUBLIC keys at day-0 (GL-4). One key per
|
|
# line: "operational <authorized_keys line>" or "recovery <authorized_keys
|
|
# line>" (# comments/blank ok; the authorized_keys comment field is the
|
|
# key_id and is required). Overrides the script's OPERATOR_KEY_* constants.
|
|
# Written to authz.signers in the agent config; no keys resolved = a
|
|
# preserved config's signers are KEPT, else self-update stays dormant.
|
|
# --preserve-from PATH merge non-Day-0 sections (privileged/storage/backup/
|
|
# local_api/authz/lan_resolver) from an existing config
|
|
# --preserve-state-from PATH carry the prior agent leaf+key+token-store (local-api.crt/key,
|
|
# local-tokens.log) over so the pinned fingerprint STAYS STABLE across a
|
|
# reinstall (no controller re-bootstrap). Use an aside copy of the old
|
|
# /var/lib/felhom-agent.
|
|
# --allow-new-leaf opt in to REGENERATE the agent leaf on a host that already has guests
|
|
# (the populated-host guard otherwise refuses; every guest must then be
|
|
# re-bootstrapped — only use intentionally).
|
|
# --force allow provisioning over an EXISTING vmid (destructive)
|
|
# --skip-provision install + configure + verify the agent, but do NOT
|
|
# provision a guest (re-install/upgrade an agent on a host
|
|
# that already has live guests; also the agent-only path)
|
|
# --dry-run print every mutating command without executing
|
|
# --resume skip steps already recorded in the state file
|
|
# -h, --help this help
|
|
#
|
|
# Uninstall (local host teardown — no hub contact, no passphrase):
|
|
# --uninstall cleanly revert an install: destroy the Felhom guest, remove the agent
|
|
# (unit/sudoers/binary/state/config/user + runtime artifacts: shared-parent
|
|
# unit, mkfs + pbs-apply wrappers, hook snippet, dnsmasq snippets), the pveum
|
|
# role/user/token/ACL, and the install state file. Refuses a non-Felhom guest (no
|
|
# /etc/felhom-bootstrap mount) and skips host-level removal if OTHER Felhom
|
|
# guests remain (both overridable with --force). Typed vmid confirmation
|
|
# required. Reuses --vmid (else the recorded provisioned_vmid), --force,
|
|
# --archive-storage, --golden-vmid, --dry-run.
|
|
# --remove-golden with --uninstall, also delete the golden vzdump from the archive storage
|
|
#
|
|
# Retrofit (local, non-destructive — no hub contact, no passphrase):
|
|
# --adopt-pool add an EXISTING Felhom guest to the `felhom` pool (creates the pool if
|
|
# needed). Resolves the guest from --vmid else the recorded provisioned_vmid;
|
|
# refuses a non-Felhom guest unless --force. Touches ONLY pool membership —
|
|
# never reconfigures/restarts the guest. (A fresh provision joins the pool
|
|
# automatically; this retrofits already-installed boxes.)
|
|
# --rescope-acl migrate an existing install from the pre-3b broad-`/` token grant to the
|
|
# pool-scoped ACL (Guest@/pool/felhom + Store@each storage + Sys.Audit/SDN.Use@/).
|
|
# ACL-only (no data touched). SUPERVISED: run with felhom-agent STOPPED, then
|
|
# deploy agent >= v0.53.0, then start (the scoped ACL + pool-param agent are
|
|
# mutually dependent).
|
|
# --acl-storages "a b c" override the storages the scoped ACL grants Datastore.* on
|
|
# (default: "local local-lvm felhom-pbs"). Used by fresh install + --rescope-acl.
|
|
#
|
|
# State (idempotent/resumable): /var/lib/felhom-install/state.json
|
|
# Agent config written 0600 to the systemd unit's -config path
|
|
# (auto-detected; else /etc/felhom-agent/agent.json).
|
|
#
|
|
# SECURITY: the passphrase is read no-echo or from a 0600 file — never a CLI arg,
|
|
# never echoed, never written to the state file or logs. The minted pve-token
|
|
# secret + per-host hub api_key live ONLY in the agent config (0600, root).
|
|
#===============================================================================
|
|
|
|
set -euo pipefail
|
|
|
|
SCRIPT_VERSION="1.23.0" # the SINGLE version source (F-1): -h and the run banners follow it.
|
|
# The hub used to carry a copy for its Setup tab; R-94 DELETED it
|
|
# (2026-08-02) because the hub cannot know which version a box runs —
|
|
# the Setup command fetches this script at run time. scripts/
|
|
# hostinstall_gates.py now asserts the hub carries NO version literal.
|
|
# 1.16.0: the FELHOM_ESCROW sudoers alias (controller-driven escrow
|
|
# ceremony, agent v0.88.0) ships on every install/update — it rides the
|
|
# CANONICAL sudoers fetch below (configs/felhom-agent.sudoers from the
|
|
# agent repo, visudo-gated), no separate installer step.
|
|
|
|
# Operator signing keys pinned at day-0 (GL-4; doc 04 §3 two-key model). PINNED below to the real
|
|
# ceremony keypairs (the pin ceremony is DONE — felhom-op-1 / felhom-rec-1); --operator-pubkey-file
|
|
# at install time overrides them. If these were ever cleared to empty, no authz.signers is written
|
|
# and agent self-update stays DORMANT (the safe fallback; the verify step warns). PUBLIC keys only —
|
|
# this script never generates, reads, or references private key material.
|
|
OPERATOR_KEY_OPERATIONAL_ID="felhom-op-1"
|
|
OPERATOR_KEY_OPERATIONAL_LINE="ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIL8z0qCNgA3x2xxAB0Qj5ro8waFjGZ8Ta/sWB63tlLw+ felhom-op-1"
|
|
OPERATOR_KEY_RECOVERY_ID="felhom-rec-1"
|
|
OPERATOR_KEY_RECOVERY_LINE="ssh-ed25519 AAAAC3NzaC1lZDI1NTE5AAAAIC/gZI5mvXmUzImh81iHdjdVYwmQ4fRGKqn+9td0fI5/ felhom-rec-1"
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# Logging (mirrors felhom-controller/scripts/docker-setup.sh)
|
|
#-------------------------------------------------------------------------------
|
|
RED='\033[0;31m'; GREEN='\033[0;32m'; YELLOW='\033[1;33m'
|
|
BLUE='\033[0;34m'; CYAN='\033[0;36m'; NC='\033[0m'
|
|
log_info() { echo -e "${GREEN}[INFO]${NC} $1"; }
|
|
log_warn() { echo -e "${YELLOW}[WARN]${NC} $1"; }
|
|
log_error() { echo -e "${RED}[ERROR]${NC} $1" >&2; }
|
|
log_step() { echo -e "${BLUE}[STEP]${NC} $1"; }
|
|
log_success() { echo -e "${GREEN}[OK]${NC} $1"; }
|
|
log_skip() { echo -e "${CYAN}[SKIP]${NC} $1"; }
|
|
log_dry() { echo -e "${CYAN}[DRY-RUN]${NC} $1"; }
|
|
die() { log_error "$1"; exit 1; }
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# Defaults
|
|
#-------------------------------------------------------------------------------
|
|
CUSTOMER_ID=""
|
|
MODE="" # --mode appliance|byo (GL-2: REQUIRED for install/--resume, no default; "dr" = the 10D stub)
|
|
PREFLIGHT_ONLY=false # --preflight-only: run the mode-aware step-1 checks only; no state writes, no later step
|
|
HUB_URL="https://hub.felhom.eu"
|
|
VMID="9201"
|
|
VMID_EXPLICIT=false # set true when --vmid is given; gates the auto-pick-a-free-vmid behavior
|
|
GOLDEN_VOLID=""
|
|
GOLDEN_VMID="9100"
|
|
ARCHIVE_STORAGE="local"
|
|
NODE=""
|
|
NODE_EXPLICIT=false # set true when --node is given; gates the multi-node wrong-node guard
|
|
BRIDGE_ADDR=""
|
|
# R-50 island control plane (SPIKE-island-bridge-2026-07-25, GO). The appliance default binds the
|
|
# local API on a host-internal, portless bridge with a fixed private /30 so the controller→agent
|
|
# channel survives any LAN/DHCP/site move (the F1 fix). These are the spike's VALIDATED parameters —
|
|
# do not change them. Set ISLAND_ON=1 only in appliance mode with no explicit --bridge-ip override
|
|
# (byo and an explicit --bridge-ip keep the historical LAN bind). --no-island opts an appliance out.
|
|
ISLAND_BRIDGE="vmbr9"
|
|
ISLAND_HOST_CIDR="169.254.253.1/30"
|
|
ISLAND_HOST_IP="169.254.253.1"
|
|
ISLAND_GUEST_CIDR="169.254.253.2/30"
|
|
ISLAND_PORT="8443"
|
|
ISLAND_ON=0
|
|
WANT_ISLAND=true # appliance default; --no-island sets false
|
|
LAN_IP="" # vmbr0 IPv4 (computed in step_agent_config; the LAN anchor for lan_resolver)
|
|
ROOTFS_GROW=""
|
|
DATAVOL_GROW=""
|
|
SYSDATA_GROW=""
|
|
CPU_CORES="" # --cores: optional appliance CPU-core cap (empty/unset = golden default)
|
|
MEM_MIB="" # --memory: optional appliance RAM cap in MiB (empty/unset = golden default)
|
|
PASSPHRASE_FILE=""
|
|
OPERATOR_PUBKEY_FILE="" # --operator-pubkey-file: "operational|recovery <authorized_keys line>" per line; overrides the OPERATOR_KEY_* constants (GL-4)
|
|
PRESERVE_FROM=""
|
|
PRESERVE_STATE_FROM="" # dir holding a prior local-api.{crt,key} + local-tokens.log to carry over (keeps the pin stable across a reinstall)
|
|
ALLOW_NEW_LEAF=false # opt-in to intentionally regenerate the agent leaf on a populated host (else the guard refuses)
|
|
FORCE=false
|
|
FORCE_GITEA_GOLDEN=false
|
|
SKIP_PROVISION=false
|
|
DRY_RUN=false
|
|
RESUME=false
|
|
UNINSTALL=false # --uninstall: local host teardown (destroy guest + remove agent/pveum/state)
|
|
REMOVE_GOLDEN=false # --remove-golden: also delete the golden vzdump during --uninstall
|
|
ADOPT_POOL=false # --adopt-pool: retrofit an EXISTING Felhom guest into the felhom pool (non-destructive)
|
|
RESCOPE_ACL=false # --rescope-acl: migrate an existing install from the broad-/ token to the scoped ACL
|
|
ROTATE_RECOVERY=false # --rotate-recovery: regenerate + re-vault the break-glass root@pam password (TASK G1)
|
|
ENABLE_OOB=false # the OOB belt (TASK H1). DEFAULT-ON for appliance mode since v1.25.0 (F9 ruling
|
|
# 2026-07-23: an operator-owned box must be reachable OOB) — see the appliance
|
|
# default below; --enable-oob forces it, --no-oob opts an appliance box out.
|
|
NO_OOB=false # --no-oob: opt an APPLIANCE install out of the default belt (byo is off already)
|
|
|
|
# --- Gitea (artifact source) + agent install model (BUNDLE slice) ---
|
|
GITEA_BASE="https://gitea.dooplex.hu"
|
|
GITEA_OWNER="admin"
|
|
AGENT_REPO="felhom-agent" # for the raw unit/sudoers fetch (config text, canonical source)
|
|
AGENT_USER="felhom-agent" # the non-root service user the unit + sudoers name
|
|
AGENT_BIN="/usr/local/bin/felhom-agent"
|
|
AGENT_SUDOERS="/etc/sudoers.d/felhom-agent"
|
|
AGENT_UNIT="/etc/systemd/system/felhom-agent.service"
|
|
AGENT_STATE_DIR="/var/lib/felhom-agent"
|
|
|
|
PVE_USER="felhom-agent@pve"
|
|
PVE_TOKENID="agent"
|
|
PVE_POOL="felhom" # dedicated pool every Felhom-managed guest joins (fleet uniformity + ACL scope)
|
|
PVE_ROLE="FelhomAgent" # the PRE-3b single broad role (removed on rescope/uninstall if present — legacy)
|
|
|
|
# Pool-scoped ACL (3b, validated by SPIKE-pool-scoped-acl-2026-07-01): the agent's privileges are split
|
|
# across THREE roles applied at scoped paths so the token can only touch Felhom's own guests + storages
|
|
# (blast-radius containment). `pveum acl` grants a whole role per path, hence 3 roles not 1. Each role is
|
|
# granted to BOTH the user AND the token (privsep intersection). Guest privs (incl. Pool.Allocate so the
|
|
# agent restores INTO the pool) live at /pool/felhom; Datastore WRITE privs at each agent-touched storage.
|
|
# `Datastore.Audit` is box-wide in Base (3b-fix v1.7.0): the agent must ENUMERATE every storage incl. the
|
|
# dynamically-enrolled removable drives (felhom-usb/felhom-flash) it observes but never registers — a
|
|
# per-storage Audit grant hid them → false "drive detached" alerts. Audit is read-only, so box-wide Audit
|
|
# keeps WRITE containment (Allocate/AllocateSpace stay per-storage). Only Sys.Audit/SDN.Use/Datastore.Audit box-wide.
|
|
# `Pool.Audit` (v1.9.0, audit A1): the agent's stale-lock reaper reads GET /pools/felhom as its
|
|
# ownership registry (agent v0.62.0+); without it the reaper fail-safes (skips) and reports the
|
|
# `pve:pool-read` capability degraded. NOTE: Pool.Allocate does NOT satisfy the read — the spike
|
|
# (SPIKE-a1-pool-membership-read-2026-07-03 T2) 403'd with Allocate granted; Audit is required.
|
|
PVE_ROLE_GUEST="FelhomAgentGuest"
|
|
PVE_ROLE_STORE="FelhomAgentStore"
|
|
PVE_ROLE_BASE="FelhomAgentBase"
|
|
PVE_PRIVS_GUEST="VM.Allocate VM.Audit VM.Config.Disk VM.Config.CPU VM.Config.Memory VM.Config.Network VM.Config.Options VM.PowerMgmt VM.Snapshot VM.Snapshot.Rollback VM.Backup Pool.Allocate Pool.Audit"
|
|
PVE_PRIVS_STORE="Datastore.Allocate Datastore.AllocateSpace"
|
|
PVE_PRIVS_BASE="Sys.Audit SDN.Use Datastore.Audit"
|
|
# Storages the agent reads/writes (archive+dump=local, restore=local-lvm, offsite DR=felhom-pbs). The
|
|
# offsite felhom-pbs MUST be included or the agent's DR backup 403s (SPIKE residual #1). --acl-storages overrides.
|
|
PVE_STORAGES=(local local-lvm felhom-pbs)
|
|
# E-2: the whole-guest backup target storage id, and what configure_backup_target resolved to.
|
|
# BACKUP_TARGET_RESOLVED feeds agent.json's local_backup_target; "local" means DEGRADED (Case B).
|
|
BACKUP_TARGET_ID="felhom-backup"
|
|
BACKUP_TARGET_RESOLVED="local"
|
|
# F-LEAK (Campaign 8, v1.21.0): the restore-test's scratch VMID band, granted the GUEST role at each
|
|
# /vms/<id> PATH. WHY THIS EXISTS: the Guest role is granted at /pool/felhom, and a guest only joins
|
|
# that pool when its restore COMPLETES. A restore-test whose restore FAILS therefore leaves a scratch
|
|
# guest that exists but is in no pool — and the agent's own teardown then 403s on VM.Allocate, so it
|
|
# leaks a half-restored guest holding its disks until a human removes it.
|
|
# THE ALTERNATIVE THAT DOES NOT WORK: adopting the stranded guest into the pool first. Tried, shipped,
|
|
# and REFUTED live on 2026-07-28 — `PUT /pools/felhom` with vms=<id> ALSO requires VM.Allocate on
|
|
# /vms/<id>, so Pool.Allocate cannot bootstrap its own membership.
|
|
# WHY THIS IS NOT A WIDENING: PVE ACLs are path-scoped, so ten explicit /vms/<id> grants authorise the
|
|
# agent on exactly these ten scratch IDs and nowhere else. It is NOT granted at /vms — that would
|
|
# authorise destroying every guest on the box, including a co-tenant's. Verify with
|
|
# pvesh get /access/permissions --userid 'felhom-agent@pve!agent' --path /vms/<some-real-vmid>
|
|
# which must NOT list VM.Allocate. The agent keeps its own in-band check; this is the outer fence.
|
|
PVE_SCRATCH_VMID_MIN=990000
|
|
PVE_SCRATCH_VMID_MAX=990009
|
|
|
|
# FELHOM_INSTALL_STATE_DIR: test-harness-only override (hostinstall-mode-harness.sh) so dry/refusal
|
|
# cases can never touch a live install's state.json. Production runs never set it.
|
|
STATE_DIR="${FELHOM_INSTALL_STATE_DIR:-/var/lib/felhom-install}"
|
|
STATE_FILE="${STATE_DIR}/state.json"
|
|
AGENT_CONFIG="" # resolved in preflight
|
|
HARD_MIN_LVM_GIB=120 # a useful appliance won't fit below this on local-lvm
|
|
|
|
# Runtime carriers (never logged)
|
|
PASSPHRASE=""
|
|
PVE_TOKEN="" # felhom-agent@pve!agent=<secret>
|
|
HOST_ID=""
|
|
HOST_API_KEY=""
|
|
GIT_USER="" # from controller.yaml (config-retrieve) — Gitea fetch credential
|
|
GIT_TOKEN="" # from controller.yaml — NEVER logged
|
|
ART_AGENT_VER="" # hub artifact manifest: agent version + sha256
|
|
ART_AGENT_SHA=""
|
|
ART_GOLDEN_VER="" # hub artifact manifest: golden version + sha256
|
|
ART_GOLDEN_SHA=""
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# Helpers
|
|
#-------------------------------------------------------------------------------
|
|
# Print the header through the last option line. Markers, not line numbers (v1.9.1's numeric
|
|
# '2,95p' had drifted): start at the "# Day-0 host-bootstrap" line, end before "# State (".
|
|
# The version line is printed from SCRIPT_VERSION — the ONLY version source (F-1).
|
|
usage() {
|
|
echo "felhom-host-install.sh v${SCRIPT_VERSION}"
|
|
sed -n '/^# Day-0 host-bootstrap/,/^# State (/p' "$0" | sed '$d' | sed 's/^# \{0,1\}//'
|
|
exit 0
|
|
}
|
|
|
|
run() { # simple (no pipes/redirects) mutating command
|
|
if $DRY_RUN; then log_dry "$*"; else "$@"; fi
|
|
}
|
|
|
|
# used_vmids — every in-use guest id on this host. LXC (pct) and VMs (qm) SHARE the id space,
|
|
# so both are consulted; headers (non-numeric first column) are filtered out.
|
|
used_vmids() {
|
|
{ pct list 2>/dev/null; qm list 2>/dev/null; } | awk '{print $1}' | grep -E '^[0-9]+$'
|
|
}
|
|
|
|
# _vmid_in_use ID — true if ID is present in the pct+qm used-set (more complete than `pct status`,
|
|
# which only knows LXC).
|
|
_vmid_in_use() {
|
|
local target="$1" used
|
|
used=" $(used_vmids | tr '\n' ' ') "
|
|
[[ "$used" == *" $target "* ]]
|
|
}
|
|
|
|
# next_free_vmid BASE — the first id >= BASE not in the used-set, scanning upward.
|
|
next_free_vmid() {
|
|
local base="$1" used id
|
|
used=" $(used_vmids | tr '\n' ' ') "
|
|
id="$base"
|
|
while [[ "$used" == *" $id "* ]]; do id=$((id + 1)); done
|
|
echo "$id"
|
|
}
|
|
|
|
# State helpers (robust JSON via python3).
|
|
_state_has() {
|
|
[[ -f "$STATE_FILE" ]] || return 1
|
|
STATE_FILE="$STATE_FILE" python3 -c "import json,os,sys;f=os.environ['STATE_FILE'];d=json.load(open(f));sys.exit(0 if sys.argv[1] in d.get('completed',[]) else 1)" "$1" 2>/dev/null
|
|
}
|
|
_state_mark() {
|
|
$DRY_RUN && return 0
|
|
mkdir -p "$STATE_DIR"
|
|
STATE_FILE="$STATE_FILE" python3 -c "import json,os,sys;f=os.environ['STATE_FILE'];d=json.load(open(f)) if os.path.exists(f) else {'completed':[]};c=d.setdefault('completed',[]);(c.append(sys.argv[1]) if sys.argv[1] not in c else None);json.dump(d,open(f,'w'),indent=2)" "$1"
|
|
}
|
|
should_skip() { # returns 0 (skip) if --resume AND step already done
|
|
if $RESUME && _state_has "$1"; then log_skip "step '$1' already completed"; return 0; fi
|
|
return 1
|
|
}
|
|
# _state_put KEY VALUE — set a top-level string key in state.json (creates the file if absent).
|
|
# Mirrors _state_mark: dry-run no-ops (writes nothing), robust JSON via python3.
|
|
_state_put() {
|
|
$DRY_RUN && return 0
|
|
mkdir -p "$STATE_DIR"
|
|
STATE_FILE="$STATE_FILE" python3 -c "import json,os,sys;f=os.environ['STATE_FILE'];d=json.load(open(f)) if os.path.exists(f) else {'completed':[]};d[sys.argv[1]]=sys.argv[2];json.dump(d,open(f,'w'),indent=2)" "$1" "$2"
|
|
}
|
|
# _state_get KEY — print the top-level string value for KEY (empty if the file/key is absent).
|
|
_state_get() {
|
|
[[ -f "$STATE_FILE" ]] || return 0
|
|
STATE_FILE="$STATE_FILE" python3 -c "import json,os,sys;d=json.load(open(os.environ['STATE_FILE']));print(d.get(sys.argv[1],''))" "$1" 2>/dev/null
|
|
}
|
|
|
|
http_code() { # GET, prints status code only (read-only preflight)
|
|
curl -fsS -o /dev/null -w '%{http_code}' "$@" 2>/dev/null || curl -sS -o /dev/null -w '%{http_code}' "$@" 2>/dev/null
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# Artifact + Gitea helpers (BUNDLE slice)
|
|
#-------------------------------------------------------------------------------
|
|
# Resolve the hub-vouched artifact manifest (agent + golden version+sha256). Passphrase-authed,
|
|
# same trust root as config-retrieve. Sets ART_* globals. Empty fields are valid (caller falls back).
|
|
resolve_artifacts() {
|
|
local resp code body
|
|
resp=$(curl -sS -w $'\n%{http_code}' "$HUB_URL/api/v1/artifacts/$CUSTOMER_ID" \
|
|
-H "X-Retrieval-Password: $PASSPHRASE")
|
|
code=$(tail -n1 <<<"$resp"); body=$(sed '$d' <<<"$resp")
|
|
[[ "$code" == "200" ]] || die "artifact manifest fetch failed: HTTP $code"
|
|
ART_AGENT_VER=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['agent']['version'])" "$body" 2>/dev/null || echo "")
|
|
ART_AGENT_SHA=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['agent']['sha256'])" "$body" 2>/dev/null || echo "")
|
|
ART_GOLDEN_VER=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['golden']['version'])" "$body" 2>/dev/null || echo "")
|
|
ART_GOLDEN_SHA=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['golden']['sha256'])" "$body" 2>/dev/null || echo "")
|
|
}
|
|
|
|
# Resolve the Gitea fetch credential (git username + token) from the customer's controller.yaml —
|
|
# the SAME secret config-retrieve already hands out (NO new credential). Sets GIT_USER / GIT_TOKEN.
|
|
# Parses the git: block without a YAML lib (fresh PVE has no PyYAML).
|
|
# v1.11.2 (GL-6 Gate-0 operator ruling, G3): EMPTY credentials are ALLOWED — the fetches then run
|
|
# ANONYMOUSLY (the Felhom artifacts are deliberately world-readable; integrity comes from the
|
|
# hub-vouched sha256 + visudo/sh -n validation, never from transport auth). Loud warn, never die:
|
|
# a customer config without a git token must not block day-0 while the anonymous path works.
|
|
resolve_git_creds() {
|
|
local yaml
|
|
yaml=$(curl -fsS "$HUB_URL/api/v1/config/$CUSTOMER_ID" -H "X-Retrieval-Password: $PASSPHRASE") \
|
|
|| die "controller.yaml fetch failed (for the git fetch credential)"
|
|
GIT_USER=$(awk '/^[^[:space:]#]/{ingit=($1=="git:")} ingit&&$1=="username:"{print $2}' <<<"$yaml" | head -1)
|
|
GIT_TOKEN=$(awk '/^[^[:space:]#]/{ingit=($1=="git:")} ingit&&$1=="token:"{print $2}' <<<"$yaml" | head -1)
|
|
# strip any surrounding quotes
|
|
GIT_USER="${GIT_USER%\"}"; GIT_USER="${GIT_USER#\"}"
|
|
GIT_TOKEN="${GIT_TOKEN%\"}"; GIT_TOKEN="${GIT_TOKEN#\"}"
|
|
if [[ -z "$GIT_TOKEN" ]]; then
|
|
log_warn " no git credential in controller.yaml — fetching artifacts ANONYMOUSLY (they are world-readable; sha256 verification unchanged)"
|
|
fi
|
|
}
|
|
|
|
# _git_auth_args populates the caller's curl arg array with basic-auth ONLY when a credential is
|
|
# present — empty creds = anonymous fetch (v1.11.2). `curl -u user:` with an empty token would NOT
|
|
# fall back to anonymous (a wrong credential 401s even on world-readable content), hence the array.
|
|
_git_auth_args() {
|
|
local -n _out=$1
|
|
_out=()
|
|
if [[ -n "$GIT_TOKEN" ]]; then
|
|
_out=(-u "${GIT_USER}:${GIT_TOKEN}")
|
|
fi
|
|
}
|
|
|
|
# _dry_auth_hint — the dry-run print's auth stanza (drill F-2): the old hardcoded "curl -u <git>"
|
|
# lied on the anonymous-fetch branch. Requires resolve_git_creds to have run (both call sites do).
|
|
_dry_auth_hint() {
|
|
if [[ -n "$GIT_TOKEN" ]]; then echo "curl -u <git>"; else echo "curl (anonymous — no git credential)"; fi
|
|
}
|
|
|
|
# Fetch a Gitea generic-package URL to a dest (authed when a credential exists, else anonymous),
|
|
# then VERIFY its sha256 against the expected (hub-vouched) value. Aborts on any mismatch —
|
|
# verify-before-use; the sha is the integrity root either way. $1=url $2=dest $3=expected_sha
|
|
fetch_verify() {
|
|
local url="$1" dest="$2" want="$3"
|
|
[[ -n "$want" ]] || die "refusing to install an artifact with no expected sha256 (manifest incomplete): $url"
|
|
local -a _auth; _git_auth_args _auth
|
|
curl -fsS "${_auth[@]}" -o "$dest" "$url" || die "fetch failed: $url"
|
|
local got; got=$(sha256sum "$dest" | awk '{print $1}')
|
|
if [[ "$got" != "$want" ]]; then
|
|
rm -f "$dest"
|
|
die "sha256 MISMATCH for $url — expected $want got $got. Refusing to install (verify-before-use)."
|
|
fi
|
|
log_success " verified sha256 ${got:0:16}… matches the hub manifest"
|
|
}
|
|
|
|
# Fetch a raw config file (the canonical unit/sudoers) from the agent repo (authed when a credential
|
|
# exists, else anonymous). These are non-executable text (not the integrity-checked binary); the
|
|
# sudoers is `visudo -cf`-validated before install, which catches corruption/tampering that would
|
|
# matter. $1=repo-path $2=dest
|
|
#
|
|
# R-110 / R-183: PINNED TO THE AGENT VERSION BEING INSTALLED, never to a branch.
|
|
#
|
|
# These sixteen files are the AGENT's configs — its systemd unit, its sudoers, its guarded wrappers —
|
|
# so the ref that is correct for them is the agent version this run is installing, which the hub has
|
|
# vouched and whose binary sha this script verifies. It is NOT the installer's own SCRIPT_VERSION:
|
|
# these files do not live in the installer's repo and have no relationship to its version line.
|
|
#
|
|
# Before this they came from `raw/branch/main`, which is a REAL SKEW and not only a publish-channel
|
|
# defect (R-183): a fresh install fetched the vouched agent BINARY while taking its unit file and
|
|
# sudoers from whatever `main` happened to hold — two refs, one install, and nothing compared them.
|
|
#
|
|
# NO FALLBACK TO A BRANCH. A vouched version whose tag is missing must fail loudly here rather than
|
|
# quietly serving `main`, because a silent fallback is exactly the "appearance of control with none of
|
|
# it" this change exists to remove. `agent_gates.py`'s published-version gate keeps the tag and the
|
|
# vouched version in step, so this die is a backstop and not the primary control.
|
|
fetch_raw() {
|
|
local path="$1" dest="$2"
|
|
# Late steps (mgmt-watchdog, OOB) can run without step 5 having resolved the manifest.
|
|
[[ -n "$ART_AGENT_VER" ]] || resolve_artifacts
|
|
[[ -n "$ART_AGENT_VER" ]] || die "cannot pin $path: no agent version resolved from the hub manifest"
|
|
local -a _auth; _git_auth_args _auth
|
|
curl -fsS "${_auth[@]}" -o "$dest" \
|
|
"$GITEA_BASE/$GITEA_OWNER/$AGENT_REPO/raw/tag/v$ART_AGENT_VER/$path" \
|
|
|| die "raw fetch failed: $path (agent tag v$ART_AGENT_VER — is that version tagged in $AGENT_REPO?)"
|
|
[[ -s "$dest" ]] || die "raw fetch empty: $path"
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# Uninstall (local host teardown) — reverse of install; no hub contact, no passphrase
|
|
#-------------------------------------------------------------------------------
|
|
# felhom_guests — every vmid on this host that carries the /etc/felhom-bootstrap bind mount (the
|
|
# read-only bootstrap mount an agent-provisioned guest always has). Matched by the CONSTANT guest
|
|
# PATH, not a hardcoded mpN slot (the slot drifts; on the demo host it's mp9).
|
|
felhom_guests() {
|
|
local id
|
|
for id in $(used_vmids); do
|
|
pct config "$id" 2>/dev/null | grep -q 'mp=/etc/felhom-bootstrap' && echo "$id"
|
|
done
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# felhom pool (fleet uniformity) — every managed guest joins the `felhom` pool. All pool ops run as
|
|
# root@pam from the installer, so NO agent/token/ACL change is involved (that is the separate 3b spike).
|
|
# API shapes confirmed on PVE 9: `pvesh get /pools` → [{poolid,comment}]; `pvesh get /pools/<id>` →
|
|
# {poolid,comment,members:[{vmid,...}]}. Pool ops: `pveum pool add|delete <id>`, `pveum pool modify
|
|
# <id> --vms <ids>` (additive).
|
|
#-------------------------------------------------------------------------------
|
|
# pool_exists — true if the felhom pool is present.
|
|
pool_exists() {
|
|
pvesh get /pools --output-format json 2>/dev/null \
|
|
| python3 -c "import json,sys;sys.exit(0 if any(p.get('poolid')=='$PVE_POOL' for p in json.load(sys.stdin)) else 1)" 2>/dev/null
|
|
}
|
|
|
|
# pool_members — space-separated vmids currently in the felhom pool (empty if none / pool absent).
|
|
pool_members() {
|
|
pvesh get "/pools/$PVE_POOL" --output-format json 2>/dev/null \
|
|
| python3 -c "import json,sys
|
|
try: d=json.load(sys.stdin)
|
|
except Exception: sys.exit(0)
|
|
print(' '.join(str(m.get('vmid')) for m in d.get('members',[]) if m.get('vmid') is not None))" 2>/dev/null
|
|
}
|
|
|
|
# ensure_felhom_pool — create the pool if absent (idempotent no-op otherwise). Via run() (dry-run-aware).
|
|
ensure_felhom_pool() {
|
|
if pool_exists; then
|
|
log_skip " pool $PVE_POOL already exists"
|
|
else
|
|
run pveum pool add "$PVE_POOL" --comment "Felhom-managed guests"
|
|
fi
|
|
}
|
|
|
|
# pool_add_guest VMID — add a guest to the felhom pool unless it is already a member (idempotent).
|
|
pool_add_guest() {
|
|
local vmid="$1" members
|
|
members=" $(pool_members) "
|
|
if [[ "$members" == *" $vmid "* ]]; then
|
|
log_skip " guest $vmid already in pool $PVE_POOL"
|
|
else
|
|
run pveum pool modify "$PVE_POOL" -vms "$vmid"
|
|
log_success " guest $vmid added to pool $PVE_POOL"
|
|
fi
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# Pool-scoped ACL helpers (3b). All ops run as root@pam (installer) — no privilege change to the agent.
|
|
#-------------------------------------------------------------------------------
|
|
# _role_exists NAME — true if a pveum role NAME exists.
|
|
_role_exists() {
|
|
pveum role list --output-format json 2>/dev/null \
|
|
| python3 -c "import json,sys;sys.exit(0 if any(r['roleid']==sys.argv[1] for r in json.load(sys.stdin)) else 1)" "$1" 2>/dev/null
|
|
}
|
|
|
|
# _ensure_role NAME "PRIVS" — create the role, or modify it to the exact priv set (idempotent).
|
|
_ensure_role() {
|
|
local name="$1" privs="$2"
|
|
if _role_exists "$name"; then
|
|
log_info " role $name exists — ensuring exact privileges"
|
|
run pveum role modify "$name" -privs "$privs"
|
|
else
|
|
run pveum role add "$name" -privs "$privs"
|
|
fi
|
|
}
|
|
|
|
# _grant PATH ROLE — grant ROLE at PATH to BOTH the user AND the token (privsep intersection). `acl
|
|
# modify` is idempotent so this is safe to repeat / re-apply after a token rotation.
|
|
_grant() {
|
|
local path="$1" role="$2"
|
|
run pveum acl modify "$path" -user "$PVE_USER" -role "$role"
|
|
run pveum acl modify "$path" -token "${PVE_USER}!${PVE_TOKENID}" -role "$role"
|
|
}
|
|
|
|
|
|
# ---- E-2 Part 2: the whole-guest backup target -------------------------------------------------
|
|
#
|
|
# CASE A an eligible secondary drive is already mounted at install time -> create `felhom-backup`
|
|
# on that drive's OWN mountpoint (F-1), is_mountpoint 1 (F-2), grant the agent (F-3), and
|
|
# point the primary tier at it.
|
|
# CASE B system drive only -> the target stays on the system drive. This is RECORDED AS DEGRADED,
|
|
# not as normal: it protects against corruption and NOT against drive loss. The install
|
|
# still succeeds — a single-drive appliance is a valid product, it just cannot survive
|
|
# drive loss.
|
|
#
|
|
# PHASE-0 REALITY CHECK (E-2, 2026-07-29): the installer has NO drive-enrollment step —
|
|
# /mnt/felhom-drives is agent-installed at runtime and drives are formatted/mounted by the dashboard
|
|
# later. So on a FRESH appliance Case A almost never fires; the common case is Case B, and the drive
|
|
# arrives months later through the wizard (E-2 Part 3). Case A exists for the reinstall/recovery
|
|
# shape, where an agent-generated .mount unit already brings the drive up by fs-UUID.
|
|
#
|
|
# eligible_backup_drive — echoes the mountpoint of a usable secondary drive, or nothing.
|
|
# Attributes SUGGEST and REFUSE THE ABSURD; they never decide by transport. The reference hardware
|
|
# forbids that: demo-felhom's backup drive is an external USB HDD, and BOTH demo boxes' drives
|
|
# report removable=0 — a transport rule would disqualify the reference drive, a removable rule
|
|
# would find no candidate at all.
|
|
eligible_backup_drive() {
|
|
local root_dev mp dev sz
|
|
root_dev=$(findmnt -no SOURCE / 2>/dev/null || true)
|
|
# Only real, already-mounted filesystems under /mnt — never / , never a bare directory.
|
|
while read -r mp dev; do
|
|
[[ -n "$mp" && -n "$dev" ]] || continue
|
|
[[ "$mp" == /mnt/* ]] || continue
|
|
[[ "$dev" != "$root_dev" ]] || continue
|
|
[[ "$dev" == /dev/* ]] || continue # skip NAS/tmpfs/overlay sources
|
|
# Refuse the absurd: a thumb-drive-sized volume is not a whole-guest backup target. A guest
|
|
# archive measured 1.6-6.0 GB on the demo fleet, and keep-last=3 multiplies it.
|
|
sz=$(findmnt -nbo SIZE "$mp" 2>/dev/null || echo 0)
|
|
[[ "$sz" -ge $((64 * 1024 * 1024 * 1024)) ]] || continue
|
|
echo "$mp"; return 0
|
|
done < <(findmnt -rn -o TARGET,SOURCE -t ext4,xfs,btrfs 2>/dev/null)
|
|
return 1
|
|
}
|
|
|
|
# configure_backup_target — Case A/B. Idempotent and SCENARIO-F SAFE.
|
|
configure_backup_target() {
|
|
# SCENARIO F, the hard rule: never touch a box that already has a target. The agent.json write
|
|
# below is already protected structurally (`setdefault` preserves an existing `backup` block
|
|
# WHOLE), and this is the second, explicit guard for the STORAGE half — `pvesm add` on an
|
|
# existing id is an error, and repointing a live target is exactly the silent retarget this
|
|
# whole arc closes.
|
|
if pvesm status --storage "$BACKUP_TARGET_ID" >/dev/null 2>&1; then
|
|
log_skip " backup target '$BACKUP_TARGET_ID' already exists — leaving it exactly as it is (Scenario F)"
|
|
BACKUP_TARGET_RESOLVED="$BACKUP_TARGET_ID"
|
|
return 0
|
|
fi
|
|
local mp
|
|
if mp=$(eligible_backup_drive); then
|
|
# CASE A. Run through the SAME guarded wrapper the agent uses, so the installer and the
|
|
# runtime path cannot drift and the laws are enforced in exactly one place.
|
|
if $DRY_RUN; then
|
|
log_dry "felhom-backup-target-apply create $BACKUP_TARGET_ID $mp ; grant"
|
|
else
|
|
/usr/local/sbin/felhom-backup-target-apply create "$BACKUP_TARGET_ID" "$mp" \
|
|
|| die "backup target creation failed on $mp"
|
|
/usr/local/sbin/felhom-backup-target-apply grant "$BACKUP_TARGET_ID" \
|
|
|| die "backup target grant failed — every backup would 403"
|
|
fi
|
|
BACKUP_TARGET_RESOLVED="$BACKUP_TARGET_ID"
|
|
log_success " backup target: $BACKUP_TARGET_ID on $mp (separate device — survives a system-drive failure)"
|
|
else
|
|
# CASE B — degraded, and SAID SO. Not a failure; a single-drive appliance is valid.
|
|
BACKUP_TARGET_RESOLVED="local"
|
|
log_warn " backup target: DEGRADED — no eligible second drive, so the whole-system backup stays on the SYSTEM drive."
|
|
log_warn " It protects against file corruption but NOT against a disk failure. Attach a second drive and assign it in the dashboard."
|
|
fi
|
|
}
|
|
|
|
# apply_scoped_acl — create the 3 scoped roles and grant each at its path(s). Requires the pool to exist.
|
|
# ORDER (3b-fix): Base (which holds box-wide Datastore.Audit) is ensured + granted BEFORE Store, so a
|
|
# RE-APPLY on a live box adds Audit@/ before Store drops its per-storage Audit → the agent never loses
|
|
# storage-enumeration visibility mid-apply (gap-free).
|
|
apply_scoped_acl() {
|
|
_ensure_role "$PVE_ROLE_BASE" "$PVE_PRIVS_BASE"
|
|
_ensure_role "$PVE_ROLE_GUEST" "$PVE_PRIVS_GUEST"
|
|
_ensure_role "$PVE_ROLE_STORE" "$PVE_PRIVS_STORE"
|
|
_grant / "$PVE_ROLE_BASE"
|
|
_grant "/pool/$PVE_POOL" "$PVE_ROLE_GUEST"
|
|
local s
|
|
for s in "${PVE_STORAGES[@]}"; do
|
|
_grant "/storage/$s" "$PVE_ROLE_STORE"
|
|
done
|
|
# F-LEAK: the scratch band, per-VMID. A failed restore-test's scratch never joins the pool, so the
|
|
# pool grant cannot reach it; these path-scoped grants can, and reach nothing else.
|
|
local v
|
|
for ((v = PVE_SCRATCH_VMID_MIN; v <= PVE_SCRATCH_VMID_MAX; v++)); do
|
|
_grant "/vms/$v" "$PVE_ROLE_GUEST"
|
|
done
|
|
log_success " scoped ACL applied (Base@/, Guest@/pool/$PVE_POOL + /vms/${PVE_SCRATCH_VMID_MIN}..${PVE_SCRATCH_VMID_MAX}, Store@[${PVE_STORAGES[*]}])"
|
|
}
|
|
|
|
# _acl_grant_present PATH TYPE UGID ROLE — true if that exact ACL grant exists.
|
|
_acl_grant_present() {
|
|
pveum acl list --output-format json 2>/dev/null | python3 -c "import json,sys
|
|
p,t,u,r=sys.argv[1:5]
|
|
sys.exit(0 if any(e.get('path')==p and e.get('type')==t and e.get('ugid')==u and e.get('roleid')==r for e in json.load(sys.stdin)) else 1)" "$1" "$2" "$3" "$4" 2>/dev/null
|
|
}
|
|
|
|
# remove_scoped_acl — delete the 3-role scoped grants (user+token at each path), then the 3 roles
|
|
# (roles last — PVE refuses to delete a referenced role). Presence-checked, tolerate-absent.
|
|
remove_scoped_acl() {
|
|
local s
|
|
if _acl_grant_present "/pool/$PVE_POOL" user "$PVE_USER" "$PVE_ROLE_GUEST"; then run pveum acl delete "/pool/$PVE_POOL" --users "$PVE_USER" --roles "$PVE_ROLE_GUEST"; fi
|
|
if _acl_grant_present "/pool/$PVE_POOL" token "${PVE_USER}!${PVE_TOKENID}" "$PVE_ROLE_GUEST"; then run pveum acl delete "/pool/$PVE_POOL" --tokens "${PVE_USER}!${PVE_TOKENID}" --roles "$PVE_ROLE_GUEST"; fi
|
|
# F-LEAK scratch-band grants. These MUST be removed before the role delete below — PVE refuses to
|
|
# delete a role that is still referenced by any ACL, so leaving them behind breaks the uninstall.
|
|
local v
|
|
for ((v = PVE_SCRATCH_VMID_MIN; v <= PVE_SCRATCH_VMID_MAX; v++)); do
|
|
if _acl_grant_present "/vms/$v" user "$PVE_USER" "$PVE_ROLE_GUEST"; then run pveum acl delete "/vms/$v" --users "$PVE_USER" --roles "$PVE_ROLE_GUEST"; fi
|
|
if _acl_grant_present "/vms/$v" token "${PVE_USER}!${PVE_TOKENID}" "$PVE_ROLE_GUEST"; then run pveum acl delete "/vms/$v" --tokens "${PVE_USER}!${PVE_TOKENID}" --roles "$PVE_ROLE_GUEST"; fi
|
|
done
|
|
for s in "${PVE_STORAGES[@]}"; do
|
|
if _acl_grant_present "/storage/$s" user "$PVE_USER" "$PVE_ROLE_STORE"; then run pveum acl delete "/storage/$s" --users "$PVE_USER" --roles "$PVE_ROLE_STORE"; fi
|
|
if _acl_grant_present "/storage/$s" token "${PVE_USER}!${PVE_TOKENID}" "$PVE_ROLE_STORE"; then run pveum acl delete "/storage/$s" --tokens "${PVE_USER}!${PVE_TOKENID}" --roles "$PVE_ROLE_STORE"; fi
|
|
done
|
|
if _acl_grant_present / user "$PVE_USER" "$PVE_ROLE_BASE"; then run pveum acl delete / --users "$PVE_USER" --roles "$PVE_ROLE_BASE"; fi
|
|
if _acl_grant_present / token "${PVE_USER}!${PVE_TOKENID}" "$PVE_ROLE_BASE"; then run pveum acl delete / --tokens "${PVE_USER}!${PVE_TOKENID}" --roles "$PVE_ROLE_BASE"; fi
|
|
local name
|
|
for name in "$PVE_ROLE_GUEST" "$PVE_ROLE_STORE" "$PVE_ROLE_BASE"; do
|
|
if _role_exists "$name"; then run pveum role delete "$name"; else log_skip " role $name already absent"; fi
|
|
done
|
|
}
|
|
|
|
# remove_old_broad_acl — remove the PRE-3b single FelhomAgent role granted at / (user+token) + the role.
|
|
# Tolerate-absent (fresh 3b installs have none). Used by --rescope-acl (migration) and step_token (so a
|
|
# re-install can't leave the old broad grant unioned with the new scoped one).
|
|
remove_old_broad_acl() {
|
|
if _acl_grant_present / user "$PVE_USER" "$PVE_ROLE"; then run pveum acl delete / --users "$PVE_USER" --roles "$PVE_ROLE"; fi
|
|
if _acl_grant_present / token "${PVE_USER}!${PVE_TOKENID}" "$PVE_ROLE"; then run pveum acl delete / --tokens "${PVE_USER}!${PVE_TOKENID}" --roles "$PVE_ROLE"; fi
|
|
if _role_exists "$PVE_ROLE"; then run pveum role delete "$PVE_ROLE"; else log_skip " old broad role $PVE_ROLE already absent"; fi
|
|
}
|
|
|
|
# _guest_drive_note VMID — GL-4 (guest-only mode): best-effort list of the drives the agent's bind
|
|
# store records for THIS vmid (/var/lib/felhom-agent/guest-binds.json, vmid -> durable-ids); store
|
|
# unreadable/absent → the generic note. Read-only, never dies.
|
|
_guest_drive_note() {
|
|
local vmid="$1" binds=""
|
|
binds=$(python3 -c "import json
|
|
try:
|
|
d=json.load(open('$AGENT_STATE_DIR/guest-binds.json'))
|
|
print(', '.join(d.get('$vmid',[])))
|
|
except Exception:
|
|
pass" 2>/dev/null || true)
|
|
if [[ -n "$binds" ]]; then
|
|
log_info " drives recorded as bound to guest $vmid: $binds"
|
|
log_info " they stay mounted (other guests may share the host) — eject via the dashboard BEFORE uninstalling if this guest owned them."
|
|
else
|
|
log_info " enrolled drives (if any) stay mounted; eject a drive via the dashboard BEFORE uninstall if it belonged to this guest."
|
|
fi
|
|
}
|
|
|
|
# _uninstall_statement full|guest-only — GL-4: the explicit end-of-teardown KEPT-vs-WIPED statement.
|
|
# WIPED mirrors what THIS mode actually ran; KEPT names everything that deliberately lives on.
|
|
# Statement-only: this script NEVER deletes PBS backups, hub records, escrow blobs, or drive data.
|
|
# Reads run_uninstall's locals (vmid, pool_removed, _busy_mounts, _had_break_glass) via bash's
|
|
# dynamic scoping — call it from run_uninstall only.
|
|
_uninstall_statement() {
|
|
local scope="$1"
|
|
echo ""
|
|
log_step "kept vs wiped — read before pulling drives or closing the customer out"
|
|
echo " WIPED (this run):"
|
|
echo " - guest $vmid (container + its OS/Docker/user-data volumes)"
|
|
if [[ "$scope" == "full" ]]; then
|
|
echo " - the felhom-agent: binary, unit, sudoers, config (+ its .bak backups), state dir, service user"
|
|
echo " - self-update artifacts: guarded wrapper, A/B slots (.prev/.new.*), rollback unit, start-limit drop-in"
|
|
echo " - break-glass watchdog + OOB artifacts (where present); guest-hook snippet; dnsmasq snippets; the mkfs + pbs-apply wrappers"
|
|
echo " - pveum: the Felhom roles/user/token/scoped ACL$( $pool_removed && printf '; the emptied %s pool' "$PVE_POOL")"
|
|
echo " - the install state file"
|
|
if $REMOVE_GOLDEN; then echo " - the golden vzdump (--remove-golden)"; fi
|
|
else
|
|
echo " - NOTHING host-level (other Felhom guests remain: agent, token/ACL, pool, state all stay)"
|
|
fi
|
|
echo " KEPT (lives on deliberately — remove/rotate these out-of-band if the customer is leaving):"
|
|
if [[ "$scope" == "full" ]]; then
|
|
echo " - the enrolled drives + ALL data under /mnt/felhom-drives — unmounted only, NEVER wiped;"
|
|
if [[ ${#_busy_mounts[@]} -gt 0 ]]; then
|
|
echo " physically removable now, EXCEPT still mounted (busy — stop the apps and retry): ${_busy_mounts[*]}"
|
|
else
|
|
echo " the drives are physically removable now."
|
|
fi
|
|
else
|
|
echo " - the enrolled drives + ALL data under /mnt/felhom-drives — left MOUNTED (remaining guests may use them)"
|
|
fi
|
|
if pvesm status 2>/dev/null | awk '$2=="pbs"{found=1} END{exit !found}'; then
|
|
echo " - the PBS backups + this customer's namespace on the PBS side — delete there if wanted"
|
|
fi
|
|
echo " - the hub host/customer record + report history (operator UI / DB)"
|
|
echo " - the escrow blob in the hub, if one exists (operator UI)"
|
|
if $_had_break_glass; then
|
|
echo " - the hub-vaulted root@pam recovery credential — the box KEEPS the password step 4b set; rotate it if the box leaves Felhom management"
|
|
fi
|
|
echo ""
|
|
}
|
|
|
|
# run_uninstall — the full guarded teardown. Every mutation goes through run() so --dry-run prints it
|
|
# and executes nothing. Ordering is the reverse of install: guest -> agent -> pveum(ACL,token,user,
|
|
# role) -> golden(opt-in) -> state file. See the TASK spec §7/§8.
|
|
run_uninstall() {
|
|
log_step "UNINSTALL — local host teardown"
|
|
|
|
# GL-4: snapshot state facts BEFORE any removal (the closing statement needs them; the state
|
|
# file itself is deleted in step 7).
|
|
local _had_break_glass=false _busy_mounts=()
|
|
_state_has break_glass && _had_break_glass=true
|
|
|
|
# 1. Resolve the target vmid: --vmid, else the recorded provisioned_vmid, else die.
|
|
local state_vmid vmid pool_removed=false
|
|
state_vmid=$(_state_get provisioned_vmid)
|
|
if $VMID_EXPLICIT; then
|
|
vmid="$VMID"
|
|
elif [[ -n "$state_vmid" ]]; then
|
|
vmid="$state_vmid"
|
|
log_info " no --vmid given; using recorded provisioned_vmid=$vmid from $STATE_FILE"
|
|
else
|
|
die "pass --vmid N (state has no recorded vmid)"
|
|
fi
|
|
|
|
# state-mismatch: an explicit --vmid that disagrees with the recorded one needs --force.
|
|
if $VMID_EXPLICIT && [[ -n "$state_vmid" && "$state_vmid" != "$vmid" ]]; then
|
|
if $FORCE; then
|
|
log_warn " --vmid $vmid differs from the recorded provisioned_vmid=$state_vmid — --force given, proceeding"
|
|
else
|
|
die "--vmid $vmid differs from the recorded provisioned_vmid=$state_vmid. Pass --force to override."
|
|
fi
|
|
fi
|
|
|
|
# 2. Guest teardown (guarded: ours-check + typed confirm).
|
|
if _vmid_in_use "$vmid"; then
|
|
# ours-check: a Felhom guest carries the /etc/felhom-bootstrap bind mount (constant guest path).
|
|
if pct config "$vmid" 2>/dev/null | grep -q 'mp=/etc/felhom-bootstrap'; then
|
|
log_info " vmid $vmid looks like a Felhom guest (has the /etc/felhom-bootstrap mount)"
|
|
elif $FORCE; then
|
|
log_warn " vmid $vmid has NO /etc/felhom-bootstrap mount — --force given, destroying anyway"
|
|
else
|
|
die "vmid $vmid does not look like a Felhom-provisioned guest (no /etc/felhom-bootstrap mount).
|
|
Refusing to destroy. Pass --force to override."
|
|
fi
|
|
|
|
# show the config so the operator can eyeball what is about to be destroyed
|
|
log_info " pct config $vmid:"
|
|
pct config "$vmid" 2>/dev/null | sed 's/^/ /'
|
|
|
|
# typed confirmation — mandatory, never skipped except in --dry-run (nothing is destroyed there).
|
|
if $DRY_RUN; then
|
|
log_dry "would prompt: Type the vmid ($vmid) to confirm PERMANENT destruction"
|
|
else
|
|
local ans
|
|
read -rp "Type the vmid ($vmid) to confirm PERMANENT destruction: " ans < /dev/tty
|
|
[[ "$ans" == "$vmid" ]] || die "confirmation mismatch (got '$ans', expected '$vmid') — aborting, nothing destroyed"
|
|
fi
|
|
|
|
# stop (tolerate already-stopped) then destroy
|
|
local gstat; gstat=$(pct status "$vmid" 2>/dev/null | awk '{print $2}')
|
|
if [[ "$gstat" == "running" ]]; then
|
|
run pct stop "$vmid"
|
|
else
|
|
log_skip " guest $vmid not running (status: ${gstat:-unknown}) — skip stop"
|
|
fi
|
|
run pct destroy "$vmid"
|
|
log_success " guest $vmid destroyed"
|
|
else
|
|
log_skip " guest $vmid already absent — skipping guest teardown"
|
|
# host-level removal is still ours-gated: allowed if state's provisioned_vmid matches; else --force.
|
|
if [[ -n "$state_vmid" && "$state_vmid" == "$vmid" ]]; then
|
|
log_info " recorded provisioned_vmid matches $vmid — host-level removal permitted"
|
|
elif ! $FORCE; then
|
|
die "guest $vmid is absent and is not the recorded provisioned_vmid ('${state_vmid:-none}') — refusing host-level removal without --force."
|
|
fi
|
|
fi
|
|
|
|
# 3. Other-Felhom-guests detector — the safe default. If any OTHER Felhom guest remains and no
|
|
# --force, stop after the guest teardown and leave every host-level component in place.
|
|
local others others_csv
|
|
others=$(felhom_guests | grep -vx "$vmid" || true)
|
|
if [[ -n "$others" ]] && ! $FORCE; then
|
|
others_csv=$(echo "$others" | tr '\n' ' ' | sed 's/ */ /g;s/^ //;s/ $//;s/ /, /g')
|
|
echo ""
|
|
log_warn "Other Felhom guests remain (${others_csv}); leaving the agent + PVE token + state in place."
|
|
log_warn "Re-run --uninstall --force to remove host-level components anyway (this orphans ${others_csv})."
|
|
# GL-4 (Scenario B): NO umounts in guest-only mode — drives may serve the remaining guests.
|
|
_guest_drive_note "$vmid"
|
|
_uninstall_statement guest-only
|
|
log_success "UNINSTALL (guest-only) complete — removed guest $vmid; host-level components preserved."
|
|
log_info " NOTE: the host record still exists in the hub — remove it there if desired."
|
|
$DRY_RUN && log_warn " DRY-RUN: nothing above was actually executed."
|
|
return 0
|
|
fi
|
|
|
|
# ── host-level removal (reverse of install) ──────────────────────────────────────────────────
|
|
log_step "host-level removal"
|
|
|
|
# 4. Agent removal — service, unit(+.bak), sudoers, binary(+.bak), state dir, config, user.
|
|
# NEVER `sudo`. Resolve the agent config path BEFORE the unit is removed (mirrors preflight —
|
|
# the unit's -config arg is the truth, else the default); the config holds the per-host hub
|
|
# api_key and must not survive an uninstall (drill finding R1).
|
|
local agent_cfg=""
|
|
if systemctl cat felhom-agent >/dev/null 2>&1; then
|
|
agent_cfg=$(systemctl cat felhom-agent 2>/dev/null | grep -oP '(?<=-config )\S+' | head -1)
|
|
fi
|
|
[[ -n "$agent_cfg" ]] || agent_cfg="/etc/felhom-agent/agent.json"
|
|
if systemctl list-unit-files felhom-agent.service >/dev/null 2>&1; then
|
|
systemctl is-active --quiet felhom-agent 2>/dev/null && run systemctl stop felhom-agent
|
|
systemctl is-enabled --quiet felhom-agent 2>/dev/null && run systemctl disable felhom-agent
|
|
else
|
|
log_skip " felhom-agent unit not loaded — skip stop/disable"
|
|
fi
|
|
if [[ -f "$AGENT_UNIT" ]]; then run rm -f "$AGENT_UNIT"; else log_skip " $AGENT_UNIT already absent"; fi
|
|
local bak
|
|
for bak in "${AGENT_UNIT}".bak-*; do [[ -e "$bak" ]] && run rm -f "$bak"; done
|
|
run systemctl daemon-reload
|
|
if [[ -f "$AGENT_SUDOERS" ]]; then run rm -f "$AGENT_SUDOERS"; else log_skip " $AGENT_SUDOERS already absent"; fi
|
|
if [[ -f "$AGENT_BIN" ]]; then run rm -f "$AGENT_BIN"; else log_skip " $AGENT_BIN already absent"; fi
|
|
for bak in "${AGENT_BIN}".bak-*; do [[ -e "$bak" ]] && run rm -f "$bak"; done
|
|
if [[ -d "$AGENT_STATE_DIR" ]]; then run rm -rf "$AGENT_STATE_DIR"; else log_skip " $AGENT_STATE_DIR already absent"; fi
|
|
if id "$AGENT_USER" >/dev/null 2>&1; then run userdel "$AGENT_USER"; else log_skip " service user $AGENT_USER already absent"; fi
|
|
|
|
# 4b. Agent config + its backups (pve token + per-host hub api_key — secrets must not survive;
|
|
# drill R1 / GL-6 F1). The config write leaves `${agent_cfg}.bak*` siblings (e.g. .bak-<ver>,
|
|
# .bak-ceremony-*, .bak-pre064) — one GL-6 residue still held a LIVE hub api_key. Remove the
|
|
# config AND every `.bak*` sibling, then the (now-empty) dir. Paths logged, contents never.
|
|
if [[ -f "$agent_cfg" ]]; then run rm -f "$agent_cfg"; else log_skip " $agent_cfg already absent"; fi
|
|
local _cfgbak
|
|
for _cfgbak in "${agent_cfg}".bak*; do [[ -e "$_cfgbak" ]] && run rm -f "$_cfgbak"; done
|
|
run rmdir "$(dirname "$agent_cfg")" 2>/dev/null || true
|
|
|
|
# 4b2. Management-plane break-glass (TASK G1): timer+oneshot+script+tmpfiles. Stop/disable the
|
|
# timer, remove all four artifacts + the runtime heal-marker. We do NOT `rmdir /run/sshd` —
|
|
# the stock sshd needs it; leaving the (now unit-less) dir in place is correct (a bare kernel
|
|
# /run tmpfs recreates it empty on next boot anyway). Tolerate-absent throughout.
|
|
if systemctl list-unit-files felhom-mgmt-watchdog.timer >/dev/null 2>&1; then
|
|
systemctl is-active --quiet felhom-mgmt-watchdog.timer 2>/dev/null && run systemctl stop felhom-mgmt-watchdog.timer
|
|
systemctl is-enabled --quiet felhom-mgmt-watchdog.timer 2>/dev/null && run systemctl disable felhom-mgmt-watchdog.timer
|
|
else
|
|
log_skip " felhom-mgmt-watchdog.timer not loaded — skip stop/disable"
|
|
fi
|
|
run systemctl reset-failed felhom-mgmt-watchdog.service 2>/dev/null || true
|
|
local wda
|
|
for wda in /etc/systemd/system/felhom-mgmt-watchdog.service /etc/systemd/system/felhom-mgmt-watchdog.timer \
|
|
/usr/local/sbin/felhom-mgmt-watchdog /etc/tmpfiles.d/felhom-privsep.conf /run/felhom-mgmt-watchdog.healed; do
|
|
if [[ -e "$wda" ]]; then run rm -f "$wda"; fi
|
|
done
|
|
|
|
# 4b3. OOB felhom-sshd instance + belt (TASK H1). Stop/disable felhom-sshd + the belt loader,
|
|
# remove units + config tree + belt table + felhom-op user/sudoers. NEVER touch the stock
|
|
# sshd/:22/etc/ssh. Tolerate-absent throughout.
|
|
if systemctl list-unit-files felhom-sshd.service >/dev/null 2>&1; then
|
|
systemctl is-active --quiet felhom-sshd 2>/dev/null && run systemctl stop felhom-sshd
|
|
systemctl is-enabled --quiet felhom-sshd 2>/dev/null && run systemctl disable felhom-sshd
|
|
fi
|
|
if systemctl list-unit-files felhom-oob-nft.service >/dev/null 2>&1; then
|
|
systemctl is-active --quiet felhom-oob-nft 2>/dev/null && run systemctl stop felhom-oob-nft
|
|
systemctl is-enabled --quiet felhom-oob-nft 2>/dev/null && run systemctl disable felhom-oob-nft
|
|
fi
|
|
run systemctl reset-failed felhom-sshd felhom-oob-nft 2>/dev/null || true
|
|
nft list table inet felhom_oob >/dev/null 2>&1 && run nft delete table inet felhom_oob
|
|
local oa
|
|
for oa in /etc/systemd/system/felhom-sshd.service /etc/systemd/system/felhom-oob-nft.service \
|
|
/etc/felhom-oob.nft /etc/sudoers.d/felhom-op /run/felhom-sshd.pid /run/felhom-sshd.healed; do
|
|
if [[ -e "$oa" ]]; then run rm -f "$oa"; fi
|
|
done
|
|
if [[ -d /etc/felhom-sshd ]]; then run rm -rf /etc/felhom-sshd; fi
|
|
if id felhom-op >/dev/null 2>&1; then run userdel -r felhom-op 2>/dev/null || run userdel felhom-op; fi
|
|
|
|
# 4b4. Self-update artifacts (TASK D1; GL-4 gap-closure). Paths derived from the AUTHORITATIVE
|
|
# list in felhom-agent configs/felhom-selfupdate-guarded: the wrapper itself, the A/B slot
|
|
# files it creates next to the live binary (.prev snapshot + orphaned .new.* temps), plus
|
|
# the rollback unit + start-limit drop-in step 5 installs alongside it. pending.json lives
|
|
# under $AGENT_STATE_DIR (already removed in 4). Tolerate-absent throughout.
|
|
if systemctl list-unit-files felhom-agent-rollback.service >/dev/null 2>&1; then
|
|
systemctl is-active --quiet felhom-agent-rollback 2>/dev/null && run systemctl stop felhom-agent-rollback
|
|
systemctl is-enabled --quiet felhom-agent-rollback 2>/dev/null && run systemctl disable felhom-agent-rollback
|
|
fi
|
|
run systemctl reset-failed felhom-agent-rollback.service 2>/dev/null || true
|
|
local sua
|
|
for sua in /usr/local/sbin/felhom-selfupdate-guarded /etc/systemd/system/felhom-agent-rollback.service \
|
|
"${AGENT_BIN}.prev"; do
|
|
if [[ -e "$sua" ]]; then run rm -f "$sua"; fi
|
|
done
|
|
for sua in "${AGENT_BIN}".new.*; do [[ -e "$sua" ]] && run rm -f "$sua"; done
|
|
if [[ -d "${AGENT_UNIT}.d" ]]; then
|
|
if [[ -f "${AGENT_UNIT}.d/felhom-agent-limits.conf" ]]; then run rm -f "${AGENT_UNIT}.d/felhom-agent-limits.conf"; fi
|
|
run rmdir "${AGENT_UNIT}.d" 2>/dev/null || true
|
|
fi
|
|
|
|
# 4c. Shared-parent unit + wrapper + /mnt/felhom-drives (agent-installed at runtime; drill R2).
|
|
# Stop/disable, remove unit + script, unbind + remove the (empty) parent dir. Tolerate-absent.
|
|
if systemctl list-unit-files felhom-shared-parent.service 2>/dev/null | grep -q felhom-shared-parent; then
|
|
systemctl is-active --quiet felhom-shared-parent 2>/dev/null && run systemctl stop felhom-shared-parent
|
|
systemctl is-enabled --quiet felhom-shared-parent 2>/dev/null && run systemctl disable felhom-shared-parent
|
|
else
|
|
log_skip " felhom-shared-parent unit not loaded — skip stop/disable"
|
|
fi
|
|
if [[ -f /etc/systemd/system/felhom-shared-parent.service ]]; then run rm -f /etc/systemd/system/felhom-shared-parent.service; else log_skip " felhom-shared-parent.service already absent"; fi
|
|
if [[ -f /usr/local/sbin/felhom-shared-parent.sh ]]; then run rm -f /usr/local/sbin/felhom-shared-parent.sh; fi
|
|
run systemctl daemon-reload
|
|
# GL-4: unmount every enrolled/network drive mounted UNDER /mnt/felhom-drives (deepest first)
|
|
# BEFORE the root self-bind. Plain umount ONLY — NEVER -l/-f: a lazy/forced unmount on a busy
|
|
# data mount risks the customer's data; a busy mount gets a warning + guidance instead. The
|
|
# DATA STAYS ON THE DRIVE — nothing here (or anywhere in this script) wipes or formats it.
|
|
local dmnt
|
|
while IFS= read -r dmnt; do
|
|
[[ -n "$dmnt" ]] || continue
|
|
if $DRY_RUN; then
|
|
log_dry "umount $dmnt # data stays on the drive"
|
|
elif umount "$dmnt" 2>/dev/null; then
|
|
log_success " unmounted $dmnt (data stays on the drive)"
|
|
else
|
|
log_warn " $dmnt is busy — NOT forcing (never umount -l/-f). Eject the drive via the dashboard, or stop the apps using it and re-run."
|
|
_busy_mounts+=("$dmnt")
|
|
continue
|
|
fi
|
|
run rmdir "$dmnt" 2>/dev/null || true
|
|
done < <(findmnt -rn -o TARGET 2>/dev/null | grep '^/mnt/felhom-drives/' | sort -r || true)
|
|
if mountpoint -q /mnt/felhom-drives 2>/dev/null; then
|
|
if [[ ${#_busy_mounts[@]} -gt 0 ]]; then
|
|
log_warn " /mnt/felhom-drives root bind left mounted (busy drive mounts above must go first)"
|
|
else
|
|
run umount /mnt/felhom-drives
|
|
fi
|
|
fi
|
|
if [[ -d /mnt/felhom-drives ]]; then run rmdir /mnt/felhom-drives 2>/dev/null || true; fi
|
|
|
|
# 4d. Guarded-mkfs + pbs-apply wrappers, guest-hook snippet, lan-resolver dnsmasq snippets (drill R3-R5).
|
|
if [[ -f /usr/local/sbin/felhom-mkfs-guarded ]]; then run rm -f /usr/local/sbin/felhom-mkfs-guarded; else log_skip " felhom-mkfs-guarded already absent"; fi
|
|
if [[ -f /usr/local/sbin/felhom-pbs-apply ]]; then run rm -f /usr/local/sbin/felhom-pbs-apply; else log_skip " felhom-pbs-apply already absent"; fi
|
|
if [[ -f /usr/local/sbin/felhom-backup-target-apply ]]; then run rm -f /usr/local/sbin/felhom-backup-target-apply; else log_skip " felhom-backup-target-apply already absent"; fi
|
|
if [[ -f /var/lib/vz/snippets/felhom-guest-hook.sh ]]; then run rm -f /var/lib/vz/snippets/felhom-guest-hook.sh; fi
|
|
local dconf _dnsmasq_touched=false
|
|
for dconf in /etc/dnsmasq.d/felhom-*.conf; do
|
|
[[ -e "$dconf" ]] || continue
|
|
run rm -f "$dconf"; _dnsmasq_touched=true
|
|
done
|
|
if $_dnsmasq_touched && systemctl is-active --quiet dnsmasq 2>/dev/null; then
|
|
run systemctl restart dnsmasq || true
|
|
fi
|
|
|
|
# 5. pveum removal (presence-checked; tolerate-absent; roles deleted only after their grants).
|
|
# Remove the 3-role scoped grants+roles (3b) AND the pre-3b single-role broad grant if present —
|
|
# both tolerate-absent so --uninstall works on a box of either shape.
|
|
remove_scoped_acl
|
|
remove_old_broad_acl
|
|
# token then user (token-remove purges its ACL; user-delete purges anything else).
|
|
if pveum user token list "$PVE_USER" --output-format json 2>/dev/null | python3 -c "import json,sys;sys.exit(0 if any(t['tokenid']=='$PVE_TOKENID' for t in json.load(sys.stdin)) else 1)" 2>/dev/null; then
|
|
run pveum user token remove "$PVE_USER" "$PVE_TOKENID"
|
|
else
|
|
log_skip " token ${PVE_USER}!${PVE_TOKENID} already absent"
|
|
fi
|
|
if pveum user list --output-format json 2>/dev/null | python3 -c "import json,sys;sys.exit(0 if any(u['userid']=='$PVE_USER' for u in json.load(sys.stdin)) else 1)" 2>/dev/null; then
|
|
run pveum user delete "$PVE_USER"
|
|
else
|
|
log_skip " user $PVE_USER already absent"
|
|
fi
|
|
|
|
# 5b. felhom pool — delete ONLY if empty (a destroyed guest is auto-removed from its pool). Never
|
|
# delete a pool that still holds members (someone else's guests, or another Felhom guest kept
|
|
# under --force).
|
|
if pool_exists; then
|
|
local pool_left; pool_left=$(pool_members)
|
|
if [[ -z "$pool_left" ]]; then
|
|
run pveum pool delete "$PVE_POOL"
|
|
pool_removed=true
|
|
else
|
|
log_skip " pool $PVE_POOL not empty (members: $pool_left) — leaving it"
|
|
fi
|
|
else
|
|
log_skip " pool $PVE_POOL already absent"
|
|
fi
|
|
|
|
# 6. Golden vzdump (opt-in via --remove-golden; else left in place).
|
|
if $REMOVE_GOLDEN; then
|
|
local gvols gv
|
|
gvols=$(pvesm list "$ARCHIVE_STORAGE" --content backup 2>/dev/null | awk -v v="$GOLDEN_VMID" '$0 ~ ("vzdump-lxc-" v "-"){print $1}')
|
|
if [[ -n "$gvols" ]]; then
|
|
while IFS= read -r gv; do
|
|
[[ -n "$gv" ]] || continue
|
|
run pvesm free "$gv"
|
|
done <<<"$gvols"
|
|
log_success " removed golden vzdump(s) from $ARCHIVE_STORAGE"
|
|
else
|
|
log_skip " no golden vzdump (vzdump-lxc-${GOLDEN_VMID}-*) on $ARCHIVE_STORAGE"
|
|
fi
|
|
else
|
|
log_skip " golden vzdump left in place (pass --remove-golden to remove)"
|
|
fi
|
|
|
|
# 7. Install state file (only reached when host-level removal ran — safe-skip returned earlier).
|
|
if [[ -f "$STATE_FILE" ]]; then run rm -f "$STATE_FILE"; else log_skip " $STATE_FILE already absent"; fi
|
|
run rmdir "$STATE_DIR" 2>/dev/null || true
|
|
|
|
# 8. Summary + the GL-4 kept-vs-wiped statement.
|
|
_uninstall_statement full
|
|
log_success "UNINSTALL complete — removed: guest $vmid, the felhom-agent (unit/sudoers/binary/state/config+baks/user + selfupdate-artifacts/shared-parent/mkfs-wrapper/pbs-apply-wrapper/hook-snippet/dnsmasq-snippets), the pveum role/user/token/ACL,$( $pool_removed && printf ' the %s pool,' "$PVE_POOL") and $STATE_FILE."
|
|
if $REMOVE_GOLDEN; then log_info " golden vzdump: removed."; else log_info " golden vzdump: left in place (--remove-golden to remove)."; fi
|
|
log_info " NOTE: the 'sudo' and 'dnsmasq' packages were left installed (system packages); the host record still exists in the hub — remove it there if desired."
|
|
$DRY_RUN && log_warn " DRY-RUN: nothing above was actually executed."
|
|
return 0
|
|
}
|
|
|
|
# run_adopt_pool — retrofit an EXISTING Felhom guest into the felhom pool. Non-destructive: creates the
|
|
# pool if absent + adds the guest; never reconfigures/restarts the guest, never contacts the hub. Guest
|
|
# resolves from --vmid else the recorded provisioned_vmid (mirrors run_uninstall).
|
|
run_adopt_pool() {
|
|
log_step "ADOPT-POOL — add an existing Felhom guest to the $PVE_POOL pool"
|
|
|
|
local state_vmid vmid
|
|
state_vmid=$(_state_get provisioned_vmid)
|
|
if $VMID_EXPLICIT; then
|
|
vmid="$VMID"
|
|
elif [[ -n "$state_vmid" ]]; then
|
|
vmid="$state_vmid"
|
|
log_info " no --vmid given; using recorded provisioned_vmid=$vmid from $STATE_FILE"
|
|
else
|
|
die "pass --vmid N (state has no recorded vmid)"
|
|
fi
|
|
|
|
_vmid_in_use "$vmid" || die "guest $vmid not found on this host (nothing to adopt)"
|
|
|
|
# ours-check: only adopt a Felhom guest (has the /etc/felhom-bootstrap mount) unless --force.
|
|
if pct config "$vmid" 2>/dev/null | grep -q 'mp=/etc/felhom-bootstrap'; then
|
|
log_info " vmid $vmid looks like a Felhom guest (has the /etc/felhom-bootstrap mount)"
|
|
elif $FORCE; then
|
|
log_warn " vmid $vmid has NO /etc/felhom-bootstrap mount — --force given, adopting anyway"
|
|
else
|
|
die "vmid $vmid does not look like a Felhom-provisioned guest (no /etc/felhom-bootstrap mount).
|
|
Refusing to adopt. Pass --force to override."
|
|
fi
|
|
|
|
ensure_felhom_pool
|
|
pool_add_guest "$vmid"
|
|
|
|
echo ""
|
|
log_success "ADOPT-POOL complete — guest $vmid is in pool $PVE_POOL (guest not otherwise modified)."
|
|
$DRY_RUN && log_warn " DRY-RUN: nothing above was actually executed."
|
|
return 0
|
|
}
|
|
|
|
# run_rescope_acl — migrate an EXISTING install from the pre-3b broad-/ token to the pool-scoped ACL.
|
|
# Non-destructive to data (ACL-only): ensure the pool + the guest is a member, apply the 3-role scoped
|
|
# grants, THEN remove the old broad grant. Idempotent + dry-run-aware. Does NOT touch the guest or hub.
|
|
# ORDERING (see §13): run this with the agent STOPPED, then deploy agent >= v0.53.0, then start — the
|
|
# scoped ACL and the pool-param agent are mutually dependent.
|
|
run_rescope_acl() {
|
|
log_step "RESCOPE-ACL — migrate to the pool-scoped token ACL"
|
|
|
|
local state_vmid vmid
|
|
state_vmid=$(_state_get provisioned_vmid)
|
|
if $VMID_EXPLICIT; then
|
|
vmid="$VMID"
|
|
elif [[ -n "$state_vmid" ]]; then
|
|
vmid="$state_vmid"
|
|
log_info " no --vmid given; using recorded provisioned_vmid=$vmid from $STATE_FILE"
|
|
else
|
|
die "pass --vmid N (state has no recorded vmid)"
|
|
fi
|
|
|
|
_vmid_in_use "$vmid" || die "guest $vmid not found on this host"
|
|
if pct config "$vmid" 2>/dev/null | grep -q 'mp=/etc/felhom-bootstrap'; then
|
|
log_info " vmid $vmid looks like a Felhom guest (has the /etc/felhom-bootstrap mount)"
|
|
elif $FORCE; then
|
|
log_warn " vmid $vmid has NO /etc/felhom-bootstrap mount — --force given, rescoping anyway"
|
|
else
|
|
die "vmid $vmid does not look like a Felhom-provisioned guest (no /etc/felhom-bootstrap mount).
|
|
Refusing to rescope. Pass --force to override."
|
|
fi
|
|
|
|
# The guest MUST be a pool member before the scoped token can touch it — ensure it first.
|
|
ensure_felhom_pool
|
|
pool_add_guest "$vmid"
|
|
# Apply the scoped grants, THEN remove the old broad grant (add-before-remove: never leave the token
|
|
# with NO grant mid-migration).
|
|
apply_scoped_acl
|
|
remove_old_broad_acl
|
|
|
|
echo ""
|
|
log_success "RESCOPE-ACL complete — token scoped to /pool/$PVE_POOL + /storage/[${PVE_STORAGES[*]}] + Sys.Audit/SDN.Use@/."
|
|
log_warn " NOW deploy agent >= v0.53.0 (restore-into-pool) and (re)start felhom-agent — the scoped ACL needs it."
|
|
$DRY_RUN && log_warn " DRY-RUN: nothing above was actually executed."
|
|
return 0
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# Arg parse
|
|
#-------------------------------------------------------------------------------
|
|
while [[ $# -gt 0 ]]; do
|
|
case "$1" in
|
|
--customer-id) CUSTOMER_ID="$2"; shift 2 ;;
|
|
--mode) MODE="$2"; shift 2 ;;
|
|
--hub-url) HUB_URL="$2"; shift 2 ;;
|
|
--vmid) VMID="$2"; VMID_EXPLICIT=true; shift 2 ;;
|
|
--golden) GOLDEN_VOLID="$2"; shift 2 ;;
|
|
--golden-vmid) GOLDEN_VMID="$2"; shift 2 ;;
|
|
--archive-storage) ARCHIVE_STORAGE="$2"; shift 2 ;;
|
|
--node) NODE="$2"; NODE_EXPLICIT=true; shift 2 ;;
|
|
--bridge-ip) BRIDGE_ADDR="$2"; shift 2 ;;
|
|
--no-island) WANT_ISLAND=false; shift ;; # R-50: keep the historical LAN bind on an appliance
|
|
--rootfs-grow) ROOTFS_GROW="$2"; shift 2 ;;
|
|
--datavol-grow) DATAVOL_GROW="$2"; shift 2 ;;
|
|
--sysdata-grow) SYSDATA_GROW="$2"; shift 2 ;;
|
|
--cores) CPU_CORES="$2"; shift 2 ;;
|
|
--memory) MEM_MIB="$2"; shift 2 ;;
|
|
--passphrase-file) PASSPHRASE_FILE="$2"; shift 2 ;;
|
|
--operator-pubkey-file) OPERATOR_PUBKEY_FILE="$2"; shift 2 ;;
|
|
--preserve-from) PRESERVE_FROM="$2"; shift 2 ;;
|
|
--preserve-state-from) PRESERVE_STATE_FROM="$2"; shift 2 ;;
|
|
--allow-new-leaf) ALLOW_NEW_LEAF=true; shift ;;
|
|
--force) FORCE=true; shift ;;
|
|
--force-gitea-golden) FORCE_GITEA_GOLDEN=true; shift ;;
|
|
--skip-provision) SKIP_PROVISION=true; shift ;;
|
|
--uninstall) UNINSTALL=true; shift ;;
|
|
--remove-golden) REMOVE_GOLDEN=true; shift ;;
|
|
--adopt-pool) ADOPT_POOL=true; shift ;;
|
|
--rescope-acl) RESCOPE_ACL=true; shift ;;
|
|
--rotate-recovery) ROTATE_RECOVERY=true; shift ;;
|
|
--enable-oob) ENABLE_OOB=true; shift ;;
|
|
--no-oob) NO_OOB=true; shift ;;
|
|
--acl-storages) read -ra PVE_STORAGES <<< "$2"; shift 2 ;;
|
|
--preflight-only) PREFLIGHT_ONLY=true; shift ;;
|
|
--dry-run) DRY_RUN=true; shift ;;
|
|
--resume) RESUME=true; shift ;;
|
|
-h|--help) usage ;;
|
|
*) die "Unknown option: $1 (use -h)" ;;
|
|
esac
|
|
done
|
|
|
|
#===============================================================================
|
|
# UNINSTALL MODE — local host teardown (no hub contact, no passphrase). Dispatched early,
|
|
# before any provision/DR logic, and does not require --customer-id.
|
|
#===============================================================================
|
|
if $UNINSTALL; then
|
|
[[ $EUID -eq 0 ]] || die "must run as root"
|
|
echo ""
|
|
log_info "felhom-host-install v${SCRIPT_VERSION} — mode=uninstall"
|
|
$DRY_RUN && log_warn "DRY-RUN: no mutations will be performed"
|
|
echo ""
|
|
run_uninstall
|
|
exit 0
|
|
fi
|
|
|
|
#===============================================================================
|
|
# ADOPT-POOL MODE — retrofit an EXISTING Felhom guest into the felhom pool (non-destructive; no hub
|
|
# contact, no passphrase, no guest reconfigure beyond pool membership). Dispatched early.
|
|
#===============================================================================
|
|
if $ADOPT_POOL; then
|
|
[[ $EUID -eq 0 ]] || die "must run as root"
|
|
echo ""
|
|
log_info "felhom-host-install v${SCRIPT_VERSION} — mode=adopt-pool"
|
|
$DRY_RUN && log_warn "DRY-RUN: no mutations will be performed"
|
|
echo ""
|
|
run_adopt_pool
|
|
exit 0
|
|
fi
|
|
|
|
#===============================================================================
|
|
# RESCOPE-ACL MODE — migrate an existing install to the pool-scoped token ACL (ACL-only, no hub, no
|
|
# passphrase). Supervised: run with the agent stopped, then deploy agent >= v0.53.0 (see §13).
|
|
#===============================================================================
|
|
if $RESCOPE_ACL; then
|
|
[[ $EUID -eq 0 ]] || die "must run as root"
|
|
echo ""
|
|
log_info "felhom-host-install v${SCRIPT_VERSION} — mode=rescope-acl"
|
|
$DRY_RUN && log_warn "DRY-RUN: no mutations will be performed"
|
|
echo ""
|
|
run_rescope_acl
|
|
exit 0
|
|
fi
|
|
|
|
#===============================================================================
|
|
# DR MODE — documented seam only (10D). NOT implemented.
|
|
#===============================================================================
|
|
if [[ "$MODE" == "dr" ]]; then
|
|
log_error "DR mode not yet implemented (10D)."
|
|
cat >&2 <<'EOF'
|
|
The DR step skeleton (for the future implementer) mirrors provision EXCEPT
|
|
the restore source:
|
|
1. pre-flight (root, PVE, hub reachable, customer+passphrase valid)
|
|
2. pveum token (identical to provision)
|
|
3. host-enroll (mint-once-reuse — the lost host re-binds to its customer)
|
|
4. agent config write (identical)
|
|
5. RESTORE: instead of the golden, restore the customer's OWN whole-CT PBS
|
|
snapshot (continuity preserved) — agent --selftest=bring-up -mode dr
|
|
-archive <pbs-snapshot>. Identity/keys come from escrow + the hub recipe.
|
|
6. verify (identical)
|
|
EOF
|
|
exit 2
|
|
fi
|
|
#===============================================================================
|
|
# GL-2 install-profile validation — everything below runs a real install (or its --preflight-only /
|
|
# --resume variant), so the profile is decided HERE, before the passphrase prompt and any step. The
|
|
# refusals are deliberately argv-time (the harness relies on them firing on a non-PVE machine too).
|
|
#===============================================================================
|
|
|
|
# --- F5 (VALIDATION-n100): appliance guest auto-sizing --------------------------
|
|
# The golden default (2 GB RAM) was too small on a 16 GB host. In APPLIANCE mode, when the operator
|
|
# passes no explicit cap, size the guest generously from the host — LXC limits are cheap, so err
|
|
# generous. An explicit --cores/--memory ALWAYS wins (never overwritten). byo mode requires explicit
|
|
# caps (never auto-sized). The host reads are overridable for the mode harness (no real /proc dep).
|
|
host_total_mib() {
|
|
if [[ -n "${FELHOM_FAKE_MEMTOTAL_MIB:-}" ]]; then echo "$FELHOM_FAKE_MEMTOTAL_MIB"; return; fi
|
|
awk '/^MemTotal:/{print int($2/1024)}' /proc/meminfo 2>/dev/null || echo 0
|
|
}
|
|
host_core_count() {
|
|
if [[ -n "${FELHOM_FAKE_NPROC:-}" ]]; then echo "$FELHOM_FAKE_NPROC"; return; fi
|
|
nproc 2>/dev/null || echo 0
|
|
}
|
|
autosize_guest_caps() {
|
|
[[ "$MODE" == "appliance" ]] || return 0
|
|
if [[ -z "$MEM_MIB" ]]; then
|
|
local total; total=$(host_total_mib)
|
|
if [[ "${total:-0}" -gt 0 ]]; then
|
|
# clamp(host-4096, min 4096, max host-2048), then the hard ceiling host-1024 (never
|
|
# over-commit). Order matters: the min floor is applied AFTER the generous max so that on
|
|
# a host too small for both the MIN wins (edge rule), bounded only by the host-1024 ceiling.
|
|
local mem=$(( total - 4096 ))
|
|
local max2=$(( total - 2048 ))
|
|
(( mem > max2 )) && mem=$max2
|
|
(( mem < 4096 )) && mem=4096
|
|
local ceil=$(( total - 1024 ))
|
|
(( mem > ceil )) && mem=$ceil
|
|
MEM_MIB=$mem
|
|
log_info " auto-sized guest RAM: ${MEM_MIB} MiB (host ${total} MiB; clamp(host-4096, min 4096, max host-2048), ceiling host-1024)"
|
|
fi
|
|
fi
|
|
if [[ -z "$CPU_CORES" ]]; then
|
|
local cores; cores=$(host_core_count)
|
|
if [[ "${cores:-0}" -gt 0 ]]; then
|
|
local c=$(( cores - 1 ))
|
|
(( c < 2 )) && c=2
|
|
CPU_CORES=$c
|
|
log_info " auto-sized guest cores: ${CPU_CORES} (host ${cores} cores; host-1, min 2)"
|
|
fi
|
|
fi
|
|
}
|
|
|
|
case "$MODE" in
|
|
appliance|byo) ;;
|
|
"")
|
|
if $PREFLIGHT_ONLY; then
|
|
log_warn "--preflight-only without --mode: running the mode-agnostic base checks only"
|
|
else
|
|
die "--mode is required: pass --mode appliance (a Felhom-owned box) or --mode byo (a host you do not own). See -h."
|
|
fi
|
|
;;
|
|
provision) die "--mode provision was retired in v1.10.0 — pass --mode appliance (Felhom-owned box) or --mode byo (a host you do not own)." ;;
|
|
*) die "Unknown --mode: $MODE (appliance|byo)" ;;
|
|
esac
|
|
|
|
# F5: fill the appliance guest caps from the host when the operator gave none (explicit flags win).
|
|
autosize_guest_caps
|
|
|
|
# BYO argument refusals (C1/C2) — before the passphrase prompt, before any step.
|
|
if [[ "$MODE" == "byo" ]]; then
|
|
if [[ -z "$CPU_CORES" || -z "$MEM_MIB" ]]; then
|
|
die "byo mode requires explicit --cores and --memory (noisy-neighbor caps on a host you do not own)"
|
|
fi
|
|
if $ENABLE_OOB; then
|
|
die "--enable-oob is not allowed in byo mode (OOB access on a host the operator does not own is the owner's call; provision it as a separate, owner-consented step)"
|
|
fi
|
|
if $ROTATE_RECOVERY; then
|
|
die "--rotate-recovery is not allowed in byo mode (byo never touches root@pam — there is no recovery credential to rotate)"
|
|
fi
|
|
fi
|
|
|
|
# H1 / F9 (ruling 2026-07-23 "install everywhere"): the OOB belt is now a DEFAULT appliance leg — a box
|
|
# the operator owns must have an out-of-band way in (the 07-22 power-outage audit found the belt on
|
|
# NEITHER fleet box because the universal ISO never passed --enable-oob). Appliance installs it unless
|
|
# --no-oob; byo stays refuse-by-default above (owner-consented, never automatic). The belt is
|
|
# lockout-SAFE if the fleet has no operator peer yet: empty sets + no felhom-sshd key = no access
|
|
# granted and nothing on :22 touched, until the operator OOB peer is registered hub-side.
|
|
if [[ "$MODE" == "appliance" ]] && ! $NO_OOB; then
|
|
ENABLE_OOB=true
|
|
fi
|
|
|
|
# GL-4: resolve the operator signing keys — script constants by default; --operator-pubkey-file
|
|
# OVERRIDES them (C5). Validated HERE (before the passphrase prompt, before any step) so a bad key
|
|
# file dies fast and the harness can exercise it on a non-PVE machine. File format: one key per
|
|
# line, "operational <authorized_keys line>" or "recovery <authorized_keys line>"; '#' comments and
|
|
# blank lines allowed. The key_id is the authorized_keys COMMENT field — required (the agent's
|
|
# signed-op verifier addresses keys by key_id).
|
|
RESOLVED_OP_ID="$OPERATOR_KEY_OPERATIONAL_ID"; RESOLVED_OP_LINE="$OPERATOR_KEY_OPERATIONAL_LINE"
|
|
RESOLVED_REC_ID="$OPERATOR_KEY_RECOVERY_ID"; RESOLVED_REC_LINE="$OPERATOR_KEY_RECOVERY_LINE"
|
|
if [[ -n "$OPERATOR_PUBKEY_FILE" ]]; then
|
|
[[ -f "$OPERATOR_PUBKEY_FILE" ]] || die "--operator-pubkey-file not found: $OPERATOR_PUBKEY_FILE"
|
|
if [[ -n "$RESOLVED_OP_LINE$RESOLVED_REC_LINE" ]]; then
|
|
log_info "--operator-pubkey-file overrides the script's built-in operator key constants"
|
|
fi
|
|
RESOLVED_OP_ID=""; RESOLVED_OP_LINE=""; RESOLVED_REC_ID=""; RESOLVED_REC_LINE=""
|
|
_kseen=false
|
|
while IFS= read -r _kline || [[ -n "$_kline" ]]; do
|
|
_kline="${_kline%$'\r'}"
|
|
[[ -z "$_kline" || "$_kline" == \#* ]] && continue
|
|
_krole="${_kline%% *}"; _krest="${_kline#* }"
|
|
case "$_krole" in
|
|
operational|recovery) ;;
|
|
*) die "--operator-pubkey-file: unknown role '$_krole' (want operational|recovery) in line: $_kline" ;;
|
|
esac
|
|
read -r _ktype _kb64 _kid _ <<<"$_krest"
|
|
[[ "$_ktype" == ssh-* || "$_ktype" == sk-ssh-* || "$_ktype" == ecdsa-* ]] \
|
|
|| die "--operator-pubkey-file: not an authorized_keys line (bad key type '${_ktype:-<empty>}') in line: $_kline"
|
|
[[ -n "$_kb64" && "$_kb64" =~ ^[A-Za-z0-9+/=]+$ ]] \
|
|
|| die "--operator-pubkey-file: not an authorized_keys line (missing/invalid key material) in line: $_kline"
|
|
[[ -n "$_kid" ]] \
|
|
|| die "--operator-pubkey-file: key line has no comment field — the comment IS the key_id; add one (e.g. felhom-op-1): $_kline"
|
|
if [[ "$_krole" == operational ]]; then
|
|
RESOLVED_OP_ID="$_kid"; RESOLVED_OP_LINE="$_krest"
|
|
else
|
|
RESOLVED_REC_ID="$_kid"; RESOLVED_REC_LINE="$_krest"
|
|
fi
|
|
_kseen=true
|
|
done < "$OPERATOR_PUBKEY_FILE"
|
|
$_kseen || die "--operator-pubkey-file has no key lines (empty/comments only): $OPERATOR_PUBKEY_FILE"
|
|
fi
|
|
|
|
# Resume mode-mismatch (C4): an install keeps the mode it started with — mode-flipping a half-done
|
|
# install would skip already-completed steps under the WRONG profile's guarantees.
|
|
if $RESUME; then
|
|
_recorded_mode=$(_state_get mode || true)
|
|
if [[ -n "$_recorded_mode" && "$_recorded_mode" != "$MODE" ]]; then
|
|
die "install started as $_recorded_mode; resume with --mode $_recorded_mode or start over"
|
|
fi
|
|
fi
|
|
|
|
#===============================================================================
|
|
# PROVISION MODE
|
|
#===============================================================================
|
|
[[ -n "$CUSTOMER_ID" ]] || die "--customer-id is required (use -h)"
|
|
|
|
echo ""
|
|
log_info "felhom-host-install v${SCRIPT_VERSION} — mode=${MODE:-<preflight-only, no mode>} customer=${CUSTOMER_ID} vmid=${VMID}"
|
|
$DRY_RUN && log_warn "DRY-RUN: no mutations will be performed"
|
|
echo ""
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# Read passphrase (no-echo prompt or 0600 file) — never on argv/logs
|
|
#-------------------------------------------------------------------------------
|
|
read_passphrase() {
|
|
if [[ -n "$PASSPHRASE_FILE" ]]; then
|
|
[[ -f "$PASSPHRASE_FILE" ]] || die "--passphrase-file not found: $PASSPHRASE_FILE"
|
|
local perm; perm=$(stat -c '%a' "$PASSPHRASE_FILE")
|
|
[[ "$perm" == "600" || "$perm" == "400" ]] || log_warn "passphrase file $PASSPHRASE_FILE is mode $perm (want 600)"
|
|
PASSPHRASE="$(< "$PASSPHRASE_FILE")"; PASSPHRASE="${PASSPHRASE%$'\n'}"
|
|
else
|
|
# Read from the terminal explicitly (not stdin), so the no-echo prompt works whether the
|
|
# script is run from a file OR piped to bash (curl … | sudo bash) — where stdin is the pipe.
|
|
read -rsp "Retrieval passphrase for customer '${CUSTOMER_ID}': " PASSPHRASE < /dev/tty; echo ""
|
|
fi
|
|
[[ -n "$PASSPHRASE" ]] || die "empty passphrase"
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# BYO host-mutation disclosure + typed acknowledgement (GL-2). Printed at the END of preflight so
|
|
# every value shown (vmid, storages, config path) is final. The list follows the CODE (steps 2-8 +
|
|
# the agent's runtime installs), not memory — keep it in sync when a step grows a new host artifact.
|
|
# The ack reads /dev/tty (the script is often `curl | bash`) and mirrors the uninstall confirm's
|
|
# dry-run branch. NO credential is ever printed here — artifact PATHS only.
|
|
#-------------------------------------------------------------------------------
|
|
_byo_disclosure_ack() {
|
|
echo ""
|
|
log_step "BYO disclosure — installing Felhom on a host the operator does NOT own will create:"
|
|
cat <<EOF
|
|
pveum: roles ${PVE_ROLE_BASE}/${PVE_ROLE_GUEST}/${PVE_ROLE_STORE} + user ${PVE_USER} + token ${PVE_USER}!${PVE_TOKENID}
|
|
+ scoped ACL grants (Base@/, Guest@/pool/${PVE_POOL}, Store@[${PVE_STORAGES[*]}]) + the '${PVE_POOL}' pool
|
|
agent: service user ${AGENT_USER} + ${AGENT_BIN} (+ .bak-* backups; A/B slots via the guarded self-update)
|
|
+ ${AGENT_UNIT} (+ felhom-agent-rollback.service + start-limit drop-in)
|
|
+ ${AGENT_SUDOERS} + state dir ${AGENT_STATE_DIR} + config ${AGENT_CONFIG} (0600)
|
|
host: /usr/local/sbin/felhom-mkfs-guarded + /usr/local/sbin/felhom-selfupdate-guarded
|
|
+ /usr/local/sbin/felhom-pbs-apply (PBS-DR apply wrapper — DR capability is baked
|
|
on every install; ACTIVATION stays a hub flag, off = zero effect on this host)
|
|
+ felhom-mgmt-watchdog service+timer+script + /etc/tmpfiles.d/felhom-privsep.conf
|
|
+ guest-hook snippet under /var/lib/vz/snippets/ (agent-installed at runtime)
|
|
+ the 'sudo' and 'age' packages if absent + install state dir ${STATE_DIR}
|
|
wg: an OUTBOUND WireGuard tunnel to the Felhom hub (wg_tunnel.enabled=true — base
|
|
infrastructure like the cloudflared tunnel; hands-free peer registration; the
|
|
operator can disable it per-customer hub-side on request)
|
|
EOF
|
|
if $SKIP_PROVISION; then
|
|
echo " guest: none (--skip-provision)"
|
|
else
|
|
echo " guest: the provisioned Felhom LXC (vmid ${VMID}, capped ${CPU_CORES} cores / ${MEM_MIB} MiB) + its volumes"
|
|
echo " + the golden vzdump imported onto storage '${ARCHIVE_STORAGE}'"
|
|
fi
|
|
echo " update: operator-signed self-update authority: ${RESOLVED_OP_ID:-NONE (self-update stays dormant)}"
|
|
echo " NOT touched in byo mode: root@pam (no break-glass), host DNS (:53), OOB sshd."
|
|
echo ""
|
|
if $DRY_RUN; then
|
|
log_dry "would prompt: Type this host's short hostname ($(hostname -s)) to acknowledge the byo install"
|
|
elif $PREFLIGHT_ONLY; then
|
|
log_info " (--preflight-only: acknowledgement prompt skipped — nothing will be mutated)"
|
|
else
|
|
local _hn _ans; _hn=$(hostname -s)
|
|
read -rp "Type this host's short hostname (${_hn}) to acknowledge the byo install: " _ans < /dev/tty
|
|
[[ "$_ans" == "$_hn" ]] || die "acknowledgement mismatch (got '${_ans}', expected '${_hn}') — aborting, nothing installed"
|
|
log_success " byo install acknowledged for host ${_hn}"
|
|
fi
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# STEP 1 — pre-flight (fail fast before any mutation)
|
|
#-------------------------------------------------------------------------------
|
|
step_preflight() {
|
|
log_step "1/8 pre-flight"
|
|
[[ $EUID -eq 0 ]] || die "must run as root"
|
|
command -v pveum >/dev/null || die "pveum not found — is this a Proxmox VE host?"
|
|
command -v pct >/dev/null || die "pct not found — is this a Proxmox VE host?"
|
|
command -v pvesh >/dev/null || die "pvesh not found"
|
|
command -v curl >/dev/null || die "curl not found"
|
|
command -v python3>/dev/null || die "python3 not found"
|
|
|
|
local pvever; pvever=$(pveversion | head -1)
|
|
# GL-2 PVE-major gate: byo installs are validated on PVE 9.x ONLY — on a host we don't own, an
|
|
# unvalidated major is a refusal, not a warning. Appliance keeps the historical warn-only.
|
|
if [[ "$pvever" != *"/9."* && "$MODE" == "byo" ]] && ! $FORCE; then
|
|
die "byo mode is validated on PVE 9.x only (got: $pvever). Pass --force to override deliberately."
|
|
fi
|
|
[[ "$pvever" == *"/9."* ]] || log_warn "expected PVE 9.x, got: $pvever"
|
|
log_info " $pvever"
|
|
|
|
# node — on a MULTI-NODE cluster, auto-selecting nodes[0] is a wrong-node footgun. Require an
|
|
# explicit --node unless there is exactly one node (or no guest will be provisioned).
|
|
local nodes_json node_count node_names
|
|
nodes_json=$(pvesh get /nodes --output-format json 2>/dev/null || echo "[]")
|
|
node_count=$(python3 -c "import json,sys;print(len(json.loads(sys.argv[1])))" "$nodes_json" 2>/dev/null || echo 0)
|
|
if [[ "${node_count:-0}" -gt 1 ]] && ! $NODE_EXPLICIT && ! $SKIP_PROVISION; then
|
|
node_names=$(python3 -c "import json,sys;print(', '.join(n['node'] for n in json.loads(sys.argv[1])))" "$nodes_json" 2>/dev/null || echo "?")
|
|
die "this is a ${node_count}-node cluster (${node_names}); pass --node explicitly — auto-selecting nodes[0] risks provisioning on the wrong node."
|
|
fi
|
|
if [[ -z "$NODE" ]]; then
|
|
NODE=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])[0]['node'])" "$nodes_json" 2>/dev/null || hostname)
|
|
fi
|
|
if $NODE_EXPLICIT; then log_info " node: $NODE (explicit)"; else log_info " node: $NODE (auto)"; fi
|
|
|
|
# agent config path: honor the existing systemd unit's -config, else default
|
|
if systemctl cat felhom-agent >/dev/null 2>&1; then
|
|
AGENT_CONFIG=$(systemctl cat felhom-agent 2>/dev/null | grep -oP '(?<=-config )\S+' | head -1)
|
|
fi
|
|
[[ -n "$AGENT_CONFIG" ]] || AGENT_CONFIG="/etc/felhom-agent/agent.json"
|
|
log_info " agent config: $AGENT_CONFIG"
|
|
# v1.1.0: the agent binary is no longer a prerequisite — the agent-install step (5/8) fetches it
|
|
# from Gitea + verifies it. Just report what's present (if anything).
|
|
if command -v felhom-agent >/dev/null 2>&1; then
|
|
log_info " agent (existing): $(felhom-agent --version 2>&1 | head -1)"
|
|
else
|
|
log_info " agent: not installed yet — will be fetched + installed in step 5/8"
|
|
fi
|
|
|
|
# local-lvm free space
|
|
local free_gib
|
|
free_gib=$(lvs --noheadings --units g -o lv_size,data_percent /dev/pve/data 2>/dev/null | awk '{gsub(/[^0-9.]/,"",$1); used=$2; print int($1*(100-used)/100)}' 2>/dev/null || echo 0)
|
|
if [[ "${free_gib:-0}" -gt 0 ]]; then
|
|
log_info " local-lvm free: ~${free_gib} GiB"
|
|
[[ "$free_gib" -ge "$HARD_MIN_LVM_GIB" ]] || log_warn "local-lvm free ~${free_gib} GiB < hard min ${HARD_MIN_LVM_GIB} GiB"
|
|
else
|
|
log_warn " could not read local-lvm free space (continuing)"
|
|
fi
|
|
|
|
# RAM floor (soft): a big appliance guest can pressure existing guests on a small box. WARN only.
|
|
local mem_avail_mib
|
|
mem_avail_mib=$(awk '/^MemAvailable:/{print int($2/1024)}' /proc/meminfo 2>/dev/null || echo 0)
|
|
if [[ "${mem_avail_mib:-0}" -gt 0 ]]; then
|
|
if [[ "$mem_avail_mib" -lt 2048 ]]; then
|
|
log_warn " low free RAM (~${mem_avail_mib} MiB); the appliance guest may pressure existing guests — consider the CPU/mem cap."
|
|
else
|
|
log_info " free RAM: ~${mem_avail_mib} MiB"
|
|
fi
|
|
fi
|
|
|
|
# Appliance-cap sanity (soft): a cap that EXCEEDS host resources won't protect other guests. WARN,
|
|
# never die — the operator may know better (e.g. capping below a future hardware upgrade).
|
|
if ! $SKIP_PROVISION; then
|
|
if [[ -n "$CPU_CORES" ]]; then
|
|
local host_cores; host_cores=$(nproc 2>/dev/null || echo 0)
|
|
if [[ "${host_cores:-0}" -gt 0 && "$CPU_CORES" -gt "$host_cores" ]]; then
|
|
log_warn " requested cap (${CPU_CORES} cores) exceeds host cores (${host_cores}); the cap won't protect other guests."
|
|
fi
|
|
fi
|
|
if [[ -n "$MEM_MIB" ]]; then
|
|
local host_mem_mib; host_mem_mib=$(awk '/^MemTotal:/{print int($2/1024)}' /proc/meminfo 2>/dev/null || echo 0)
|
|
if [[ "${host_mem_mib:-0}" -gt 0 && "$MEM_MIB" -gt "$host_mem_mib" ]]; then
|
|
log_warn " requested cap (${MEM_MIB} MiB) exceeds host RAM (~${host_mem_mib} MiB); the cap won't protect other guests."
|
|
fi
|
|
fi
|
|
fi
|
|
|
|
# Existing-guest context (GL-2, both modes): cheap situational awareness — on a byo/shared box
|
|
# this is the population the caps + scoped ACL protect.
|
|
local _guest_count; _guest_count=$( { used_vmids || true; } | wc -l)
|
|
log_info " existing guests on this host: ${_guest_count} (pct+qm)"
|
|
|
|
# archive-storage-exists guard (provision only — the golden lives there + the restore reads it).
|
|
if ! $SKIP_PROVISION; then
|
|
if pvesm status --storage "$ARCHIVE_STORAGE" >/dev/null 2>&1; then
|
|
log_info " archive storage '$ARCHIVE_STORAGE' present"
|
|
else
|
|
die "archive storage '$ARCHIVE_STORAGE' not found (pvesm status). Pass --archive-storage NAME."
|
|
fi
|
|
fi
|
|
|
|
# --acl-storages existence (GL-2): the scoped ACL grants Datastore.* at each of these paths — a
|
|
# grant on a storage this box doesn't have silently protects nothing (campaign-2 R2 class). byo
|
|
# refuses; appliance warns. EXCEPTION (v1.15.0, drill retrofit-ACL lesson): `felhom-pbs` is
|
|
# EXPECTED absent on a fresh box — the PBS-DR tier apply creates it later, and the grant MUST
|
|
# be pre-positioned or the apply-bridge 403s (Datastore.Allocate on /storage/felhom-pbs — the
|
|
# drill hit exactly this after narrowing --acl-storages). Do NOT narrow the set to silence
|
|
# the info line.
|
|
local _missing_storages=() _acls
|
|
for _acls in "${PVE_STORAGES[@]}"; do
|
|
if ! pvesm status --storage "$_acls" >/dev/null 2>&1; then
|
|
if [[ "$_acls" == "felhom-pbs" ]]; then
|
|
log_info " acl storage 'felhom-pbs' not present yet — expected: the PBS-DR tier creates it; the grant is pre-positioned deliberately"
|
|
else
|
|
_missing_storages+=("$_acls")
|
|
fi
|
|
fi
|
|
done
|
|
if [[ ${#_missing_storages[@]} -gt 0 ]]; then
|
|
if [[ "$MODE" == "byo" ]]; then
|
|
die "acl storage(s) not found on this box: ${_missing_storages[*]} — pass --acl-storages with this box's real storages"
|
|
fi
|
|
for _acls in "${_missing_storages[@]}"; do
|
|
log_warn " acl storage '$_acls' not found (pvesm status) — its grant will point at nothing"
|
|
done
|
|
elif [[ "$MODE" == "byo" ]]; then
|
|
log_info " acl storages all present or pre-positioned: ${PVE_STORAGES[*]}"
|
|
fi
|
|
|
|
# Host DNS :53 gate (byo only) — GL-8/F6. In byo the agent's own lan-resolver stays OFF (a config
|
|
# assert enforces it), so ANY process bound to :53 is the OWNER's — and Felhom needs the guest
|
|
# reachable by name on the LAN. Felhom must NOT stop/mask/kill a service on a host it does not own
|
|
# (the dnsmasq analog of the break-glass rule), so this REFUSES with the remediation instead of
|
|
# mutating anything. `ss` failure/absence degrades to a warn (never block on a missing tool).
|
|
# Appliance mode is untouched — there the agent's lan_resolver owns :53 by design.
|
|
if [[ "$MODE" == "byo" ]]; then
|
|
local _dns53
|
|
if _dns53=$(ss -H -ltnup 'sport = :53' 2>/dev/null); then
|
|
if [[ -n "$_dns53" ]]; then
|
|
log_error " a resolver is already bound to :53 on this host:"
|
|
echo "$_dns53" | tr -s '[:space:]' ' ' | cut -c1-200 | sed 's/^/ /' >&2
|
|
die "a resolver is already bound to :53 on this host — Felhom needs the guest reachable by name on your LAN.
|
|
Stop or reconfigure that resolver, OR point your LAN DNS at the guest's address, then re-run.
|
|
(Felhom does NOT touch DNS services on a host it does not own — this is a refusal, not a change.)"
|
|
fi
|
|
log_info " host DNS (:53): free"
|
|
else
|
|
log_warn " host DNS (:53): could not query (ss failed/absent) — skipping the :53 gate"
|
|
fi
|
|
fi
|
|
|
|
# hub reachable
|
|
local hc; hc=$(http_code "$HUB_URL/api/v1/config/$CUSTOMER_ID" -H "X-Retrieval-Password: preflight-no-op" || echo 000)
|
|
[[ "$hc" != "000" ]] || die "hub unreachable at $HUB_URL"
|
|
log_info " hub reachable ($HUB_URL)"
|
|
|
|
# customer + passphrase valid (read-only GET /config/{id})
|
|
read_passphrase
|
|
local code; code=$(http_code "$HUB_URL/api/v1/config/$CUSTOMER_ID" -H "X-Retrieval-Password: $PASSPHRASE")
|
|
case "$code" in
|
|
200) log_success " customer '$CUSTOMER_ID' exists + passphrase valid" ;;
|
|
401) die "passphrase REJECTED (401) for customer '$CUSTOMER_ID'" ;;
|
|
404) die "customer '$CUSTOMER_ID' not found in hub (404) — create it in the hub first" ;;
|
|
*) die "unexpected hub status $code on config preflight" ;;
|
|
esac
|
|
|
|
# golden archive — auto-discover a LOCAL one for info; the golden step (7/8) ensures one exists
|
|
# (local else Gitea-fetched + verified), so a missing local golden is no longer fatal here.
|
|
if [[ -z "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then
|
|
GOLDEN_VOLID=$(pvesm list "$ARCHIVE_STORAGE" --content backup 2>/dev/null | awk -v v="$GOLDEN_VMID" '$0 ~ ("vzdump-lxc-" v "-"){print $1}' | sort | tail -1)
|
|
fi
|
|
if [[ -n "$GOLDEN_VOLID" ]]; then
|
|
pvesm list "$ARCHIVE_STORAGE" --content backup 2>/dev/null | grep -q "$(basename "$GOLDEN_VOLID")" || die "golden volid not resolvable: $GOLDEN_VOLID"
|
|
log_info " golden (local): $GOLDEN_VOLID"
|
|
else
|
|
log_info " golden: none local — will fetch + verify from Gitea in step 7/8"
|
|
fi
|
|
|
|
# vmid guard (irrelevant when --skip-provision: we never touch a guest). "In use" is checked
|
|
# against the pct+qm id-set (LXC and VMs share the space), not just `pct status`.
|
|
if $SKIP_PROVISION; then
|
|
log_info " --skip-provision: agent install/config only, no guest will be provisioned"
|
|
elif _vmid_in_use "$VMID"; then
|
|
if $VMID_EXPLICIT; then
|
|
# Explicit --vmid stays deterministic: die unless --force (which over-provisions, destructive).
|
|
if $FORCE; then
|
|
log_warn " vmid $VMID already exists — --force given, it WILL be destroyed by provision"
|
|
else
|
|
die "vmid $VMID already exists. Refusing to clobber a live guest. Pass --force to provision over it."
|
|
fi
|
|
elif $FORCE; then
|
|
# Default vmid + --force: honor the destructive over-provision without prompting.
|
|
log_warn " vmid $VMID already exists — --force given, it WILL be destroyed by provision"
|
|
else
|
|
# Default vmid in use, no --force: auto-pick the next free id and CONFIRM (never silent).
|
|
local free_vmid; free_vmid=$(next_free_vmid "$VMID")
|
|
log_info " vmid $VMID is in use; next free vmid is $free_vmid"
|
|
local ans; read -rp "VMID $VMID is in use. Use next free VMID $free_vmid? [y/N] " ans < /dev/tty
|
|
[[ "$ans" == "y" || "$ans" == "Y" ]] || die "no free vmid confirmed"
|
|
VMID="$free_vmid"
|
|
log_success " using auto-selected vmid $VMID"
|
|
fi
|
|
fi
|
|
# BYO disclosure + typed ack — last, so every disclosed value (vmid incl. auto-pick) is final.
|
|
if [[ "$MODE" == "byo" ]]; then
|
|
_byo_disclosure_ack
|
|
fi
|
|
|
|
# Record the customer + install profile into the state (foundation for a later automatic
|
|
# --uninstall; the mode key backs the C4 resume-mismatch refusal). --preflight-only writes NO
|
|
# state — it must stay re-runnable and must never let a later --resume skip the real preflight.
|
|
if ! $PREFLIGHT_ONLY; then
|
|
_state_put customer_id "$CUSTOMER_ID"
|
|
_state_put mode "$MODE"
|
|
_state_mark preflight
|
|
fi
|
|
log_success "pre-flight passed"
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# STEP 2 — Proxmox API token (idempotent pveum; reuse-if-working else rotate)
|
|
#-------------------------------------------------------------------------------
|
|
step_token() {
|
|
log_step "2/8 Proxmox API token"
|
|
# Fast-path only when the token is ALREADY in memory this process (avoids a double reuse-check if
|
|
# called twice). On --resume PVE_TOKEN starts empty, so this correctly re-derives it from the
|
|
# on-disk config (reuse) or a rotation — GL6-F4. Uses _state_has (no misleading SKIP log here).
|
|
if _state_has token && [[ -n "$PVE_TOKEN" ]]; then return 0; fi
|
|
|
|
# Pool BEFORE the ACL: /pool/felhom must exist before apply_scoped_acl grants on it (3b). Always —
|
|
# even under --skip-provision (the token exists now; a later provision-into-pool needs pool + grant).
|
|
# The 3 scoped roles + grants are created by apply_scoped_acl below (AFTER the token exists).
|
|
ensure_felhom_pool
|
|
|
|
# user: tolerate-exists
|
|
if pveum user list --output-format json 2>/dev/null | python3 -c "import json,sys;sys.exit(0 if any(u['userid']=='$PVE_USER' for u in json.load(sys.stdin)) else 1)"; then
|
|
log_info " user $PVE_USER exists"
|
|
else
|
|
run pveum user add "$PVE_USER"
|
|
fi
|
|
|
|
# token: reuse if the existing agent config token still authenticates, else rotate
|
|
local reused=false
|
|
if [[ -f "$AGENT_CONFIG" ]] && python3 -c "import json,sys;d=json.load(open('$AGENT_CONFIG'));sys.exit(0 if d.get('proxmox',{}).get('token') else 1)" 2>/dev/null; then
|
|
log_info " existing agent config has a token — testing it (read-only --selftest)"
|
|
if felhom-agent --config "$AGENT_CONFIG" --selftest >/dev/null 2>&1; then
|
|
log_success " existing token authenticates — REUSING (no rotation)"
|
|
PVE_TOKEN=$(python3 -c "import json;print(json.load(open('$AGENT_CONFIG'))['proxmox']['token'])")
|
|
reused=true
|
|
else
|
|
log_warn " existing token failed selftest — will rotate"
|
|
fi
|
|
fi
|
|
|
|
if ! $reused; then
|
|
if $DRY_RUN; then
|
|
log_dry "pveum user token remove $PVE_USER $PVE_TOKENID # if present"
|
|
log_dry "pveum user token add $PVE_USER $PVE_TOKENID --privsep 1 --output-format json # capture .value"
|
|
PVE_TOKEN="${PVE_USER}!${PVE_TOKENID}=<DRY-RUN-SECRET>"
|
|
else
|
|
if pveum user token list "$PVE_USER" --output-format json 2>/dev/null | python3 -c "import json,sys;sys.exit(0 if any(t['tokenid']=='$PVE_TOKENID' for t in json.load(sys.stdin)) else 1)"; then
|
|
log_info " removing stale token $PVE_TOKENID (secret unrecoverable — rotating)"
|
|
pveum user token remove "$PVE_USER" "$PVE_TOKENID"
|
|
fi
|
|
local secret
|
|
secret=$(pveum user token add "$PVE_USER" "$PVE_TOKENID" --privsep 1 --output-format json | python3 -c "import json,sys;print(json.load(sys.stdin)['value'])")
|
|
[[ -n "$secret" ]] || die "failed to capture new token secret"
|
|
PVE_TOKEN="${PVE_USER}!${PVE_TOKENID}=${secret}"
|
|
log_success " token minted (secret captured, not logged)"
|
|
fi
|
|
fi
|
|
|
|
# Scoped ACL grants — AFTER the token exists (`pveum user token remove` purges the token's ACL, so
|
|
# re-applying post-rotate is mandatory; `acl modify` is idempotent so this is safe on the reuse path).
|
|
apply_scoped_acl
|
|
# If this box previously ran the pre-3b broad grant (re-install/upgrade), remove it — else the old
|
|
# FelhomAgent role at / would UNION with the scoped grant and defeat containment. Tolerate-absent.
|
|
remove_old_broad_acl
|
|
_state_mark token
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# STEP 3 — compute grows (floors) if not passed
|
|
#-------------------------------------------------------------------------------
|
|
step_grows() {
|
|
log_step "3/8 compute volume grows"
|
|
# Golden base since build-golden.sh v3.0.0 (R-165): rootfs 32G + ONE data volume 24G. The separate
|
|
# 8G user-data volume was MERGED AWAY — one volume, one free-space figure, no ceiling — so there is
|
|
# one number to compute here instead of two.
|
|
#
|
|
# THE SIZE IS DERIVED FROM THE PHYSICAL DISK, which is what makes the merge safe to ship: an
|
|
# unflagged install does NOT get the golden's 24G, it gets a share of the thin pool's real free
|
|
# space. (Before R-165 this same block already did the deriving; the merge only collapsed its
|
|
# 80/20 docker-vs-sysdata split into a single total.)
|
|
if [[ -z "$ROOTFS_GROW$DATAVOL_GROW$SYSDATA_GROW" ]]; then
|
|
local free_gib
|
|
free_gib=$(lvs --noheadings --units g -o lv_size,data_percent /dev/pve/data 2>/dev/null | awk '{gsub(/[^0-9.]/,"",$1); used=$2; print int($1*(100-used)/100)}' 2>/dev/null || echo 0)
|
|
# Reserve headroom; the totals below are the pre-merge pair SUMMED, so an appliance gets the
|
|
# same capacity it did before — it is simply no longer split by a wall.
|
|
ROOTFS_GROW=0
|
|
if [[ "${free_gib:-0}" -ge 300 ]]; then
|
|
DATAVOL_GROW=226 # 184+42 -> the standard 250G appliance (was 200G+50G)
|
|
elif [[ "${free_gib:-0}" -ge 150 ]]; then
|
|
DATAVOL_GROW=106 # 84+22
|
|
else
|
|
DATAVOL_GROW=46 # 34+12 — minimal floor
|
|
fi
|
|
SYSDATA_GROW=0
|
|
log_info " auto-computed from ~${free_gib} GiB free (ONE volume since R-165)"
|
|
fi
|
|
ROOTFS_GROW="${ROOTFS_GROW:-0}"; DATAVOL_GROW="${DATAVOL_GROW:-0}"; SYSDATA_GROW="${SYSDATA_GROW:-0}"
|
|
# A hand-passed --sysdata-grow is still ACCEPTED and still counts: the agent folds it into the one
|
|
# volume (bringup.go 4b), so an operator reproducing an old command line gets the same total.
|
|
log_info " grows: rootfs +${ROOTFS_GROW}G (->$((32+ROOTFS_GROW))G), data +$((DATAVOL_GROW+SYSDATA_GROW))G (->$((24+DATAVOL_GROW+SYSDATA_GROW))G, ONE volume)"
|
|
_state_mark grows
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# STEP 4 — host enroll (option C; single secret, no global key)
|
|
#-------------------------------------------------------------------------------
|
|
step_enroll() {
|
|
log_step "4/8 host enrollment (POST /host-enroll)"
|
|
if $DRY_RUN; then
|
|
log_dry "curl -fsS -X POST $HUB_URL/api/v1/host-enroll -H 'X-Retrieval-Password: <pass>' -d '{\"customer_id\":\"$CUSTOMER_ID\"}'"
|
|
HOST_ID="<dry-run-host-id>"; HOST_API_KEY="<dry-run-key>"; _state_mark enroll; return 0
|
|
fi
|
|
local resp code body
|
|
resp=$(curl -sS -w $'\n%{http_code}' -X POST "$HUB_URL/api/v1/host-enroll" \
|
|
-H "X-Retrieval-Password: $PASSPHRASE" -H 'Content-Type: application/json' \
|
|
-d "{\"customer_id\":\"$CUSTOMER_ID\"}")
|
|
code=$(tail -n1 <<<"$resp"); body=$(sed '$d' <<<"$resp")
|
|
case "$code" in
|
|
201) log_success " host MINTED (first enroll)" ;;
|
|
200) log_success " host REUSED (idempotent — existing credential)" ;;
|
|
401) die "host-enroll 401 (passphrase) — should have been caught in preflight" ;;
|
|
404) die "host-enroll 404 (unknown customer)" ;;
|
|
*) die "host-enroll unexpected $code: $body" ;;
|
|
esac
|
|
HOST_ID=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['host_id'])" "$body")
|
|
HOST_API_KEY=$(python3 -c "import json,sys;print(json.loads(sys.argv[1])['api_key'])" "$body")
|
|
[[ -n "$HOST_ID" && -n "$HOST_API_KEY" ]] || die "host-enroll: malformed response"
|
|
log_info " host_id: $HOST_ID (api_key captured, not logged)"
|
|
_state_mark enroll
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# STEP 4b — break-glass credential (TASK G1): generate + set + vault the root@pam console password
|
|
#-------------------------------------------------------------------------------
|
|
# The human fallback for when BOTH the sshd path AND the agent-independent auto-heal (layers 1+2) have
|
|
# failed: a strong root@pam password lets the operator reach the PVE WEB CONSOLE (pveproxy :8006 — a
|
|
# failure domain distinct from sshd) and run the one-line /run/sshd fix. Generated with strong entropy,
|
|
# set via chpasswd, and vaulted to the hub over the enroll-authenticated channel (host api_key). The
|
|
# password is NEVER logged, printed, or written to any file — it goes stdin→chpasswd and stdin→curl
|
|
# only (SPIKE-felhom-sshd finding #9 / TASK G1 trap 3). Idempotent: skipped if already vaulted unless
|
|
# --rotate-recovery (a re-set would strand the operator's saved copy).
|
|
step_break_glass() {
|
|
log_step "4b/8 break-glass credential (root@pam console password → hub vault)"
|
|
if [[ -z "${HOST_ID:-}" || -z "${HOST_API_KEY:-}" ]]; then
|
|
log_warn " no host_id/api_key (enroll skipped?) — cannot vault a recovery credential; skipping"
|
|
return 0
|
|
fi
|
|
if $DRY_RUN; then
|
|
log_dry "openssl rand → strong root@pam password (never logged) ; chpasswd ; PUT $HUB_URL/api/v1/hosts/$HOST_ID/recovery-credential (Bearer host key)"
|
|
_state_mark break_glass; return 0
|
|
fi
|
|
if _state_has break_glass && ! $ROTATE_RECOVERY; then
|
|
log_skip " recovery credential already vaulted (use --rotate-recovery to regenerate)"
|
|
return 0
|
|
fi
|
|
# Strong password: 24 url-safe bytes (~144 bits). Kept ONLY in a local shell var, never on disk.
|
|
local newpw
|
|
newpw=$(openssl rand -base64 24 2>/dev/null | tr -d '\n' | tr '+/' '-_')
|
|
[[ ${#newpw} -ge 24 ]] || die "failed to generate a strong recovery password"
|
|
# Set root@pam (= the Linux root user on PVE) via chpasswd on STDIN — no argv, no log.
|
|
if ! printf 'root:%s\n' "$newpw" | chpasswd 2>/dev/null; then
|
|
newpw="" # scrub
|
|
die "chpasswd failed to set the root@pam recovery password"
|
|
fi
|
|
# Vault to the hub over the host-key-authenticated channel; password only on stdin (-d @-).
|
|
local code
|
|
code=$(printf '{"username":"root@pam","password":"%s"}' "$newpw" \
|
|
| curl -sS -o /dev/null -w '%{http_code}' -X PUT \
|
|
"$HUB_URL/api/v1/hosts/$HOST_ID/recovery-credential" \
|
|
-H "Authorization: Bearer $HOST_API_KEY" -H 'Content-Type: application/json' -d @- 2>/dev/null)
|
|
newpw="" # scrub the plaintext from the shell var the moment it is vaulted
|
|
case "$code" in
|
|
200) log_success " root@pam password set + vaulted to the hub (retrieve via the operator /admin path; never logged here)"
|
|
# Drill F-8: say this OUT LOUD — the operator was live-surprised by a 401 at the PVE
|
|
# GUI. Pointer only, never the credential.
|
|
log_warn " NOTE: the root@pam password just CHANGED — the old one now fails at the PVE web GUI (:8006)."
|
|
log_warn " Retrieve the new one at hub → host page (vaulted recovery credential)." ;;
|
|
401|403) die "recovery-credential vault rejected ($code) — host key/authorization problem" ;;
|
|
*) die "recovery-credential vault failed (HTTP $code)" ;;
|
|
esac
|
|
_state_mark break_glass
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# STEP 5 — agent install: fetch+verify the binary, ensure the service user, sudoers, unit
|
|
#-------------------------------------------------------------------------------
|
|
# Closes the old prerequisite "install the agent binary + unit manually". Fetches the binary from
|
|
# Gitea (git token from controller.yaml), VERIFIES its sha256 against the hub manifest, then installs
|
|
# the non-root felhom-agent user + binary + sudoers + unit. The SERVICE is started in step 6 (after the
|
|
# config is written) — here we only install + daemon-reload + enable.
|
|
# Part 2 (v1.25.0): a fresh appliance install boots on the stock PVE ENTERPRISE apt repo with no
|
|
# subscription key, so every `apt-get update` 401s and the box can never pull security updates. Switch
|
|
# the Proxmox-managed repos to no-subscription — idempotent, APPLIANCE-ONLY, and scoped to the repos
|
|
# Proxmox itself ships (pve-enterprise + ceph enterprise, deb822 `.sources` on trixie/PVE9, or the
|
|
# legacy `.list`). The customer's own sources are never touched. No upgrades are run — repo alignment
|
|
# only. Runs BEFORE the first apt operation (the sudo/age installs below) so they don't hit the 401.
|
|
align_apt_repos() {
|
|
[[ "$MODE" == "appliance" ]] || { log_skip " apt repo alignment skipped (byo — the owner manages repos)"; return 0; }
|
|
if $DRY_RUN; then log_dry "disable pve/ceph enterprise .sources (Enabled: no) + legacy .list; add pve-no-subscription.sources; apt-get update"; return 0; fi
|
|
local d=/etc/apt/sources.list.d changed=0 suite="" keyring="" f
|
|
if [[ -f "$d/pve-enterprise.sources" ]]; then
|
|
suite=$(awk -F': *' '/^Suites:/{print $2; exit}' "$d/pve-enterprise.sources")
|
|
keyring=$(awk -F': *' '/^Signed-By:/{print $2; exit}' "$d/pve-enterprise.sources")
|
|
fi
|
|
[[ -n "$suite" ]] || suite="trixie"
|
|
[[ -n "$keyring" ]] || keyring="/usr/share/keyrings/proxmox-archive-keyring.gpg"
|
|
# 1a. disable enterprise deb822 .sources (append `Enabled: no`, idempotent)
|
|
for f in "$d/pve-enterprise.sources" "$d/ceph.sources"; do
|
|
[[ -f "$f" ]] || continue
|
|
grep -qiE '^Enabled:[[:space:]]*no' "$f" && continue
|
|
sed -i '/^[Ee]nabled:/d' "$f"; printf 'Enabled: no\n' >> "$f"; changed=1
|
|
log_info " disabled $(basename "$f") (Enabled: no)"
|
|
done
|
|
# 1b. legacy .list enterprise line → commented
|
|
for f in "$d/pve-enterprise.list"; do
|
|
[[ -f "$f" ]] || continue
|
|
grep -qE '^[[:space:]]*deb[[:space:]].*enterprise\.proxmox\.com' "$f" || continue
|
|
sed -i 's|^[[:space:]]*\(deb[[:space:]].*enterprise\.proxmox\.com.*\)$|# \1 # felhom: no-subscription|' "$f"; changed=1
|
|
log_info " commented enterprise line in $(basename "$f")"
|
|
done
|
|
# 2. add pve-no-subscription if it is not already present in any list/sources
|
|
if ! grep -rqsE 'download\.proxmox\.com/debian/pve([[:space:]/]|$)' "$d"/*.sources "$d"/*.list 2>/dev/null; then
|
|
cat > "$d/pve-no-subscription.sources" <<SRC
|
|
Types: deb
|
|
URIs: http://download.proxmox.com/debian/pve
|
|
Suites: $suite
|
|
Components: pve-no-subscription
|
|
Signed-By: $keyring
|
|
SRC
|
|
changed=1; log_info " added pve-no-subscription.sources (suite=$suite)"
|
|
fi
|
|
# 3. refresh — must be clean now (a stale enterprise line would 401)
|
|
if apt-get update -q >/dev/null 2>&1; then
|
|
log_success " apt repos aligned to no-subscription ($([[ $changed == 1 ]] && echo changed || echo 'already aligned'); apt-get update OK)"
|
|
else
|
|
log_warn " apt-get update non-zero after repo alignment — inspect $d (non-fatal to the install)"
|
|
fi
|
|
}
|
|
|
|
step_agent_install() {
|
|
log_step "5/8 agent install (fetch + verify + install)"
|
|
align_apt_repos # Part 2: switch enterprise→no-subscription before any apt-get (appliance only)
|
|
|
|
# Manifest + git fetch credential (both passphrase / config-retrieve — NO new credential).
|
|
resolve_artifacts
|
|
resolve_git_creds
|
|
[[ -n "$ART_AGENT_VER" ]] || die "hub artifact manifest has no agent version — set it in the operator UI (Configs → Day-0 artifacts)"
|
|
log_info " manifest: agent v$ART_AGENT_VER (sha ${ART_AGENT_SHA:0:16}…), golden v${ART_GOLDEN_VER:-<unset>}"
|
|
|
|
# Idempotent skip: same version already installed AND the service is healthy.
|
|
local cur=""
|
|
[[ -x "$AGENT_BIN" ]] && cur=$("$AGENT_BIN" --version 2>/dev/null | awk '{print $2}')
|
|
if [[ "$cur" == "$ART_AGENT_VER" ]] && systemctl is-active --quiet felhom-agent 2>/dev/null; then
|
|
log_skip " agent v$cur already installed + service active — skipping binary install"
|
|
else
|
|
local url="$GITEA_BASE/api/packages/$GITEA_OWNER/generic/felhom-agent/$ART_AGENT_VER/felhom-agent"
|
|
log_info " fetching agent binary v$ART_AGENT_VER from Gitea …"
|
|
if $DRY_RUN; then
|
|
log_dry "$(_dry_auth_hint) -o /tmp/felhom-agent.new $url ; verify sha256=$ART_AGENT_SHA ; install -m0755 -> $AGENT_BIN"
|
|
else
|
|
local tmp; tmp=$(mktemp -t felhom-agent.XXXXXX)
|
|
fetch_verify "$url" "$tmp" "$ART_AGENT_SHA"
|
|
# back up any existing binary before replacing
|
|
if [[ -f "$AGENT_BIN" ]]; then
|
|
cp -a "$AGENT_BIN" "${AGENT_BIN}.bak-$(date +%s)" 2>/dev/null || true
|
|
fi
|
|
install -m 0755 -o root -g root "$tmp" "$AGENT_BIN"
|
|
rm -f "$tmp"
|
|
log_success " installed $AGENT_BIN ($("$AGENT_BIN" --version 2>&1 | head -1))"
|
|
fi
|
|
fi
|
|
|
|
# The non-root model REQUIRES the `sudo` package (provides both `sudo` and `visudo`). A host that
|
|
# previously ran the agent as root+`direct` won't have it installed. Install it idempotently before
|
|
# the sudoers (visudo validates it) and before the daemon starts (it shells out via `sudo -n`).
|
|
if ! command -v sudo >/dev/null 2>&1 || ! command -v visudo >/dev/null 2>&1; then
|
|
if $DRY_RUN; then
|
|
log_dry "apt-get install -y sudo # required for the non-root agent (provides sudo + visudo)"
|
|
else
|
|
log_info " installing the 'sudo' package (required for the non-root agent model) …"
|
|
DEBIAN_FRONTEND=noninteractive apt-get install -y -q sudo >/dev/null 2>&1 \
|
|
|| { apt-get update -q >/dev/null 2>&1; DEBIAN_FRONTEND=noninteractive apt-get install -y -q sudo >/dev/null 2>&1; } \
|
|
|| die "failed to install the 'sudo' package (needed for the non-root agent)"
|
|
log_success " sudo installed ($(sudo --version 2>/dev/null | head -1))"
|
|
fi
|
|
fi
|
|
# `age` (drill F-10) — the escrow ceremony's identity-wrap dependency. NOT installed by PVE by
|
|
# default; a fresh-box ceremony died on it live (the demo host only had it from its spike-era
|
|
# install). DR-tier-by-default: baked on every install, both modes. Idempotent.
|
|
if ! command -v age >/dev/null 2>&1; then
|
|
if $DRY_RUN; then
|
|
log_dry "apt-get install -y age # escrow ceremony identity-wrap dependency (F-10)"
|
|
else
|
|
log_info " installing the 'age' package (escrow ceremony identity-wrap dependency) …"
|
|
DEBIAN_FRONTEND=noninteractive apt-get install -y -q age >/dev/null 2>&1 \
|
|
|| { apt-get update -q >/dev/null 2>&1; DEBIAN_FRONTEND=noninteractive apt-get install -y -q age >/dev/null 2>&1; } \
|
|
|| die "failed to install the 'age' package (the escrow ceremony needs it)"
|
|
log_success " age installed ($(age --version 2>/dev/null | head -1))"
|
|
fi
|
|
fi
|
|
|
|
# Resolve visudo by absolute path too (non-login SSH PATH can miss /usr/sbin).
|
|
local VISUDO; VISUDO=$(command -v visudo 2>/dev/null || echo /usr/sbin/visudo)
|
|
|
|
# Service user (system, no login, no home dir creation needed beyond state).
|
|
if $DRY_RUN; then
|
|
log_dry "useradd --system --no-create-home --shell /usr/sbin/nologin $AGENT_USER # if absent"
|
|
elif id "$AGENT_USER" >/dev/null 2>&1; then
|
|
log_info " service user $AGENT_USER exists"
|
|
else
|
|
useradd --system --no-create-home --shell /usr/sbin/nologin "$AGENT_USER"
|
|
log_success " created service user $AGENT_USER"
|
|
fi
|
|
|
|
# systemd-journal group: the NAS verify pipeline (agent v0.81.0) classifies mount failures from
|
|
# an UNPRIVILEGED `journalctl -u <unit>` read — group membership instead of a sudoers grant
|
|
# (rule: no journalctl via sudo). Idempotent. Existing installs: run the usermod by hand, then
|
|
# `systemctl restart felhom-agent`.
|
|
if $DRY_RUN; then
|
|
log_dry "usermod -aG systemd-journal $AGENT_USER"
|
|
elif id -nG "$AGENT_USER" | tr ' ' '\n' | grep -qx systemd-journal; then
|
|
log_info " $AGENT_USER already in systemd-journal"
|
|
else
|
|
usermod -aG systemd-journal "$AGENT_USER"
|
|
log_success " added $AGENT_USER to systemd-journal (unprivileged journal read for NAS verify)"
|
|
fi
|
|
|
|
# State dir (the old root deployment may have created it root-owned; StateDirectory= also adjusts
|
|
# on start, but chown here so the very first start has a writable dir).
|
|
run mkdir -p "$AGENT_STATE_DIR"
|
|
run chown -R "${AGENT_USER}:${AGENT_USER}" "$AGENT_STATE_DIR"
|
|
run chmod 0750 "$AGENT_STATE_DIR"
|
|
|
|
# ── Agent local-API leaf lifecycle (B.2) ──────────────────────────────────────────────────────
|
|
# The leaf's SHA-256 is pinned into EVERY guest's bootstrap. A reinstall that REGENERATES the leaf
|
|
# invalidates every controller's pin (the 2026-06-28 root→non-root incident → controller↔agent dead
|
|
# for days). Two protections:
|
|
# (a) --preserve-state-from DIR: carry the prior leaf+key+token-store over → the fp stays STABLE,
|
|
# no re-bootstrap needed. (Distinct from --preserve-from, which merges config sections only.)
|
|
# (b) populated-host guard: REFUSE to proceed leaf-less on a host that already has guests, unless
|
|
# --preserve-state-from or an explicit --allow-new-leaf is given. Converts the silent footgun
|
|
# into a hard stop.
|
|
local _have_leaf=false
|
|
[[ -f "$AGENT_STATE_DIR/local-api.crt" && -f "$AGENT_STATE_DIR/local-api.key" ]] && _have_leaf=true
|
|
|
|
if [[ -n "$PRESERVE_STATE_FROM" ]]; then
|
|
[[ -f "$PRESERVE_STATE_FROM/local-api.crt" && -f "$PRESERVE_STATE_FROM/local-api.key" ]] \
|
|
|| die "--preserve-state-from $PRESERVE_STATE_FROM: local-api.crt/key not found there"
|
|
openssl x509 -in "$PRESERVE_STATE_FROM/local-api.crt" -noout >/dev/null 2>&1 \
|
|
|| die "--preserve-state-from: $PRESERVE_STATE_FROM/local-api.crt does not parse as an X.509 cert — refusing"
|
|
if $DRY_RUN; then
|
|
log_dry "preserve agent state: copy local-api.{crt,key}+local-tokens.log from $PRESERVE_STATE_FROM -> $AGENT_STATE_DIR (chown $AGENT_USER; 644/600/600)"
|
|
else
|
|
install -o "$AGENT_USER" -g "$AGENT_USER" -m 0644 "$PRESERVE_STATE_FROM/local-api.crt" "$AGENT_STATE_DIR/local-api.crt"
|
|
install -o "$AGENT_USER" -g "$AGENT_USER" -m 0600 "$PRESERVE_STATE_FROM/local-api.key" "$AGENT_STATE_DIR/local-api.key"
|
|
[[ -f "$PRESERVE_STATE_FROM/local-tokens.log" ]] && \
|
|
install -o "$AGENT_USER" -g "$AGENT_USER" -m 0600 "$PRESERVE_STATE_FROM/local-tokens.log" "$AGENT_STATE_DIR/local-tokens.log"
|
|
log_success " preserved agent leaf+token store from $PRESERVE_STATE_FROM (pin stays stable — no re-bootstrap)"
|
|
fi
|
|
_have_leaf=true
|
|
fi
|
|
|
|
# GL-8/F3: the guard fires only when a FELHOM guest is present — regenerating the leaf orphans
|
|
# THAT guest's pinned fingerprint (the 2026-06-28 incident). A BYO host's own non-Felhom guests
|
|
# have no Felhom pin to break, so `felhom_guests` (the /etc/felhom-bootstrap detector) — NOT
|
|
# `pct list` (any guest) — is the correct populated-host test: a fresh byo install on a host that
|
|
# only runs the owner's guests now proceeds without `--allow-new-leaf`.
|
|
if ! $_have_leaf && ! $ALLOW_NEW_LEAF; then
|
|
if [[ -n "$(felhom_guests)" ]]; then
|
|
die "this host already has a Felhom guest but $AGENT_STATE_DIR has no agent leaf to preserve.
|
|
Re-running here will REGENERATE the leaf and invalidate every controller's pin (the 2026-06-28 incident).
|
|
Pass --preserve-state-from <old state dir> to keep the pin stable, or --allow-new-leaf to regenerate
|
|
intentionally (every guest must then be re-bootstrapped)."
|
|
fi
|
|
fi
|
|
|
|
# Guarded-mkfs wrapper (Impl-1 Part B) — the ONLY mkfs path the sudoers permits. Install it BEFORE
|
|
# the sudoers (which allowlists it), 0755 root:root under /usr/local/sbin. bash -n before install.
|
|
if $DRY_RUN; then
|
|
log_dry "fetch configs/felhom-mkfs-guarded.sh ; bash -n ; install 0755 -> /usr/local/sbin/felhom-mkfs-guarded"
|
|
else
|
|
local wtmp; wtmp=$(mktemp -t felhom-mkfs.XXXXXX)
|
|
fetch_raw "configs/felhom-mkfs-guarded.sh" "$wtmp"
|
|
bash -n "$wtmp" || { rm -f "$wtmp"; die "fetched felhom-mkfs-guarded.sh failed bash -n — refusing to install"; }
|
|
install -m 0755 -o root -g root "$wtmp" /usr/local/sbin/felhom-mkfs-guarded
|
|
rm -f "$wtmp"
|
|
log_success " installed /usr/local/sbin/felhom-mkfs-guarded (0755, the guarded mkfs path)"
|
|
fi
|
|
|
|
# Self-update guarded wrapper (TASK D1) — the ONLY binary-swap path the sudoers permits. Install
|
|
# it BEFORE the sudoers (which allowlists it), 0755 root:root under /usr/local/sbin. sh -n first.
|
|
if $DRY_RUN; then
|
|
log_dry "fetch configs/felhom-selfupdate-guarded ; sh -n ; install 0755 -> /usr/local/sbin/felhom-selfupdate-guarded"
|
|
else
|
|
local sutmp; sutmp=$(mktemp -t felhom-selfupd.XXXXXX)
|
|
fetch_raw "configs/felhom-selfupdate-guarded" "$sutmp"
|
|
sh -n "$sutmp" || { rm -f "$sutmp"; die "fetched felhom-selfupdate-guarded failed sh -n — refusing to install"; }
|
|
install -m 0755 -o root -g root "$sutmp" /usr/local/sbin/felhom-selfupdate-guarded
|
|
rm -f "$sutmp"
|
|
log_success " installed /usr/local/sbin/felhom-selfupdate-guarded (0755, the guarded A/B binary-swap path)"
|
|
fi
|
|
|
|
# PBS-DR apply wrapper (drill F-7) — the ONLY pbs-apply path the FELHOM_PBSDR sudoers alias
|
|
# permits. The sudoers has shipped the alias since the PBS-DR slice; the binary itself did not
|
|
# ship, so the 3 pbsdr-* capabilities were born DEGRADED ("binary not found") on every fresh
|
|
# box. DR-tier-by-default (2026-07-12): capability is baked on EVERY install (activation is a
|
|
# hub flag). Install BEFORE the sudoers, 0755 root:root, bash -n first — the mkfs pattern.
|
|
if $DRY_RUN; then
|
|
log_dry "fetch configs/felhom-pbs-apply ; bash -n ; install 0755 -> /usr/local/sbin/felhom-pbs-apply"
|
|
else
|
|
local patmp; patmp=$(mktemp -t felhom-pbsapply.XXXXXX)
|
|
fetch_raw "configs/felhom-pbs-apply" "$patmp"
|
|
bash -n "$patmp" || { rm -f "$patmp"; die "fetched felhom-pbs-apply failed bash -n — refusing to install"; }
|
|
install -m 0755 -o root -g root "$patmp" /usr/local/sbin/felhom-pbs-apply
|
|
rm -f "$patmp"
|
|
log_success " installed /usr/local/sbin/felhom-pbs-apply (0755, the guarded PBS-DR apply path)"
|
|
fi
|
|
|
|
# E-2a: backup-target apply wrapper — the ONLY path the FELHOM_BACKUPTARGET sudoers alias permits.
|
|
# The agent cannot create a PVE storage (Datastore.Allocate at /storage) or grant an ACL
|
|
# (Permissions.Modify); both are deliberately outside its role, so the move runs through this
|
|
# fenced shim. Same pattern as pbs-apply: install BEFORE the sudoers, 0755 root:root, bash -n first.
|
|
if $DRY_RUN; then
|
|
log_dry "fetch configs/felhom-backup-target-apply ; bash -n ; install 0755 -> /usr/local/sbin/felhom-backup-target-apply"
|
|
else
|
|
local bttmp; bttmp=$(mktemp -t felhom-bt.XXXXXX)
|
|
fetch_raw "configs/felhom-backup-target-apply" "$bttmp"
|
|
bash -n "$bttmp" || { rm -f "$bttmp"; die "fetched felhom-backup-target-apply failed bash -n — refusing to install"; }
|
|
install -m 0755 -o root -g root "$bttmp" /usr/local/sbin/felhom-backup-target-apply
|
|
rm -f "$bttmp"
|
|
log_success " installed /usr/local/sbin/felhom-backup-target-apply (0755, the guarded backup-target path)"
|
|
fi
|
|
|
|
# Sudoers — fetch the canonical file, validate with visudo -cf BEFORE installing (0440 root:root).
|
|
if $DRY_RUN; then
|
|
log_dry "fetch configs/felhom-agent.sudoers ; visudo -cf ; install 0440 -> $AGENT_SUDOERS"
|
|
else
|
|
local sdtmp; sdtmp=$(mktemp -t felhom-sudoers.XXXXXX)
|
|
fetch_raw "configs/felhom-agent.sudoers" "$sdtmp"
|
|
"$VISUDO" -cf "$sdtmp" >/dev/null || { rm -f "$sdtmp"; die "fetched sudoers failed visudo -cf — refusing to install"; }
|
|
install -m 0440 -o root -g root "$sdtmp" "$AGENT_SUDOERS"
|
|
rm -f "$sdtmp"
|
|
# re-validate the live drop-in in the full sudoers context
|
|
"$VISUDO" -cf /etc/sudoers >/dev/null || die "sudoers invalid after installing $AGENT_SUDOERS"
|
|
log_success " installed $AGENT_SUDOERS (0440, visudo-validated)"
|
|
fi
|
|
|
|
# systemd unit — fetch the canonical unit, install, daemon-reload, enable (NOT start — no config yet).
|
|
if $DRY_RUN; then
|
|
log_dry "fetch configs/felhom-agent.service -> $AGENT_UNIT ; systemctl daemon-reload ; systemctl enable felhom-agent"
|
|
else
|
|
local untmp; untmp=$(mktemp -t felhom-unit.XXXXXX)
|
|
fetch_raw "configs/felhom-agent.service" "$untmp"
|
|
grep -q "User=$AGENT_USER" "$untmp" || { rm -f "$untmp"; die "fetched unit does not run as $AGENT_USER — refusing"; }
|
|
if [[ -f "$AGENT_UNIT" ]]; then cp -a "$AGENT_UNIT" "${AGENT_UNIT}.bak-$(date +%s)" 2>/dev/null || true; fi
|
|
install -m 0644 -o root -g root "$untmp" "$AGENT_UNIT"
|
|
rm -f "$untmp"
|
|
systemctl daemon-reload
|
|
systemctl enable felhom-agent >/dev/null 2>&1 || true
|
|
log_success " installed $AGENT_UNIT + enabled (started in step 6 after config)"
|
|
fi
|
|
|
|
# Self-update rollback unit + start-limit drop-in (TASK D1). The drop-in gives felhom-agent the
|
|
# tuned start-limit (SPIKE-agent-selfupdate: [Unit] StartLimitIntervalSec=120 + Burst=4 → a
|
|
# crash-looping update reaches terminal `failed` in ~20s instead of looping forever) and the
|
|
# OnFailure= that auto-reverts. Both are idempotent installs (+ daemon-reload). Non-fatal on a
|
|
# box whose agent repo predates them (the raw fetch would 404) — self-update just stays manual.
|
|
if $DRY_RUN; then
|
|
log_dry "fetch configs/felhom-agent-rollback.service -> /etc/systemd/system/ ; fetch configs/felhom-agent-limits.conf -> $AGENT_UNIT.d/ ; daemon-reload"
|
|
else
|
|
local rbtmp; rbtmp=$(mktemp -t felhom-rollback.XXXXXX)
|
|
if fetch_raw "configs/felhom-agent-rollback.service" "$rbtmp" 2>/dev/null; then
|
|
install -m 0644 -o root -g root "$rbtmp" /etc/systemd/system/felhom-agent-rollback.service
|
|
local ddir="${AGENT_UNIT}.d"; mkdir -p "$ddir"
|
|
local lmtmp; lmtmp=$(mktemp -t felhom-limits.XXXXXX)
|
|
fetch_raw "configs/felhom-agent-limits.conf" "$lmtmp"
|
|
# [SF-3] the keys MUST be in [Unit] — refuse a drop-in that would half-apply in [Service].
|
|
grep -q '^\[Unit\]' "$lmtmp" || { rm -f "$lmtmp" "$rbtmp"; die "self-update limits drop-in missing [Unit] section — refusing"; }
|
|
install -m 0644 -o root -g root "$lmtmp" "$ddir/felhom-agent-limits.conf"
|
|
rm -f "$lmtmp"
|
|
systemctl daemon-reload
|
|
log_success " installed self-update rollback unit + start-limit drop-in (auto-rollback armed)"
|
|
else
|
|
log_skip " self-update rollback artifacts not in the agent repo yet — self-update stays manual"
|
|
fi
|
|
rm -f "$rbtmp"
|
|
fi
|
|
|
|
# Management-plane break-glass layers 1+2 (TASK G1). Three artifacts that keep the host reachable
|
|
# even if a second sshd (H1) removes the SHARED /run/sshd privsep dir (SPIKE-felhom-sshd §8):
|
|
# • felhom-privsep.conf (tmpfiles) — layer 1: /run/sshd is boot-persistent, owned by no unit.
|
|
# • felhom-mgmt-watchdog (script) — layer 2 heal action (recreate dir + reset-failed sshd).
|
|
# • .service + .timer — run it every ~60s, AGENT-INDEPENDENTLY (heals with the
|
|
# agent down — the whole point; trap 1).
|
|
# Non-fatal if the agent repo predates them (raw fetch 404s → break-glass just stays manual).
|
|
# HARD GUARD: refuse ANY fetched unit that declares RuntimeDirectory= — that directive is the very
|
|
# incident G1 closes (a second sshd's `RuntimeDirectory=sshd` removed the shared /run/sshd).
|
|
install_mgmt_watchdog
|
|
|
|
# H1: dedicated felhom-sshd OOB instance + static belt (appliance default since v1.25.0; --no-oob opts out).
|
|
install_oob
|
|
|
|
_state_mark agent_install
|
|
}
|
|
|
|
# install_mgmt_watchdog fetches + installs the G1 break-glass host artifacts (idempotent; enables the
|
|
# timer). Split out for readability; called from step_agent_install. Every unit is RuntimeDirectory-
|
|
# guarded (trap 2). Non-fatal on a repo that predates the artifacts.
|
|
install_mgmt_watchdog() {
|
|
if $DRY_RUN; then
|
|
log_dry "fetch configs/felhom-privsep.tmpfiles -> /etc/tmpfiles.d/felhom-privsep.conf ; systemd-tmpfiles --create"
|
|
log_dry "fetch configs/felhom-mgmt-watchdog.sh -> /usr/local/sbin/felhom-mgmt-watchdog (0755)"
|
|
log_dry "fetch configs/felhom-mgmt-watchdog.{service,timer} -> /etc/systemd/system/ ; enable --now felhom-mgmt-watchdog.timer"
|
|
return 0
|
|
fi
|
|
local wdtmp; wdtmp=$(mktemp -t felhom-mgmt-wd.XXXXXX)
|
|
if ! fetch_raw "configs/felhom-mgmt-watchdog.sh" "$wdtmp" 2>/dev/null; then
|
|
log_skip " mgmt-watchdog artifacts not in the agent repo yet — break-glass auto-heal stays manual"
|
|
rm -f "$wdtmp"; return 0
|
|
fi
|
|
sh -n "$wdtmp" || { rm -f "$wdtmp"; die "fetched felhom-mgmt-watchdog.sh failed sh -n — refusing to install"; }
|
|
install -m 0755 -o root -g root "$wdtmp" /usr/local/sbin/felhom-mgmt-watchdog
|
|
rm -f "$wdtmp"
|
|
|
|
# tmpfiles (layer 1) — install + create now so /run/sshd is guaranteed present immediately.
|
|
local tftmp; tftmp=$(mktemp -t felhom-privsep.XXXXXX)
|
|
fetch_raw "configs/felhom-privsep.tmpfiles" "$tftmp"
|
|
install -m 0644 -o root -g root "$tftmp" /etc/tmpfiles.d/felhom-privsep.conf
|
|
rm -f "$tftmp"
|
|
systemd-tmpfiles --create /etc/tmpfiles.d/felhom-privsep.conf 2>/dev/null || true
|
|
|
|
# units (layer 2) — RuntimeDirectory guard on BOTH before install (trap 2 / the incident cause).
|
|
local svctmp tmrtmp; svctmp=$(mktemp -t felhom-wd-svc.XXXXXX); tmrtmp=$(mktemp -t felhom-wd-tmr.XXXXXX)
|
|
fetch_raw "configs/felhom-mgmt-watchdog.service" "$svctmp"
|
|
fetch_raw "configs/felhom-mgmt-watchdog.timer" "$tmrtmp"
|
|
if grep -qiE '^[[:space:]]*RuntimeDirectory[[:space:]]*=' "$svctmp" "$tmrtmp"; then
|
|
rm -f "$svctmp" "$tmrtmp"
|
|
die "mgmt-watchdog unit declares RuntimeDirectory= — that is the incident G1 fixes; refusing to install"
|
|
fi
|
|
install -m 0644 -o root -g root "$svctmp" /etc/systemd/system/felhom-mgmt-watchdog.service
|
|
install -m 0644 -o root -g root "$tmrtmp" /etc/systemd/system/felhom-mgmt-watchdog.timer
|
|
rm -f "$svctmp" "$tmrtmp"
|
|
systemctl daemon-reload
|
|
systemctl enable --now felhom-mgmt-watchdog.timer >/dev/null 2>&1 || true
|
|
log_success " installed break-glass layers 1+2 (tmpfiles /run/sshd + agent-independent watchdog timer)"
|
|
}
|
|
|
|
# install_oob installs the dedicated felhom-sshd OOB instance + the static nft belt (TASK H1). The
|
|
# STATIC parts land here (unit, host key, felhom-op user + sudoers, belt table + loader); the agent
|
|
# reconciles the DYNAMIC parts (port claim, config render, authorized_keys, belt set elements) once
|
|
# oob.enabled=true. Gated by --enable-oob. Non-fatal if the agent repo predates the artifacts.
|
|
# The felhom-sshd unit is RuntimeDirectory-guarded (the G1 incident cause) before install.
|
|
install_oob() {
|
|
$ENABLE_OOB || { log_skip " OOB (felhom-sshd) off (byo, or appliance --no-oob) — skipping"; return 0; }
|
|
if $DRY_RUN; then
|
|
log_dry "mkdir /etc/felhom-sshd{,/authorized_keys} ; ssh-keygen host key ; useradd felhom-op ; install felhom-op sudoers"
|
|
log_dry "fetch felhom-sshd.service (RuntimeDirectory-guarded) + felhom-oob.nft + felhom-oob-nft.service ; enable"
|
|
return 0
|
|
fi
|
|
install -d -o root -g root -m 0755 /etc/felhom-sshd /etc/felhom-sshd/authorized_keys
|
|
# dedicated host key (stable across reloads) — generate once.
|
|
if [[ ! -f /etc/felhom-sshd/ssh_host_ed25519_key ]]; then
|
|
ssh-keygen -t ed25519 -N "" -f /etc/felhom-sshd/ssh_host_ed25519_key -C felhom-sshd-hostkey -q
|
|
chmod 600 /etc/felhom-sshd/ssh_host_ed25519_key
|
|
fi
|
|
# operator login user (no ambient privilege; escalates only via the felhom-op sudoers verbs).
|
|
id felhom-op >/dev/null 2>&1 || useradd --create-home --shell /bin/bash felhom-op
|
|
# felhom-op scoped sudoers (visudo-validated before install).
|
|
local optmp; optmp=$(mktemp -t felhom-op-sudoers.XXXXXX)
|
|
if fetch_raw "configs/felhom-op.sudoers" "$optmp" 2>/dev/null; then
|
|
"$VISUDO" -cf "$optmp" >/dev/null || { rm -f "$optmp"; die "fetched felhom-op sudoers failed visudo -cf"; }
|
|
install -m 0440 -o root -g root "$optmp" /etc/sudoers.d/felhom-op
|
|
fi
|
|
rm -f "$optmp"
|
|
# felhom-sshd unit — RuntimeDirectory guard (the G1 incident cause) BEFORE install.
|
|
local sutmp; sutmp=$(mktemp -t felhom-sshd-unit.XXXXXX)
|
|
if ! fetch_raw "configs/felhom-sshd.service" "$sutmp" 2>/dev/null; then
|
|
log_skip " felhom-sshd.service not in the agent repo yet — OOB stays unconfigured"
|
|
rm -f "$sutmp"; return 0
|
|
fi
|
|
if grep -qiE '^[[:space:]]*RuntimeDirectory[[:space:]]*=' "$sutmp"; then
|
|
rm -f "$sutmp"; die "felhom-sshd unit declares RuntimeDirectory= — the SPIKE-felhom-sshd §8 incident cause; refusing"
|
|
fi
|
|
install -m 0644 -o root -g root "$sutmp" /etc/systemd/system/felhom-sshd.service
|
|
rm -f "$sutmp"
|
|
# static belt table + boot loader.
|
|
local ntmp utmp; ntmp=$(mktemp -t felhom-oob-nft.XXXXXX); utmp=$(mktemp -t felhom-oob-unit.XXXXXX)
|
|
fetch_raw "configs/felhom-oob.nft" "$ntmp"
|
|
fetch_raw "configs/felhom-oob-nft.service" "$utmp"
|
|
nft -c -f "$ntmp" >/dev/null 2>&1 || { rm -f "$ntmp" "$utmp"; die "felhom-oob.nft failed nft -c — refusing"; }
|
|
install -m 0644 -o root -g root "$ntmp" /etc/felhom-oob.nft
|
|
install -m 0644 -o root -g root "$utmp" /etc/systemd/system/felhom-oob-nft.service
|
|
rm -f "$ntmp" "$utmp"
|
|
systemctl daemon-reload
|
|
systemctl enable --now felhom-oob-nft.service >/dev/null 2>&1 || true # load the static belt now
|
|
systemctl enable felhom-sshd >/dev/null 2>&1 || true # NOT start — the agent renders the config first
|
|
log_success " installed OOB felhom-sshd instance + static belt (agent renders config + fills sets once oob.enabled)"
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# ensure_island_bridge — create the R-50 host-internal island bridge (vmbr9, portless, fixed /30) if
|
|
# absent, idempotently: append an ifupdown2 stanza to /etc/network/interfaces + `ifreload -a`.
|
|
# Portless (bridge-ports none) = no physical NIC, so the /30 cannot collide with any customer LAN and
|
|
# survives any LAN renumber; vmbr0 is a separate stanza and is never touched. Safe to re-run (no-ops
|
|
# when present). Spike-proven non-disruptive (SPIKE-island-bridge-2026-07-25, probe P2).
|
|
#-------------------------------------------------------------------------------
|
|
ensure_island_bridge() {
|
|
local ifaces=/etc/network/interfaces
|
|
if grep -qE "^[[:space:]]*iface[[:space:]]+${ISLAND_BRIDGE}[[:space:]]" "$ifaces" 2>/dev/null; then
|
|
log_info " island bridge ${ISLAND_BRIDGE} already present — leaving it"
|
|
return 0
|
|
fi
|
|
if $DRY_RUN; then
|
|
log_dry "append ${ISLAND_BRIDGE} stanza (portless, ${ISLAND_HOST_CIDR}) to $ifaces + ifreload -a"
|
|
return 0
|
|
fi
|
|
log_info " creating island bridge ${ISLAND_BRIDGE} (portless, ${ISLAND_HOST_CIDR})"
|
|
cp -a "$ifaces" "${ifaces}.pre-island.bak" 2>/dev/null || true
|
|
cat >> "$ifaces" <<EOF
|
|
|
|
# R-50 island control plane (felhom-host-install) — host-internal, no physical port. The agent binds
|
|
# ${ISLAND_HOST_IP}:${ISLAND_PORT} here; the guest gets ${ISLAND_GUEST_CIDR} on net1. Never add a
|
|
# bridge-port — the portlessness is what makes the /30 uncollidable with any customer LAN.
|
|
auto ${ISLAND_BRIDGE}
|
|
iface ${ISLAND_BRIDGE} inet static
|
|
address ${ISLAND_HOST_CIDR}
|
|
bridge-ports none
|
|
bridge-stp off
|
|
bridge-fd 0
|
|
EOF
|
|
if command -v ifreload >/dev/null 2>&1; then
|
|
run ifreload -a || log_warn " ifreload -a returned non-zero — verify ${ISLAND_BRIDGE} is up"
|
|
else
|
|
run ifup "${ISLAND_BRIDGE}" || log_warn " ifup ${ISLAND_BRIDGE} returned non-zero"
|
|
fi
|
|
ip -4 -o addr show "${ISLAND_BRIDGE}" 2>/dev/null | grep -q "${ISLAND_HOST_IP}" \
|
|
&& log_success " ${ISLAND_BRIDGE} up: ${ISLAND_HOST_CIDR}" \
|
|
|| log_warn " ${ISLAND_BRIDGE} did not come up with ${ISLAND_HOST_IP} — check $ifaces"
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# STEP 6 — write agent config + ensure service healthy
|
|
#-------------------------------------------------------------------------------
|
|
step_agent_config() {
|
|
log_step "6/8 agent config + service"
|
|
# E-2 Part 2: resolve the whole-guest backup target BEFORE the config is written — the resolved
|
|
# value feeds local_backup_target below. Scenario F safe: it returns early, changing nothing, when
|
|
# a target already exists, and `setdefault` preserves an existing box's whole backup block anyway.
|
|
configure_backup_target
|
|
# TLS pin: the SERVED leaf cert fingerprint (not pvesh node info — may differ)
|
|
local fp
|
|
fp=$(echo | openssl s_client -connect 127.0.0.1:8006 2>/dev/null | openssl x509 -noout -fingerprint -sha256 2>/dev/null | sed 's/.*=//')
|
|
[[ -n "$fp" ]] || log_warn " could not compute TLS fingerprint (leaving empty — agent will use system trust)"
|
|
# bridge / local-api addr. LAN_IP (vmbr0) is always resolved — it is the lan_resolver anchor even
|
|
# on an island install (Finding-1), and the historical bind on a non-island install.
|
|
LAN_IP=$(ip -4 -o addr show vmbr0 2>/dev/null | awk '{print $4}' | cut -d/ -f1 | head -1)
|
|
if [[ "$MODE" == "appliance" && "$WANT_ISLAND" == true && -z "$BRIDGE_ADDR" ]]; then
|
|
# R-50 appliance default: bind the local API on the host-internal island bridge (LAN-move
|
|
# immune). Create the bridge first so the agent can bind it; point listen_addr at it; the guest
|
|
# gets net1 from the agent (island_bridge/island_guest_addr, written below). lan_resolver.host_ip
|
|
# is pinned to the LAN IP in the config write — never let it default off the island listen_addr.
|
|
ensure_island_bridge
|
|
BRIDGE_ADDR="${ISLAND_HOST_IP}:${ISLAND_PORT}"
|
|
ISLAND_ON=1
|
|
log_info " R-50 island ON: local_api=${BRIDGE_ADDR} (${ISLAND_BRIDGE}); guest net1=${ISLAND_GUEST_CIDR}; lan_resolver.host_ip=${LAN_IP:-<none>}"
|
|
elif [[ -z "$BRIDGE_ADDR" ]]; then
|
|
BRIDGE_ADDR="${LAN_IP:-127.0.0.1}:8443" # byo / --no-island / explicit-LAN: historical bind
|
|
elif [[ "$BRIDGE_ADDR" != *:* ]]; then
|
|
BRIDGE_ADDR="${BRIDGE_ADDR}:8443"
|
|
fi
|
|
log_info " node=$NODE local_api=$BRIDGE_ADDR tls_fp=${fp:0:17}…"
|
|
|
|
if $DRY_RUN; then
|
|
log_dry "write $AGENT_CONFIG (0600): proxmox{endpoint,node=$NODE,token=<secret>,tls.fingerprint=$fp} hub{url=$HUB_URL,host_id=$HOST_ID,api_key=<secret>} local_api{$BRIDGE_ADDR}"
|
|
if [[ -n "$RESOLVED_OP_LINE$RESOLVED_REC_LINE" ]]; then
|
|
log_dry "write authz.signers: operational=${RESOLVED_OP_ID:-<none>} recovery=${RESOLVED_REC_ID:-<none>} (operator-signed self-update authority)"
|
|
else
|
|
log_dry "authz.signers: none resolved — self-update stays dormant (preserved signers, if any, are kept)"
|
|
fi
|
|
if [[ "$MODE" == "byo" ]]; then
|
|
log_dry "assert (byo) written config: lan_resolver.enable is false/absent"
|
|
log_dry "assert (byo) written config: oob.enabled is false/absent"
|
|
log_dry "(byo NOTE: wg_tunnel.enabled=true is now the DEFAULT both modes — WG is base infrastructure; consent = the typed disclosure ack)"
|
|
fi
|
|
log_dry "systemctl restart felhom-agent && felhom-agent --config $AGENT_CONFIG --selftest"
|
|
_state_mark agent_config; return 0
|
|
fi
|
|
|
|
mkdir -p "$(dirname "$AGENT_CONFIG")"
|
|
# Build config: optional preserve base + fresh-host defaults + Day-0 overrides.
|
|
# Secrets passed via env (NOT argv) to avoid ps exposure.
|
|
PVE_TOKEN="$PVE_TOKEN" HOST_API_KEY="$HOST_API_KEY" \
|
|
NODE="$NODE" FP="$fp" HUB_URL="$HUB_URL" HOST_ID="$HOST_ID" BRIDGE_ADDR="$BRIDGE_ADDR" \
|
|
ISLAND_ON="$ISLAND_ON" ISLAND_BRIDGE="$ISLAND_BRIDGE" ISLAND_GUEST_CIDR="$ISLAND_GUEST_CIDR" LAN_IP="$LAN_IP" \
|
|
OP_KEY_ID="$RESOLVED_OP_ID" OP_KEY_LINE="$RESOLVED_OP_LINE" \
|
|
REC_KEY_ID="$RESOLVED_REC_ID" REC_KEY_LINE="$RESOLVED_REC_LINE" \
|
|
PRESERVE_FROM="$PRESERVE_FROM" INSTALL_MODE="$MODE" ENABLE_OOB="$ENABLE_OOB" BACKUP_TARGET_RESOLVED="$BACKUP_TARGET_RESOLVED" OUT="$AGENT_CONFIG" python3 <<'PY'
|
|
import json, os, sys
|
|
out = os.environ['OUT']
|
|
base = {}
|
|
pf = os.environ.get('PRESERVE_FROM','')
|
|
if pf and os.path.exists(pf):
|
|
try: base = json.load(open(pf))
|
|
except Exception: base = {}
|
|
# fresh-host defaults for any section not preserved
|
|
base.setdefault('log_level','info')
|
|
# CAMPAIGN-3 Part 6: deployment_mode gates the agent's node self-heal (host networking recovery at
|
|
# boot, F12-class). Authoritative from --mode: "appliance" unlocks the remedy; "byo" (or anything
|
|
# else, including a stale preserved value) leaves it CHECK-only. Written explicitly so a byo box can
|
|
# never inherit an appliance self-heal from a --preserve-from config.
|
|
base['deployment_mode'] = 'appliance' if os.environ.get('INSTALL_MODE') == 'appliance' else 'byo'
|
|
# H1 / F9 (v1.25.0): when the belt is installed (appliance default, or explicit --enable-oob), turn the
|
|
# agent's OOB loop on so it claims a port, renders felhom-sshd's config, and fills the belt sets each
|
|
# tick. Without this the static units sit inert (the pre-v1.25.0 state that left the belt unusable). byo
|
|
# never reaches here with ENABLE_OOB true (refused above) and the byo asserts below re-check oob is off.
|
|
if os.environ.get('ENABLE_OOB') == 'true':
|
|
base.setdefault('oob', {})['enabled'] = True
|
|
# privileged.mode = "sudo": the canonical unit runs the agent as the NON-root felhom-agent user, so
|
|
# every host-root op goes through `sudo -n` against /etc/sudoers.d/felhom-agent. ("direct" was the old
|
|
# dev/CI shortcut for a root agent.) Force the mode authoritative (a stale preserved "direct" config
|
|
# would otherwise break the non-root daemon); the binary paths MUST match the sudoers allowlist.
|
|
base.setdefault('privileged', {})
|
|
base['privileged']['mode'] = 'sudo'
|
|
base['privileged'].setdefault('sudo_path','sudo')
|
|
for _k,_v in {"unit_dir":"/etc/systemd/system","stage_dir":"/var/lib/felhom-agent/units","systemctl":"/usr/bin/systemctl","install":"/usr/bin/install","smartctl":"/usr/sbin/smartctl","lvs":"/usr/sbin/lvs"}.items():
|
|
base['privileged'].setdefault(_k,_v)
|
|
base.setdefault('storage', {"watchdog_interval_seconds":5,"watchdog_debounce_seconds":15,"known_refresh_seconds":20})
|
|
# R-82: local DAILY + offsite WEEKLY. The two tiers carry SEPARATE cadences and retentions —
|
|
# keep_last=3 is three DAYS on the daily tier and three WEEKS on a weekly one, so one shared knob
|
|
# would guarantee that one of them is wrong. keep_last=2 on the offsite tier = two weeks (operator
|
|
# ruling 2026-07-26).
|
|
#
|
|
# The offsite tier is written even though `felhom-pbs` does not exist yet: that storage appears only
|
|
# when the hub provisions the DR tier. The agent DEFERS a tier whose target storage is absent
|
|
# (localapi targetStoragePresent), so the tier stays silent until it is real and goes live with no
|
|
# restart the moment it is provisioned. Writing it here means a box is never left with a DR tier that
|
|
# is provisioned but unscheduled — the exact "applied and empty" state R-82 exists to end.
|
|
#
|
|
# setdefault: an EXISTING box's backup block is preserved WHOLE. Upgrades never gain the tier here —
|
|
# they are migrated explicitly (R-82 Slice D.2), so an in-place upgrade can never silently start
|
|
# writing to an offsite datastore.
|
|
base.setdefault('backup', {"local_backup_target":os.environ.get('BACKUP_TARGET_RESOLVED','local'),"local_backup_retention":3,"restore_storage":"local-lvm","restore_test_cadence_seconds":0,"scratch_vmid_min":990000,"scratch_vmid_max":990009,"pbs_secret_dir":"/etc/pve/priv/storage","backup_cadence_seconds":0,"backup_targets":[{"target_id":"felhom-pbs","cadence_seconds":604800,"keep_last":2}]})
|
|
base.setdefault('local_api', {})
|
|
base['local_api'].setdefault('enable', True)
|
|
base['local_api']['listen_addr'] = os.environ['BRIDGE_ADDR']
|
|
# R-50: on an island install, record the island bridge + the guest net1 CIDR so the provisioner
|
|
# (bringup) attaches the guest's static island NIC. All-or-nothing — the agent's config.Validate
|
|
# rejects a half-set island. On a non-island install these keys are ABSENT (pre-R-50, LAN-only).
|
|
if os.environ.get('ISLAND_ON') == '1':
|
|
base['local_api']['island_bridge'] = os.environ['ISLAND_BRIDGE']
|
|
base['local_api']['island_guest_addr'] = os.environ['ISLAND_GUEST_CIDR']
|
|
base['local_api'].setdefault('cert_file','/var/lib/felhom-agent/local-api.crt')
|
|
base['local_api'].setdefault('key_file','/var/lib/felhom-agent/local-api.key')
|
|
base['local_api'].setdefault('token_store','/var/lib/felhom-agent/local-tokens.log')
|
|
# GL-2: byo defaults the lan-resolver OFF (the installer must never take over the owner's host DNS;
|
|
# the post-write byo asserts refuse a true value however it got in). Appliance keeps the historical
|
|
# default-on. setdefault either way — a --preserve-from section wins and the asserts judge it.
|
|
if os.environ.get('INSTALL_MODE') == 'byo':
|
|
base.setdefault('lan_resolver', {"enable": False})
|
|
else:
|
|
base.setdefault('lan_resolver', {"enable": True})
|
|
# R-50 Finding-1 (the dnsmasq trap): pin the split-horizon DNS to the LAN IP EXPLICITLY. Without
|
|
# this, config.go's LANResolverConfig.WithDefaults derives host_ip from listen_addr — which on an
|
|
# island install is the island IP (169.254.253.1) — silently moving LAN DNS onto the island bridge
|
|
# and killing it (spike-confirmed live). Set whenever the LAN IP is known (island or LAN bind).
|
|
_lan = os.environ.get('LAN_IP','').strip()
|
|
if _lan:
|
|
base['lan_resolver']['host_ip'] = _lan
|
|
# WG tunnel (drill F-9; operator decision 2026-07-12 #5): WG is BASE infrastructure like
|
|
# cloudflared — always-on for every install, both modes (it is also the future OOB-management
|
|
# transport; the PBS-DR tier merely RIDES it). Registration is hands-free (the drill proved it:
|
|
# keygen -> hub peer 10.77.0.x/32 -> handshake, no vouch). setdefault, not override: a
|
|
# --preserve-from config that deliberately disabled it is respected (the per-customer opt-out
|
|
# lives hub-side; byo consent = the disclosure ack + the tester agreement).
|
|
base.setdefault('wg_tunnel', {"enabled": True})
|
|
# Day-0 overrides (always authoritative)
|
|
base['proxmox'] = {
|
|
"endpoint":"https://127.0.0.1:8006",
|
|
"node": os.environ['NODE'],
|
|
"token": os.environ['PVE_TOKEN'],
|
|
"tls": {"fingerprint": os.environ['FP'], "insecure_skip_verify": False},
|
|
}
|
|
base['hub'] = {
|
|
"url": os.environ['HUB_URL'],
|
|
"host_id": os.environ['HOST_ID'],
|
|
"api_key": os.environ['HOST_API_KEY'],
|
|
"poll_seconds": base.get('hub',{}).get('poll_seconds',900),
|
|
"timeout_seconds": base.get('hub',{}).get('timeout_seconds',30),
|
|
}
|
|
# GL-4: operator-signed self-update authority (authz.signers, schema = agent config.go SignerKey).
|
|
# Script/file-provided keys are authoritative ONLY when non-empty — a reinstall with NO keys keeps
|
|
# a manually-pinned box's preserved signers (never silently clobber; C4). nonce_store_path is never
|
|
# invented here: the agent defaults it, and a preserved value rides along untouched in base['authz'].
|
|
signers = []
|
|
if os.environ.get('OP_KEY_LINE',''):
|
|
signers.append({"key_id": os.environ['OP_KEY_ID'], "role": "operational", "public_key": os.environ['OP_KEY_LINE']})
|
|
if os.environ.get('REC_KEY_LINE',''):
|
|
signers.append({"key_id": os.environ['REC_KEY_ID'], "role": "recovery", "public_key": os.environ['REC_KEY_LINE']})
|
|
if signers:
|
|
authz = base.setdefault('authz', {})
|
|
prev = authz.get('signers') or []
|
|
if prev:
|
|
print("[config] replacing %d preserved authz signer(s) with the script/file-provided operator key(s) (pin rotation)" % len(prev), file=sys.stderr)
|
|
authz['signers'] = signers
|
|
fd = os.open(out, os.O_WRONLY|os.O_CREAT|os.O_TRUNC, 0o600)
|
|
with os.fdopen(fd,'w') as f:
|
|
json.dump(base, f, indent=2); f.write('\n')
|
|
PY
|
|
# The non-root felhom-agent daemon must READ this config (token + hub api_key live here). Own it by
|
|
# the service user, 0600 (root still reads it for the provision one-shot).
|
|
chown "${AGENT_USER}:${AGENT_USER}" "$AGENT_CONFIG" 2>/dev/null || chmod 600 "$AGENT_CONFIG"
|
|
chmod 600 "$AGENT_CONFIG"
|
|
log_success " wrote $AGENT_CONFIG (0600 ${AGENT_USER})"
|
|
|
|
# GL-2 byo config asserts: these three features take over host DNS / open tunnels — forbidden on
|
|
# a customer/BYO box until their production endpoints exist. Parse the JUST-WRITTEN file (same
|
|
# interpreter pattern as the write above) and refuse to start the daemon on any true value —
|
|
# a --preserve-from carrying lan_resolver.enable=true is exactly what this catches.
|
|
# NOTE (v1.15.0 / decision 5): wg_tunnel.enabled was REMOVED from the byo-forbidden set — WG is
|
|
# base infrastructure (always-on both modes, like cloudflared); byo consent moved to the typed
|
|
# disclosure ack (the WG tunnel is now on that list) + the tester-agreement disclosure line.
|
|
if [[ "$MODE" == "byo" ]]; then
|
|
AGENT_CONFIG="$AGENT_CONFIG" python3 <<'PY' || die "byo config assert FAILED — see the keys above; a byo box must keep lan_resolver/oob off and deployment_mode=byo (fix the preserved config and re-run with --resume)"
|
|
import json, os, sys
|
|
d = json.load(open(os.environ['AGENT_CONFIG']))
|
|
bad = []
|
|
if d.get('lan_resolver', {}).get('enable'): bad.append('lan_resolver.enable (takes over host DNS on :53)')
|
|
if d.get('oob', {}).get('enabled'): bad.append('oob.enabled (starts a second operator sshd)')
|
|
if d.get('deployment_mode') == 'appliance': bad.append('deployment_mode=appliance (unlocks host-service self-heal on a host we do not own)')
|
|
if bad:
|
|
print('byo-forbidden config keys are TRUE: ' + '; '.join(bad), file=sys.stderr)
|
|
sys.exit(1)
|
|
PY
|
|
log_success " byo config asserts passed (lan_resolver.enable / oob.enabled off; deployment_mode=byo; wg_tunnel intentionally ON — base infrastructure)"
|
|
fi
|
|
|
|
# health: read-only selftest (proxmox) must pass before provisioning
|
|
if ! felhom-agent --config "$AGENT_CONFIG" --selftest >/dev/null 2>&1; then
|
|
felhom-agent --config "$AGENT_CONFIG" --selftest 2>&1 | tail -20 >&2
|
|
die "agent --selftest FAILED with the new config (token/ACL/TLS problem) — fix before provisioning"
|
|
fi
|
|
log_success " agent --selftest (read-only) passed"
|
|
|
|
# start the daemon (host-report loop) as the felhom-agent user and confirm it stays up. is-active is
|
|
# the real proof the NON-root user can read the 0600 config (the root selftest above can't show that).
|
|
if systemctl list-unit-files felhom-agent.service >/dev/null 2>&1; then
|
|
run systemctl enable felhom-agent >/dev/null 2>&1 || true
|
|
run systemctl restart felhom-agent
|
|
if ! $DRY_RUN; then
|
|
sleep 3
|
|
if systemctl is-active --quiet felhom-agent; then
|
|
log_success " felhom-agent service active (non-root $AGENT_USER reads the config OK)"
|
|
else
|
|
systemctl status felhom-agent --no-pager -l 2>&1 | tail -20 >&2
|
|
journalctl -u felhom-agent -n 20 --no-pager 2>&1 | tail -20 >&2
|
|
die "felhom-agent did not stay active after restart — see status/journal above"
|
|
fi
|
|
fi
|
|
else
|
|
log_warn " no felhom-agent systemd unit — daemon host-report loop not started (provision one-shot still works)"
|
|
fi
|
|
_state_mark agent_config
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# STEP 7 — golden: ensure a restorable golden archive (local else Gitea-fetched + verified)
|
|
#-------------------------------------------------------------------------------
|
|
# Local auto-discovery is the default + fallback. When no local golden exists (or --force-gitea-golden),
|
|
# fetch the golden from Gitea (git token), VERIFY its sha256 against the hub manifest, and import it
|
|
# into the archive storage's dump dir under a valid vzdump name so the provision restore can use it.
|
|
step_golden() {
|
|
log_step "7/8 golden archive"
|
|
|
|
if [[ -n "$GOLDEN_VOLID" ]] && ! $FORCE_GITEA_GOLDEN; then
|
|
log_skip " using local golden: $GOLDEN_VOLID"
|
|
_state_mark golden; return 0
|
|
fi
|
|
|
|
# Need the manifest + git creds (already resolved in step 5, but re-resolve on a fresh --resume run).
|
|
[[ -n "$ART_GOLDEN_VER" ]] || resolve_artifacts
|
|
[[ -n "$GIT_TOKEN" ]] || resolve_git_creds
|
|
[[ -n "$ART_GOLDEN_VER" && -n "$ART_GOLDEN_SHA" ]] || die "hub manifest has no golden version/sha256 — set it in the operator UI, or pass --golden VOLID"
|
|
|
|
local url="$GITEA_BASE/api/packages/$GITEA_OWNER/generic/felhom-golden/$ART_GOLDEN_VER/golden.tar.zst"
|
|
if $DRY_RUN; then
|
|
log_dry "$(_dry_auth_hint) -o <dump>/vzdump-lxc-${GOLDEN_VMID}-<ts>.tar.zst $url ; verify sha256=$ART_GOLDEN_SHA ; set GOLDEN_VOLID"
|
|
GOLDEN_VOLID="${ARCHIVE_STORAGE}:backup/vzdump-lxc-${GOLDEN_VMID}-<dry-run>.tar.zst"
|
|
_state_mark golden; return 0
|
|
fi
|
|
|
|
# Resolve the archive storage's dump dir (pvesm path maps a volid → fs path without needing it to exist).
|
|
local dump_dir fname dest
|
|
dump_dir=$(dirname "$(pvesm path "${ARCHIVE_STORAGE}:backup/vzdump-lxc-${GOLDEN_VMID}-2000_01_01-00_00_00.tar.zst" 2>/dev/null)")
|
|
[[ -d "$dump_dir" ]] || die "could not resolve dump dir for storage $ARCHIVE_STORAGE (got '$dump_dir')"
|
|
fname="vzdump-lxc-${GOLDEN_VMID}-$(date +%Y_%m_%d-%H_%M_%S).tar.zst"
|
|
dest="${dump_dir}/${fname}"
|
|
log_info " fetching golden v$ART_GOLDEN_VER from Gitea → $dest"
|
|
fetch_verify "$url" "$dest" "$ART_GOLDEN_SHA"
|
|
GOLDEN_VOLID="${ARCHIVE_STORAGE}:backup/${fname}"
|
|
pvesm list "$ARCHIVE_STORAGE" --content backup 2>/dev/null | grep -q "$fname" \
|
|
|| log_warn " imported golden not yet listed by pvesm (continuing — restore uses the volid directly)"
|
|
log_success " golden imported + verified: $GOLDEN_VOLID"
|
|
_state_mark golden
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# STEP 8 — provision (golden restore -> resize -> bootstrap.json -> onboot:1)
|
|
#-------------------------------------------------------------------------------
|
|
step_provision() {
|
|
log_step "8/8 provision guest $VMID"
|
|
# NOTE: -hub-password is passed on argv (the agent's only input for it) — briefly
|
|
# visible in ps. Tracked as an Observation (candidate: env/stdin in the agent).
|
|
# Optional operator CPU/RAM cap — passed to the agent ONLY when set (an agent < v0.52.0 would
|
|
# reject the unknown flag and die; opt-in means no one hits that until they intentionally cap).
|
|
local -a cap_args=()
|
|
[[ -n "$CPU_CORES" ]] && cap_args+=(-cores "$CPU_CORES")
|
|
[[ -n "$MEM_MIB" ]] && cap_args+=(-memory "$MEM_MIB")
|
|
# felhom pool: ensure it exists before the restore (step_token already created it; this is a
|
|
# belt-and-suspenders no-op that also covers a --resume path). The AGENT adds the guest to the pool
|
|
# atomically via restore --pool (v0.53.0) — no separate script-side pool_add_guest.
|
|
ensure_felhom_pool
|
|
if $DRY_RUN; then
|
|
log_dry "felhom-agent --config $AGENT_CONFIG --selftest=provision -archive $GOLDEN_VOLID -vmid $VMID -customer-id $CUSTOMER_ID -hub-password <pass> -rootfs-grow $ROOTFS_GROW -datavol-grow $DATAVOL_GROW -sysdata-grow $SYSDATA_GROW ${cap_args[*]} # agent restores INTO pool $PVE_POOL"
|
|
log_dry "record provisioned_vmid=$VMID in $STATE_FILE (for a later automatic --uninstall)"
|
|
_state_mark provision; return 0
|
|
fi
|
|
if ! felhom-agent --config "$AGENT_CONFIG" --selftest=provision \
|
|
-archive "$GOLDEN_VOLID" -vmid "$VMID" \
|
|
-customer-id "$CUSTOMER_ID" -hub-password "$PASSPHRASE" \
|
|
-rootfs-grow "$ROOTFS_GROW" -datavol-grow "$DATAVOL_GROW" -sysdata-grow "$SYSDATA_GROW" \
|
|
"${cap_args[@]}"; then
|
|
die "provision FAILED — see the agent error above. Fix and re-run with --resume."
|
|
fi
|
|
log_success " provision completed"
|
|
_state_mark provision
|
|
# Record the provisioned vmid so a later --uninstall resolves the target automatically + safely.
|
|
_state_put provisioned_vmid "$VMID"
|
|
# (No pool_add_guest here — the agent's restore --pool already made the guest a member.)
|
|
# Reboot the guest ONCE: the golden's controller-bootstrap unit evaluates its
|
|
# ConditionPathExists=/etc/felhom-bootstrap/bootstrap.json at BOOT, and the back-half attaches
|
|
# the bootstrap mount to the ALREADY-RUNNING guest — without a reboot the unit stays skipped and
|
|
# the controller never deploys (drill finding R6; the agent's own provision output says
|
|
# "next: reboot the guest"). On fast hosts the first boot sometimes wins the race — the reboot
|
|
# is idempotent either way (the unit no-ops when the controller already runs).
|
|
log_info " rebooting guest $VMID so the baked controller-bootstrap unit picks up the mount"
|
|
run pct reboot "$VMID"
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# STEP 7 — verify
|
|
#-------------------------------------------------------------------------------
|
|
step_verify() {
|
|
log_step "verify"
|
|
if $DRY_RUN; then log_dry "pct status/config $VMID; docker ps in-guest; host-report includes $VMID; assert guest $VMID in pool $PVE_POOL + scoped ACL grants present (user+token: Base@/, Guest@/pool/$PVE_POOL, Store@[${PVE_STORAGES[*]}])"; return 0; fi
|
|
local ok=true
|
|
local st; st=$(pct status "$VMID" 2>/dev/null | awk '{print $2}')
|
|
[[ "$st" == "running" ]] && log_success " pct status: running" || { log_error " pct status: $st"; ok=false; }
|
|
if pct config "$VMID" 2>/dev/null | grep -q '^onboot: 1'; then log_success " onboot: 1"; else log_error " onboot NOT 1"; ok=false; fi
|
|
pct config "$VMID" 2>/dev/null | grep -E '^(rootfs|mp0|mp1|mp8):' | sed 's/^/ /'
|
|
# GL-2 pool + ACL asserts (BOTH modes — campaign-2 R2 proved membership can silently drop, and a
|
|
# guest outside the pool is a guest the scoped token cannot manage).
|
|
local members; members=" $(pool_members || true) "
|
|
if [[ "$members" == *" $VMID "* ]]; then
|
|
log_success " pool: guest $VMID is a member of $PVE_POOL"
|
|
else
|
|
log_error " pool: guest $VMID NOT in pool $PVE_POOL (the scoped token cannot manage it — adopt with --adopt-pool)"; ok=false
|
|
fi
|
|
local _pairs=("/ $PVE_ROLE_BASE" "/pool/$PVE_POOL $PVE_ROLE_GUEST") _pair _apath _arole _acls2
|
|
for _acls2 in "${PVE_STORAGES[@]}"; do _pairs+=("/storage/$_acls2 $PVE_ROLE_STORE"); done
|
|
# F-LEAK: assert the scratch band too — a missing grant here is invisible until a restore-test
|
|
# FAILS, which is exactly the case that leaked a guest before v1.21.0.
|
|
local _sv; for ((_sv = PVE_SCRATCH_VMID_MIN; _sv <= PVE_SCRATCH_VMID_MAX; _sv++)); do _pairs+=("/vms/$_sv $PVE_ROLE_GUEST"); done
|
|
for _pair in "${_pairs[@]}"; do
|
|
_apath="${_pair% *}"; _arole="${_pair#* }"
|
|
if _acl_grant_present "$_apath" user "$PVE_USER" "$_arole" \
|
|
&& _acl_grant_present "$_apath" token "${PVE_USER}!${PVE_TOKENID}" "$_arole"; then
|
|
log_success " acl: $_arole@$_apath present (user+token)"
|
|
else
|
|
log_error " acl: $_arole@$_apath MISSING for the user and/or the token (re-apply with --rescope-acl)"; ok=false
|
|
fi
|
|
done
|
|
# GL-4: operator-key pin state. Dormant is the SAFE default, not an error — WARN only.
|
|
local _signers
|
|
_signers=$(python3 -c "import json;print(len(json.load(open('$AGENT_CONFIG')).get('authz',{}).get('signers') or []))" 2>/dev/null || echo 0)
|
|
if [[ "${_signers:-0}" -ge 1 ]]; then
|
|
log_success " authz signers: $_signers (operator-signed self-update armed)"
|
|
else
|
|
log_warn " no operator key pinned — agent self-update stays dormant (pin via --operator-pubkey-file or the script's OPERATOR_KEY_* constants; the key ceremony is an operator step)"
|
|
fi
|
|
# controller container healthy in-guest — bounded wait (the post-provision reboot + docker start
|
|
# take a while, especially on modest hardware; drill R6 re-verify)
|
|
local cstat="" _waited=0
|
|
while [[ -z "$cstat" && $_waited -lt 180 ]]; do
|
|
cstat=$(pct exec "$VMID" -- docker ps --filter name=felhom-controller --format '{{.Status}}' 2>/dev/null | head -1)
|
|
[[ -n "$cstat" ]] || { sleep 5; _waited=$((_waited+5)); }
|
|
done
|
|
if [[ -n "$cstat" ]]; then log_success " controller: $cstat (after ~${_waited}s)"; else log_warn " controller container not visible after ${_waited}s — check 'pct exec $VMID -- journalctl -u felhom-controller-bootstrap'"; fi
|
|
local cver; cver=$(pct exec "$VMID" -- docker ps --filter name=felhom-controller --format '{{.Image}}' 2>/dev/null | head -1)
|
|
[[ -n "$cver" ]] && log_info " controller image: $cver"
|
|
# tunnel
|
|
local tun; tun=$(pct exec "$VMID" -- docker ps --filter name=cloudflared --format '{{.Status}}' 2>/dev/null | head -1)
|
|
[[ -n "$tun" ]] && log_info " cloudflared: $tun" || log_warn " cloudflared not visible yet"
|
|
# host-report includes the guest (best-effort via the agent's hub selftest)
|
|
log_info " (confirm in the hub UI that host $HOST_ID reports guest $VMID)"
|
|
_state_mark verify
|
|
echo ""
|
|
if $ok; then
|
|
log_success "Day-0 provision SUCCESS — vmid=$VMID host_id=$HOST_ID customer=$CUSTOMER_ID golden=$GOLDEN_VOLID"
|
|
else
|
|
log_warn "Day-0 provision completed WITH WARNINGS — review the checks above"
|
|
fi
|
|
# Drill F-8: repeat the break-glass pointer in the FINAL summary (the 4b line scrolls away).
|
|
if _state_has break_glass; then
|
|
log_info "root@pam was rotated + vaulted at step 4b — retrieve at hub → host page (the old GUI password no longer works)."
|
|
fi
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# verify (agent-only, for --skip-provision): the agent is installed, runs non-root, and reports.
|
|
#-------------------------------------------------------------------------------
|
|
step_verify_agent() {
|
|
log_step "verify (agent only)"
|
|
if $DRY_RUN; then log_dry "felhom-agent --version; systemctl is-active felhom-agent; --selftest=hub (one collect+report)"; return 0; fi
|
|
local ok=true
|
|
log_info " binary: $("$AGENT_BIN" --version 2>&1 | head -1)"
|
|
log_info " runs as: $(systemctl show felhom-agent -p User --value 2>/dev/null) (want $AGENT_USER)"
|
|
if systemctl is-active --quiet felhom-agent; then log_success " service active"; else log_error " service NOT active"; ok=false; fi
|
|
# one explicit collect+report to prove the hub link end-to-end (host-report lands).
|
|
if felhom-agent --config "$AGENT_CONFIG" --selftest=hub >/dev/null 2>&1; then
|
|
log_success " --selftest=hub OK (a host-report reached the hub)"
|
|
else
|
|
log_warn " --selftest=hub did not confirm (the daemon loop still reports every poll_seconds)"
|
|
fi
|
|
_state_mark verify
|
|
echo ""
|
|
if $ok; then
|
|
log_success "Agent install SUCCESS — $("$AGENT_BIN" --version 2>&1 | head -1) as $AGENT_USER, host_id=$HOST_ID customer=$CUSTOMER_ID"
|
|
else
|
|
log_warn "Agent install completed WITH WARNINGS — review the checks above"
|
|
fi
|
|
}
|
|
|
|
#-------------------------------------------------------------------------------
|
|
# Main
|
|
#-------------------------------------------------------------------------------
|
|
# Scrub secret carriers on any exit; under --preflight-only also print the explicit FAIL verdict
|
|
# (the PASS verdict is printed by the preflight-only branch below).
|
|
# shellcheck disable=SC2154 # _rc IS assigned first inside the trap string itself
|
|
trap '_rc=$?; PASSPHRASE=""; PVE_TOKEN=""; HOST_API_KEY=""; GIT_TOKEN=""; if [[ $_rc -ne 0 ]] && $PREFLIGHT_ONLY; then log_error "PRE-FLIGHT FAIL (exit $_rc) — fix the finding above and re-run"; fi' EXIT
|
|
|
|
# --preflight-only: the full (mode-aware) step-1 checks, an explicit verdict, exit — no state marks,
|
|
# no later step (GL-2).
|
|
if $PREFLIGHT_ONLY; then
|
|
step_preflight
|
|
echo ""
|
|
log_success "PRE-FLIGHT PASS (mode=${MODE:-unset}) — no state written, no install step executed"
|
|
exit 0
|
|
fi
|
|
|
|
if $RESUME && _state_has preflight; then
|
|
# still need the passphrase for enroll/provision even on resume
|
|
read_passphrase
|
|
# re-resolve cheap derived values skipped steps would have set
|
|
[[ -n "$NODE" ]] || NODE=$(pvesh get /nodes --output-format json 2>/dev/null | python3 -c "import json,sys;print(json.load(sys.stdin)[0]['node'])" 2>/dev/null || hostname)
|
|
if [[ -z "$AGENT_CONFIG" ]] && systemctl cat felhom-agent >/dev/null 2>&1; then
|
|
AGENT_CONFIG=$(systemctl cat felhom-agent 2>/dev/null | grep -oP '(?<=-config )\S+' | head -1)
|
|
fi
|
|
[[ -n "$AGENT_CONFIG" ]] || AGENT_CONFIG="/etc/felhom-agent/agent.json"
|
|
# Backfill display values from the already-written config so the summary is complete.
|
|
[[ -f "$AGENT_CONFIG" ]] && HOST_ID=$(python3 -c "import json;print(json.load(open('$AGENT_CONFIG')).get('hub',{}).get('host_id',''))" 2>/dev/null || true)
|
|
# GL6-F4: golden's GOLDEN_VOLID feeds provision, but the resume path skips preflight (where local
|
|
# auto-discovery sets it). If the golden step already completed, the archive is on the local
|
|
# storage from run 1 — re-derive the volid so provision doesn't get an empty -archive. (When
|
|
# golden hasn't completed yet, step_golden runs and sets it.)
|
|
if [[ -z "$GOLDEN_VOLID" ]] && _state_has golden; then
|
|
GOLDEN_VOLID=$(pvesm list "$ARCHIVE_STORAGE" --content backup 2>/dev/null | awk -v v="$GOLDEN_VMID" '$0 ~ ("vzdump-lxc-" v "-"){print $1}' | sort | tail -1)
|
|
[[ -n "$GOLDEN_VOLID" ]] && log_info " golden (resumed from local): $GOLDEN_VOLID"
|
|
fi
|
|
log_skip "pre-flight (resumed)"
|
|
else
|
|
step_preflight
|
|
fi
|
|
|
|
# GL6-F4: token/enroll/grows are PRODUCERS whose IN-MEMORY outputs (PVE_TOKEN; HOST_ID/HOST_API_KEY;
|
|
# ROOTFS_GROW/DATAVOL_GROW/SYSDATA_GROW) are consumed by later steps (agent_config writes the token +
|
|
# hub.host_id into the config; provision passes the grows as flags). They MUST run every pass — even
|
|
# on --resume — or a resumed install writes a config missing hub.host_id/proxmox.token (daemon
|
|
# crash-loops) and provision gets `-rootfs-grow ""` (flag parse error). All three are idempotent:
|
|
# token reuses the on-disk token if it still authenticates (else rotates), enroll is a mint-once-
|
|
# reuse POST (200 REUSED), grows is a pure recompute — so unconditional re-run is cheap + correct.
|
|
# (This is why they are NOT behind `should_skip … ||`; step_token's own `_state_has` fast-path still
|
|
# short-circuits the reuse check when the token is already in memory within one process.)
|
|
step_token
|
|
step_grows
|
|
step_enroll
|
|
# GL-2: break-glass is gated at the CALL SITE (not inside the step) so the byo skip is auditable in
|
|
# one place. byo = a host the operator does not own: root@pam is the OWNER's credential — never
|
|
# reset, never vaulted.
|
|
if [[ "$MODE" == "appliance" ]]; then
|
|
should_skip break_glass || step_break_glass
|
|
else
|
|
log_skip "break-glass (step 4b) SKIPPED (byo: the host owner manages break-glass; root@pam untouched, nothing vaulted)"
|
|
fi
|
|
should_skip agent_install || step_agent_install
|
|
should_skip agent_config || step_agent_config
|
|
should_skip golden || step_golden
|
|
if $SKIP_PROVISION; then
|
|
log_skip "provision (--skip-provision) — agent install/config verified only"
|
|
step_verify_agent
|
|
else
|
|
should_skip provision || step_provision
|
|
step_verify
|
|
fi
|