Files
felhom.eu/scripts
admin 6088afcbed
gates / gates (push) Successful in 21s
Verify the standing picture against source: 12 downgrades, and the decay ran both ways
55 claims verified. Twelve moved, all downwards: walked 32 -> 20, built 5 -> 17.
Register ceiling R-284 -> R-290.

THE RULE DID NOT FIRE THE WAY IT WAS EXPECTED TO. Not one downgrade came from
code moving under an old proof. All twelve came from step 1 of the same rule --
the cited evidence does not exist. Measured: of the 28 capability-map rows
behind the page's claims, 8 carry a tests/ or audits/ path and 20 carry prose
only. The green dots were drawn from rows that cite an argument, not a walk
(R-290). The map, not the dataset, is what needs fixing -- it still says
PROVEN-LIVE for all twelve.

And once it ran backwards: fault.operator-email looked contradicted by R-182,
but live source shows the backup_run_failures digest allowlisted, operator-only
and templated, with recovery_unit_capture_failed now record-only. The claim is
right and the REGISTER ROW is stale (R-289). The session went looking for stale
proofs and found a stale defect.

R-281 WITHDRAWN -- wrong in both directions, settled by the operator's mailbox.
The tripwire DID fire (escrow_blob_served 10:19:41Z = 12:19 CEST) and false
error-severity alarms fired too, for deliberate attended work (R-285). The
measurement's cause is ESTABLISHED: the P7 query copied hub.db without hub.db-wal,
and the signature is exact -- it reported "2 events all day, newest 00:30:07",
and the rows at or before 00:30:07 number exactly 2. Timezone and wrong-key were
tested and refuted. The control had been drawn from the same stale snapshot as
the measurement, which is why it agreed (R-286).

Part 4: NO WORKFLOW CHANGED, deliberately. The gate is not ref-sensitive -- it
enumerates from the Gitea tags API, and both previous tag pushes passed. The red
is TRUE: run 267 saw v0.120.0 downloadable, run 284 on the same commit saw 404.
Who deleted the package is NOT established and is not guessed (R-287).

The page is now generated from where-felhom-stands.yaml by scripts/render_stands.py:
static, zero script tags, every moved status carrying a visible "changed, was X"
chip. The React bundle -- whose content was gzip+base64 inside a JS module map --
is kept as a dated snapshot. scripts/check_stands.py gates the data and convicted
51 problems in my own first draft before the staged positive control ever ran.
2026-08-09 18:40:49 +02:00
..

Felhom host scripts

Operator-side scripts for standing up a Felhom Proxmox host.

felhom-host-install.sh — Day-0 host bootstrap (operator-deploy)

Run on a freshly-PVE-installed box to fully automate Day-0: Proxmox API token → hub host enrollment (single secret) → agent install (fetch + verify + install) → agent config → golden → guest provision → verify. It composes already-proven mechanisms (the pveum role/token sequence, the hub POST /host-enroll enrollment from option C, and felhom-agent --selftest=provision). The agent renders bootstrap.json into the guest and the controller pulls its own controller.yaml in-guest — the script never fetches that.

Since v1.1.0 (BUNDLE slice) the script also installs the agent itself: it fetches the agent binary + golden from Gitea generic packages and verifies each artifact's sha256 against the hub-vouched manifest (GET /api/v1/artifacts/{id}) before installing/using it. The fetch credential is the git token already inside the customer's controller.yaml (config-retrieve) — no new credential, and the checksum trust root is the hub, not Gitea.

Grounding: documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md.

Prerequisites (manual, before running)

  1. Install Proxmox VE 9.x on the box. During the installer, use Advanced → LVM sizing so local-lvm (the pve/data thin pool) has enough room for the appliance volumes — a useful box wants ≥ ~120 GiB free on local-lvm (rootfs 32G + Docker-data ~200G + user-data ~50G after grows). The script refuses below the hard minimum.
  2. SSH into the box as root.
  3. Create the customer in the hub first (hub UI → new customer). The customer's retrieval passphrase (a 5-word Hungarian phrase) is the only secret you carry to the box.

That's it. The agent binary + golden are fetched + verified + installed by the script (provided the operator has recorded the current artifact set in the hub UI → Configs → Day-0 artifacts, and published them via felhom-agent/scripts/publish-agent.sh + configs/build-golden.sh). A local golden, if present, is still used as a fallback.

Usage

curl -fsSL https://felhom.eu/scripts/felhom-host-install.sh -o felhom-host-install.sh
chmod +x felhom-host-install.sh

# secure no-echo passphrase prompt:
sudo ./felhom-host-install.sh --customer-id <customer>

# or from a 0600 file (no prompt):
sudo ./felhom-host-install.sh --customer-id <customer> --passphrase-file /root/.pass

# preview every mutating command without executing:
sudo ./felhom-host-install.sh --customer-id <customer> --dry-run

# resume after a fixed mid-way failure (skips completed steps):
sudo ./felhom-host-install.sh --customer-id <customer> --resume

The passphrase is read no-echo or from a 0600 file — never a CLI argument, never echoed, never written to the state file or logs. The minted Proxmox-token secret and the per-host hub api_key live only in the agent config (0600, root).

Key options

Option Default Purpose
--customer-id ID (required) customer (must already exist in the hub)
--vmid N 9201 guest VMID to provision
--golden VOLID newest vzdump-lxc-<golden-vmid> golden archive
--rootfs/--datavol/--sysdata-grow N auto-compute volume grows (GiB over the golden base 32/16/8)
--passphrase-file PATH no-echo prompt read passphrase from a 0600 file
--preserve-from PATH merge non-Day-0 sections (PBS/local_api/privileged/authz) from an existing config
--dry-run / --resume / --force off preview / resume / clobber an existing vmid
--mode provision|dr provision dr is a documented 10D stub (not implemented)

Behaviour notes

  • Idempotent + resumable. A step-state file (/var/lib/felhom-install/state.json) records completed steps; --resume skips them. A plain re-run refuses to clobber an existing --vmid (pass --force to override).
  • Single-secret enrollment. POST /host-enroll mints on first call (201) and reuses the credential on later calls (200) — re-running never orphans a running agent's key. The global operator key is never used.
  • Token automation. Creates/normalises the 16-priv FelhomAgent role, the felhom-agent@pve user + privsep token, and both ACL grants (user and token — the ACL is applied after the token exists, because pveum user token remove purges it).
  • Agent install (v1.1.0). Fetches the binary from Gitea (/api/packages/admin/generic/felhom-agent/<ver>/felhom-agent), verifies its sha256 against the hub manifest, then installs the non-root felhom-agent service user + binary + sudoers (0440, visudo -cf-validated) + the canonical systemd unit. Idempotent: same version already installed + service active → skips. A sha256 mismatch aborts the install (verify-before-use). The agent runs non-root (privileged.mode: "sudo" + the sudoers allowlist), never as root.
  • Golden (v1.1.0). Uses a local golden when present; otherwise fetches it from Gitea (/api/packages/admin/generic/felhom-golden/<ver>/golden.tar.zst), verifies its sha256, and imports it into the archive storage's dump dir for the restore. --force-gitea-golden forces the Gitea path even when a local golden exists.
  • DR mode (--mode dr) is a documented seam only — it restores the customer's own PBS whole-CT snapshot instead of the golden. Not implemented (10D).

Productionization hooks (not done here)

  • Serving: place this file where the felhom.eu site serves it at https://felhom.eu/scripts/felhom-host-install.sh (a static route; verify on deploy).
  • Per-customer artifact pinning: the hub manifest currently returns the global current artifact set for every customer; per-customer pinning is a future hook (GET /api/v1/artifacts/{id} already takes the customer id).
  • Unit/sudoers integrity: the binary + golden are sha256-verified against the hub; the unit + sudoers are fetched from the agent repo main (canonical text) and the sudoers is visudo -cf-validated.