Files
felhom.eu/scripts
admin 7f11cfb36c hub v0.65.0 — PBS DR storage visibility (ep0 usage op) + Offsite tab split + dual dashboard gauges (R-5)
Makes PBS DR storage visible like the restic pool box (v0.64.0), differentiated. Scoping
correction: restic = subaccounts on the shared Hetzner Storage Box (Hetzner API); PBS DR =
the felhom-offsite PBS datastore on the ep0 endpoint VM (NO Hetzner API). Option A
(Viktor-ruled): a read-only `usage` op on the felhom-tenantsync ep0 forced command (twin of
fingerprint), polled by a new hub checker on the 15-min throttle. READ-ONLY throughout.

Phase-0 (gate PASSED): on ep0 (PBS 4.2.3), df -B1 --output=size,used,avail <datastore path>
yields bytes (39990112256/7627939840/... ~19%), read-only, existing sudo context, no admin token.

- scripts/felhom-tenantsync.sh -> v1.2.0: read-only `usage` short-circuit (df on the datastore
  path), no customer_id, no admin token, NO mutation. + a bash harness proving zero mutation.
- tenantsync.Client.Usage() + BoxUsage; unknown-op -> typed ErrUsageUnsupported (graceful).
- monitor.PBSDRBoxChecker: OffsiteBoxChecker clone over a usageReader seam; 15-min throttle,
  cached PBSBoxSnapshot, escalation-only pbsdr_box_fill on the "pbsdr-box" scope (operator only,
  no SaveEvent), recovery re-arm. Fill only. THREE states: ok / unavailable (ep0 <=v1.1.0,
  neutral no-alert) / degraded (exec failed, keep last).
- config: Alerting.PBSDRBoxFill{Warn,Crit}Percent (80/90); built with the tenantsync client,
  60s sweep, SetPBSDRBox. Hub deploy INDEPENDENT of the ep0 update (graceful degradation).
- web: /offsite splits into Restic + PBS DR hash tabs (endpoint cards under PBS DR); PBS panel;
  the single dashboard tile becomes two gauges (RESTIC pct.ratio, PBS DR pct / n/a).
- runbook offsite-endpoint.md 10: v1.2.0 update steps (no sudoers/authorized_keys change).

Tests: 10 Go + the harness; 3 red-proofs (usage mutation, escalation-only, unavailable-drives-band)
confirmed red then restored. go build/vet/test + bash -n + hub confirm gate all pass.
2026-07-17 21:13:30 +02:00
..

Felhom host scripts

Operator-side scripts for standing up a Felhom Proxmox host.

felhom-host-install.sh — Day-0 host bootstrap (operator-deploy)

Run on a freshly-PVE-installed box to fully automate Day-0: Proxmox API token → hub host enrollment (single secret) → agent install (fetch + verify + install) → agent config → golden → guest provision → verify. It composes already-proven mechanisms (the pveum role/token sequence, the hub POST /host-enroll enrollment from option C, and felhom-agent --selftest=provision). The agent renders bootstrap.json into the guest and the controller pulls its own controller.yaml in-guest — the script never fetches that.

Since v1.1.0 (BUNDLE slice) the script also installs the agent itself: it fetches the agent binary + golden from Gitea generic packages and verifies each artifact's sha256 against the hub-vouched manifest (GET /api/v1/artifacts/{id}) before installing/using it. The fetch credential is the git token already inside the customer's controller.yaml (config-retrieve) — no new credential, and the checksum trust root is the hub, not Gitea.

Grounding: documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md.

Prerequisites (manual, before running)

  1. Install Proxmox VE 9.x on the box. During the installer, use Advanced → LVM sizing so local-lvm (the pve/data thin pool) has enough room for the appliance volumes — a useful box wants ≥ ~120 GiB free on local-lvm (rootfs 32G + Docker-data ~200G + user-data ~50G after grows). The script refuses below the hard minimum.
  2. SSH into the box as root.
  3. Create the customer in the hub first (hub UI → new customer). The customer's retrieval passphrase (a 5-word Hungarian phrase) is the only secret you carry to the box.

That's it. The agent binary + golden are fetched + verified + installed by the script (provided the operator has recorded the current artifact set in the hub UI → Configs → Day-0 artifacts, and published them via felhom-agent/scripts/publish-agent.sh + configs/build-golden.sh). A local golden, if present, is still used as a fallback.

Usage

curl -fsSL https://felhom.eu/scripts/felhom-host-install.sh -o felhom-host-install.sh
chmod +x felhom-host-install.sh

# secure no-echo passphrase prompt:
sudo ./felhom-host-install.sh --customer-id <customer>

# or from a 0600 file (no prompt):
sudo ./felhom-host-install.sh --customer-id <customer> --passphrase-file /root/.pass

# preview every mutating command without executing:
sudo ./felhom-host-install.sh --customer-id <customer> --dry-run

# resume after a fixed mid-way failure (skips completed steps):
sudo ./felhom-host-install.sh --customer-id <customer> --resume

The passphrase is read no-echo or from a 0600 file — never a CLI argument, never echoed, never written to the state file or logs. The minted Proxmox-token secret and the per-host hub api_key live only in the agent config (0600, root).

Key options

Option Default Purpose
--customer-id ID (required) customer (must already exist in the hub)
--vmid N 9201 guest VMID to provision
--golden VOLID newest vzdump-lxc-<golden-vmid> golden archive
--rootfs/--datavol/--sysdata-grow N auto-compute volume grows (GiB over the golden base 32/16/8)
--passphrase-file PATH no-echo prompt read passphrase from a 0600 file
--preserve-from PATH merge non-Day-0 sections (PBS/local_api/privileged/authz) from an existing config
--dry-run / --resume / --force off preview / resume / clobber an existing vmid
--mode provision|dr provision dr is a documented 10D stub (not implemented)

Behaviour notes

  • Idempotent + resumable. A step-state file (/var/lib/felhom-install/state.json) records completed steps; --resume skips them. A plain re-run refuses to clobber an existing --vmid (pass --force to override).
  • Single-secret enrollment. POST /host-enroll mints on first call (201) and reuses the credential on later calls (200) — re-running never orphans a running agent's key. The global operator key is never used.
  • Token automation. Creates/normalises the 16-priv FelhomAgent role, the felhom-agent@pve user + privsep token, and both ACL grants (user and token — the ACL is applied after the token exists, because pveum user token remove purges it).
  • Agent install (v1.1.0). Fetches the binary from Gitea (/api/packages/admin/generic/felhom-agent/<ver>/felhom-agent), verifies its sha256 against the hub manifest, then installs the non-root felhom-agent service user + binary + sudoers (0440, visudo -cf-validated) + the canonical systemd unit. Idempotent: same version already installed + service active → skips. A sha256 mismatch aborts the install (verify-before-use). The agent runs non-root (privileged.mode: "sudo" + the sudoers allowlist), never as root.
  • Golden (v1.1.0). Uses a local golden when present; otherwise fetches it from Gitea (/api/packages/admin/generic/felhom-golden/<ver>/golden.tar.zst), verifies its sha256, and imports it into the archive storage's dump dir for the restore. --force-gitea-golden forces the Gitea path even when a local golden exists.
  • DR mode (--mode dr) is a documented seam only — it restores the customer's own PBS whole-CT snapshot instead of the golden. Not implemented (10D).

Productionization hooks (not done here)

  • Serving: place this file where the felhom.eu site serves it at https://felhom.eu/scripts/felhom-host-install.sh (a static route; verify on deploy).
  • Per-customer artifact pinning: the hub manifest currently returns the global current artifact set for every customer; per-customer pinning is a future hook (GET /api/v1/artifacts/{id} already takes the customer id).
  • Unit/sudoers integrity: the binary + golden are sha256-verified against the hub; the unit + sudoers are fetched from the agent repo main (canonical text) and the sudoers is visudo -cf-validated.