Waiting to be bound is the NORMAL state of a freshly installed box, and it must
not be reported as failure. The PAIRING poll loop used to BE systemd's
Restart=on-failure/RestartSec=30 — one poll per invocation, exiting non-zero
until the bind landed — so every 30s systemd printed "Failed to start Felhom
host bootstrap" on the physical console the CUSTOMER is watching. The
2026-07-18 N100 rehearsal measured 52 FAILED lines in ~11 minutes while nothing
was wrong (VALIDATION-n100-rehearsal-2026-07-18.md F6).
felhom-bootstrap.sh: run_pairing() is now a while-loop that sleeps
POLL_INTERVAL (30s — the hub-side rate is unchanged) between polls, so the unit
sits in `activating`. Registration split into register_appliance(), which
returns non-zero for a transient problem (no network yet, no identity, no
token) and is retried by the loop instead of taking the unit down. Cadence
constants: POLL_INTERVAL=30, BANNER_EVERY=10 (5 min), HEARTBEAT_EVERY=20
(10 min).
Quiet without going dark: a 204 is logged once on entry (worded so nobody reads
it as an error) and then only on the 10-minute heartbeat with elapsed minutes;
404 and unexpected codes degrade the same way. 410 STILL exits non-zero on
purpose — delivery consumed but no local env is a real crash window, and a
clean systemd restart is the right response.
Console banner: every 5 min instead of every cycle, single accented spelling
instead of the parositasra/párosításra double, and the reassurance the
rehearsal showed was missing ("Ez a képernyő magától frissül — nincs teendő a
doboznál").
felhom-bootstrap.service: TimeoutStartSec=infinity. This is load-bearing, not
cosmetic — a Type=oneshot ExecStart is killed at DefaultTimeoutStartSec (90s),
so without it systemd would kill the new in-script wait after 90 seconds and
Restart=on-failure would silently reinstate the exact spam this removes, after
appearing to work for the first three polls. Restart=/RestartSec= are kept
deliberately: they still cover the DIRECT path, a failed host-install, and 410.
Verified behaviourally, not assumed: driven in a throwaway Debian container
against a stub hub answering 204 five times then delivering — logged the wait
once plus one heartbeat, never exited between polls, then consumed the
delivery, wrote the 0600 env, fell through to the direct install in the same
invocation and exited 0. The old design produced five unit invocations and five
"Failed to start" console lines for that same sequence.
Hub endpoints, payloads, polling rate and one-shot delivery semantics are all
unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Nn3VgQk9iwEGgyx6QJ2NvE
Felhom host scripts
Operator-side scripts for standing up a Felhom Proxmox host.
felhom-host-install.sh — Day-0 host bootstrap (operator-deploy)
Run on a freshly-PVE-installed box to fully automate Day-0: Proxmox API token →
hub host enrollment (single secret) → agent install (fetch + verify + install) →
agent config → golden → guest provision → verify. It composes already-proven mechanisms
(the pveum role/token sequence, the hub POST /host-enroll enrollment from option C, and
felhom-agent --selftest=provision). The agent renders bootstrap.json into the guest and
the controller pulls its own controller.yaml in-guest — the script never fetches that.
Since v1.1.0 (BUNDLE slice) the script also installs the agent itself: it fetches the
agent binary + golden from Gitea generic packages and verifies each artifact's sha256 against
the hub-vouched manifest (GET /api/v1/artifacts/{id}) before installing/using it. The fetch
credential is the git token already inside the customer's controller.yaml (config-retrieve) —
no new credential, and the checksum trust root is the hub, not Gitea.
Grounding: documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md.
Prerequisites (manual, before running)
- Install Proxmox VE 9.x on the box. During the installer, use Advanced → LVM
sizing so
local-lvm(thepve/datathin pool) has enough room for the appliance volumes — a useful box wants ≥ ~120 GiB free onlocal-lvm(rootfs 32G + Docker-data ~200G + user-data ~50G after grows). The script refuses below the hard minimum. - SSH into the box as root.
- Create the customer in the hub first (hub UI → new customer). The customer's retrieval passphrase (a 5-word Hungarian phrase) is the only secret you carry to the box.
That's it. The agent binary + golden are fetched + verified + installed by the script (provided the
operator has recorded the current artifact set in the hub UI → Configs → Day-0 artifacts, and
published them via felhom-agent/scripts/publish-agent.sh + configs/build-golden.sh). A local
golden, if present, is still used as a fallback.
Usage
curl -fsSL https://felhom.eu/scripts/felhom-host-install.sh -o felhom-host-install.sh
chmod +x felhom-host-install.sh
# secure no-echo passphrase prompt:
sudo ./felhom-host-install.sh --customer-id <customer>
# or from a 0600 file (no prompt):
sudo ./felhom-host-install.sh --customer-id <customer> --passphrase-file /root/.pass
# preview every mutating command without executing:
sudo ./felhom-host-install.sh --customer-id <customer> --dry-run
# resume after a fixed mid-way failure (skips completed steps):
sudo ./felhom-host-install.sh --customer-id <customer> --resume
The passphrase is read no-echo or from a 0600 file — never a CLI argument, never
echoed, never written to the state file or logs. The minted Proxmox-token secret and the
per-host hub api_key live only in the agent config (0600, root).
Key options
| Option | Default | Purpose |
|---|---|---|
--customer-id ID |
(required) | customer (must already exist in the hub) |
--vmid N |
9201 |
guest VMID to provision |
--golden VOLID |
newest vzdump-lxc-<golden-vmid> |
golden archive |
--rootfs/--datavol/--sysdata-grow N |
auto-compute | volume grows (GiB over the golden base 32/16/8) |
--passphrase-file PATH |
no-echo prompt | read passphrase from a 0600 file |
--preserve-from PATH |
— | merge non-Day-0 sections (PBS/local_api/privileged/authz) from an existing config |
--dry-run / --resume / --force |
off | preview / resume / clobber an existing vmid |
--mode provision|dr |
provision |
dr is a documented 10D stub (not implemented) |
Behaviour notes
- Idempotent + resumable. A step-state file (
/var/lib/felhom-install/state.json) records completed steps;--resumeskips them. A plain re-run refuses to clobber an existing--vmid(pass--forceto override). - Single-secret enrollment.
POST /host-enrollmints on first call (201) and reuses the credential on later calls (200) — re-running never orphans a running agent's key. The global operator key is never used. - Token automation. Creates/normalises the 16-priv
FelhomAgentrole, thefelhom-agent@pveuser + privsep token, and both ACL grants (user and token — the ACL is applied after the token exists, becausepveum user token removepurges it). - Agent install (v1.1.0). Fetches the binary from Gitea
(
/api/packages/admin/generic/felhom-agent/<ver>/felhom-agent), verifies its sha256 against the hub manifest, then installs the non-rootfelhom-agentservice user + binary + sudoers (0440,visudo -cf-validated) + the canonical systemd unit. Idempotent: same version already installed + service active → skips. A sha256 mismatch aborts the install (verify-before-use). The agent runs non-root (privileged.mode: "sudo"+ the sudoers allowlist), never as root. - Golden (v1.1.0). Uses a local golden when present; otherwise fetches it from Gitea
(
/api/packages/admin/generic/felhom-golden/<ver>/golden.tar.zst), verifies its sha256, and imports it into the archive storage's dump dir for the restore.--force-gitea-goldenforces the Gitea path even when a local golden exists. - DR mode (
--mode dr) is a documented seam only — it restores the customer's own PBS whole-CT snapshot instead of the golden. Not implemented (10D).
Productionization hooks (not done here)
- Serving: place this file where the felhom.eu site serves it at
https://felhom.eu/scripts/felhom-host-install.sh(a static route; verify on deploy). - Per-customer artifact pinning: the hub manifest currently returns the global current artifact
set for every customer; per-customer pinning is a future hook (
GET /api/v1/artifacts/{id}already takes the customer id). - Unit/sudoers integrity: the binary + golden are sha256-verified against the hub; the unit +
sudoers are fetched from the agent repo
main(canonical text) and the sudoers isvisudo -cf-validated.