PART 1+2 — dated checks stop being wishes. R-341 booked two measurements as prose in a register row; nothing read those dates and nothing would have objected when they passed. The dates now live in a DUE-CHECKS block INSIDE OPEN-ITEMS.md (inside, so no sidecar can drift from it) and a new gate reads them. Registered as #10 in repo_gates.py, --fast, so it runs in BOTH the pre-push hook and CI. exit 0 nothing due (prints pending count + nearest date; empty block too) exit 1 a row is due/overdue (due <= today, UTC -- due TODAY counts), or a row names an item with no R-row exit 2 block absent/duplicated/unparseable -- INCONCLUSIVE, never 0 It REFUSES rather than warns, and its docstring states the limitation: it is NOT a scheduler, it fires on the next push, not on the date. 37 tests. BOTH red-proofs run and reverted -- and the first one earned its keep by catching a hollow assertion of MINE rather than confirming the gate: flipping <= to < left a due-today row in neither bucket, min() raised on an empty list, and the TRACEBACK exited 1, so "rc == 1" passed while the boundary was wrong. An exit code cannot tell a verdict from a crash. The test now asserts the conviction banner and the absence of a traceback, and the gate returns 2 rather than crashing if that partition breaks again. PART 3 — the floor raise, and the premise was WRONG. Read back from the store (not the form): min_controller_version = 0.216.0 @ 12:36:58Z, zero per-customer overrides, no "managed floor HELD" line. But read 5 shows the raise was NOT a no-op: demo-felhom had been on 0.214.0 since 12 Aug and auto-updated 0.214.0 -> 0.216.0 at 12:37:07Z -- NINE SECONDS after the save, exactly the immediate action publish-train rule 2 documents. No error events followed; it restarted clean. R-343 is therefore filed OPEN, not CLOSED: the closing condition was all five reads clean and no directive served. It went well, but a record calling it inert when it moved a customer box is what misleads the next reader. The row also states why the floor was behind -- rule 2 policy, not drift, earned by the 2026-07-11 skew onto Peti's box -- and cites ResolveManagedFloor (store.go:2068) plus the two build-felhom-iso.sh facts (build-time at :267, fails open at :78-82) rather than asserting them. Two boxes are below the floor and neither reports: drill-r50 (blocked, powered off) and peti-felhom (host row deleted). peti-felhom was NOT contacted -- its row records that a report from a deleted host 401s and is not persisted, so the raise cannot reach it. PART 4 — R-342 filed READY, quoting stop2-snapshot.txt verbatim: Hetzner server snapshot 421440873 covers /dev/sda only; /mnt/pbs-datastore is a separate Volume that snapshots exclude, so a rollback restores software state and NOT the datastore. Fine for that upgrade; the safeguard for any future procedure that could touch the datastore does not exist and is Viktor's call. Also: CLAUDE.md's gate list named 6 of 10 registered gates -- completed rather than appending a 7th to a wrong list (124 -> 128 effective, ceiling 200). Capability map deliberately unchanged; no row cites a floor or golden version. repo_gates.py fully green, 10/10.
Felhom host scripts
Operator-side scripts for standing up a Felhom Proxmox host.
felhom-host-install.sh — Day-0 host bootstrap (operator-deploy)
Run on a freshly-PVE-installed box to fully automate Day-0: Proxmox API token →
hub host enrollment (single secret) → agent install (fetch + verify + install) →
agent config → golden → guest provision → verify. It composes already-proven mechanisms
(the pveum role/token sequence, the hub POST /host-enroll enrollment from option C, and
felhom-agent --selftest=provision). The agent renders bootstrap.json into the guest and
the controller pulls its own controller.yaml in-guest — the script never fetches that.
Since v1.1.0 (BUNDLE slice) the script also installs the agent itself: it fetches the
agent binary + golden from Gitea generic packages and verifies each artifact's sha256 against
the hub-vouched manifest (GET /api/v1/artifacts/{id}) before installing/using it. The fetch
credential is the git token already inside the customer's controller.yaml (config-retrieve) —
no new credential, and the checksum trust root is the hub, not Gitea.
Grounding: documentation/audits/SPIKE-day0-firstboot-handshake-2026-06-26.md.
Prerequisites (manual, before running)
- Install Proxmox VE 9.x on the box. During the installer, use Advanced → LVM
sizing so
local-lvm(thepve/datathin pool) has enough room for the appliance volumes — a useful box wants ≥ ~120 GiB free onlocal-lvm(rootfs 32G + Docker-data ~200G + user-data ~50G after grows). The script refuses below the hard minimum. - SSH into the box as root.
- Create the customer in the hub first (hub UI → new customer). The customer's retrieval passphrase (a 5-word Hungarian phrase) is the only secret you carry to the box.
That's it. The agent binary + golden are fetched + verified + installed by the script (provided the
operator has recorded the current artifact set in the hub UI → Configs → Day-0 artifacts, and
published them via felhom-agent/scripts/publish-agent.sh + configs/build-golden.sh). A local
golden, if present, is still used as a fallback.
Usage
curl -fsSL https://felhom.eu/scripts/felhom-host-install.sh -o felhom-host-install.sh
chmod +x felhom-host-install.sh
# secure no-echo passphrase prompt:
sudo ./felhom-host-install.sh --customer-id <customer>
# or from a 0600 file (no prompt):
sudo ./felhom-host-install.sh --customer-id <customer> --passphrase-file /root/.pass
# preview every mutating command without executing:
sudo ./felhom-host-install.sh --customer-id <customer> --dry-run
# resume after a fixed mid-way failure (skips completed steps):
sudo ./felhom-host-install.sh --customer-id <customer> --resume
The passphrase is read no-echo or from a 0600 file — never a CLI argument, never
echoed, never written to the state file or logs. The minted Proxmox-token secret and the
per-host hub api_key live only in the agent config (0600, root).
Key options
| Option | Default | Purpose |
|---|---|---|
--customer-id ID |
(required) | customer (must already exist in the hub) |
--vmid N |
9201 |
guest VMID to provision |
--golden VOLID |
newest vzdump-lxc-<golden-vmid> |
golden archive |
--rootfs/--datavol/--sysdata-grow N |
auto-compute | volume grows (GiB over the golden base 32/16/8) |
--passphrase-file PATH |
no-echo prompt | read passphrase from a 0600 file |
--preserve-from PATH |
— | merge non-Day-0 sections (PBS/local_api/privileged/authz) from an existing config |
--dry-run / --resume / --force |
off | preview / resume / clobber an existing vmid |
--mode provision|dr |
provision |
dr is a documented 10D stub (not implemented) |
Behaviour notes
- Idempotent + resumable. A step-state file (
/var/lib/felhom-install/state.json) records completed steps;--resumeskips them. A plain re-run refuses to clobber an existing--vmid(pass--forceto override). - Single-secret enrollment.
POST /host-enrollmints on first call (201) and reuses the credential on later calls (200) — re-running never orphans a running agent's key. The global operator key is never used. - Token automation. Creates/normalises the 16-priv
FelhomAgentrole, thefelhom-agent@pveuser + privsep token, and both ACL grants (user and token — the ACL is applied after the token exists, becausepveum user token removepurges it). - Agent install (v1.1.0). Fetches the binary from Gitea
(
/api/packages/admin/generic/felhom-agent/<ver>/felhom-agent), verifies its sha256 against the hub manifest, then installs the non-rootfelhom-agentservice user + binary + sudoers (0440,visudo -cf-validated) + the canonical systemd unit. Idempotent: same version already installed + service active → skips. A sha256 mismatch aborts the install (verify-before-use). The agent runs non-root (privileged.mode: "sudo"+ the sudoers allowlist), never as root. - Golden (v1.1.0). Uses a local golden when present; otherwise fetches it from Gitea
(
/api/packages/admin/generic/felhom-golden/<ver>/golden.tar.zst), verifies its sha256, and imports it into the archive storage's dump dir for the restore.--force-gitea-goldenforces the Gitea path even when a local golden exists. - DR mode (
--mode dr) is a documented seam only — it restores the customer's own PBS whole-CT snapshot instead of the golden. Not implemented (10D).
Productionization hooks (not done here)
- Serving: place this file where the felhom.eu site serves it at
https://felhom.eu/scripts/felhom-host-install.sh(a static route; verify on deploy). - Per-customer artifact pinning: the hub manifest currently returns the global current artifact
set for every customer; per-customer pinning is a future hook (
GET /api/v1/artifacts/{id}already takes the customer id). - Unit/sudoers integrity: the binary + golden are sha256-verified against the hub; the unit +
sudoers are fetched from the agent repo
main(canonical text) and the sudoers isvisudo -cf-validated.