Live validation on demo-felhom caught the felhom-regen-hostkeys unit failing with 203/EXEC: ExecStart was /usr/sbin/ssh-keygen but on Debian 13 ssh-keygen is at /usr/bin/ssh-keygen. Fixed build-golden.sh, rebuilt the golden, re-validated — host keys now regenerate on first boot by the baked unit (agent issues no ssh-keygen). All three live scenarios green: provision (fresh MAC, host keys via unit, machine-id, Docker, DHCP), dr (continuity: hostname + host keys preserved), Recover (killed mid-restore -> orphan rolled back idempotently). REPORT + CHANGELOG updated. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
49 KiB
Changelog
All notable changes to felhom-agent are recorded here. Update on every code change that gets pushed.
v0.8.0 — slice 7 Phase 1: unified bring-up reconcile job (provision + guest-loss DR) (2026-06-09)
The shared FRONT HALF of provision and guest-loss DR, as a journaled reconcile job mirroring the
slice-6 restore-test's crash-safety — but it KEEPS the guest on success and applies a
scenario-specific identity policy. Agent-only; no hub/wire change (the new guest auto-appears in
the host-report via ListLXC). Grounded by the slice-7 bring-up spike findings (commit 3342993):
F1 (restore preserves the archived MAC → provision reset is unconditional), F3 (SSH host keys do
not auto-regenerate → a baked golden first-boot unit, not an agent guest-internal op), F4 (the
transient PVE config-lock 500 → bounded retry).
Added
reconcile.RunBringUp(bringup.go) —BringUpSpec(Modeprovision|dr_guest_loss, Archive, VMID, RestoreStorage, Hostname, Cores/MemoryMB, RootfsGrowGB, Mounts, KeepMAC, BootTimeout) →BringUpResult(VMID, AssignedMAC, Pass, Verified, StartWarnings/Recognized). Sequence (each mutation preceded by journaling the owning entry): restore → identity reset → size → attach mounts → start LINK-UP. Verdict is liveness (waitRunning), never the start exitstatus (reuses the v0.7.0 WARNINGS surface). Success KEEPS the guest (no teardown).- Scenario-specific identity reset (doc 03 §9): provision → fresh MAC unconditionally
(
PUT net0withhwaddromitted → PVE regenerates, F1) + hostname; machine-id + SSH host keys regenerate guest-side on first boot (golden bake + the new unit) — the agent does NOT touch guest internals. dr_guest_loss → preserve continuity (keep hostname; keep MAC unlessKeepMAC=false); never resets restic/tunnel/hub identity. - Compensating rollback — any mid-flight failure destroys the just-created guest
(
ClassGuestDestroy, benign viaProvenance{SameTxnCreated:true}, gated); on teardown failure the entry is left in-flight forRecover. New journal flagRollback+Recover'srecoverBringUpreap a half-built guest left by a mid-job crash (idempotent, viaListLXC). - F4 config-lock retry — steps 3+5 coalesced into ONE
PUT config(net0+hostname+cores+ memory+mpN); rootfs grow stays its own call.setConfigWithLockRetryretries ONLY the transient PVE config-lock 500 (pveConfigLock: 500 + "can't lock file"/"got timeout"); any other error fails immediately — never retried. --selftest=bring-up(-mode provision|dr -archive -vmid -hostname [-keep]) — runs the real journaled job (after aRecover), then tears the guest down unless-keep.configs/build-golden.sh— the validated golden recipe as a script, incl. the F3 first-bootfelhom-regen-hostkeys.serviceunit (Condition-gated: fires on provision, no-ops on DR). The slice-7 spike archive (which lacks the unit) is superseded.
Deferred (stated, not built)
- Provisioning BACK HALF (controller deploy, bootstrap, per-guest token mint) → slice 8.
- Host-loss DR + PBS escrow consumption → slice 10.
- The SOURCE of a
BringUpSpec(hub desired-state: which archive/VMID/mounts) → slice 10; this job takes the spec as input.GuestMountis defined minimally (no hub coupling).
Tests
- provision happy path (fresh MAC = net0 without hwaddr, hostname, coalesced sizing+mount, rootfs
grow separate, started, guest NOT destroyed); compensating rollback at each step (restore /
config / start-task / waitRunning — asserts the guest WAS destroyed); DR continuity (MAC kept,
hostname not reset) + DR
KeepMAC=falseresets MAC; liveness verdict (warnings+running pass / not-running fail); F4 (lock-500→retry→proceed; non-lock-500→fail without retry); owning entry journaled BEFORE restore; reserved/existing VMID refused;Recoverrolls back / clean.
Live-validated (demo-felhom)
- provision: fresh MAC + hostname; SSH host keys regenerated by the baked golden unit (agent
issued no
ssh-keygen), machine-id unique, Docker runs, clean DHCP lease → torn down. - dr: continuity preserved (hostname + host keys kept). Recover: a killed mid-restore left an
orphan; the re-run's
Recoverrolled it back (idempotent). - Live caught a bug, then fixed: the host-key unit's
ExecStartwas/usr/sbin/ssh-keygen(203/EXEC); on Debian 13 it is/usr/bin/ssh-keygen— corrected inbuild-golden.sh, golden rebuilt, re-validated. (Mocked unit tests couldn't surface this; the live run did.)
v0.7.0 — restore-test: verdict is liveness, not start-task exitstatus (2026-06-09)
Fixes a correctness bug found by the live hub-enrollment runbook: the self-restore-test reported
pass:false on every modern-distro guest. PVE's guest-start task exits "WARNINGS: 1" for the
benign systemd-nesting advisory (WARN: Systemd 257 detected. You may need to enable nesting.), and
WaitTask treated any non-"OK" exitstatus as a hard failure — so the verdict was decided by an
advisory exit code instead of by observed liveness, before the real boot check ran. A crying-wolf
test got it disabled on the demo host; this re-enables it. Single bump (0.6.0→0.7.0) covering the
agent's part of both task phases; the wire fields below are consumed by hub from v0.7.5.
Design invariant (in code): warning classification affects visibility only; pass/fail is liveness-only. A wrong/stale recognizer can at worst over-notice a benign warning — it can never false-fail and never hide a real warning.
Added
proxmox.WaitOptions.AllowWarnings— opt-in per call. When set, a task that completes"WARNINGS: N"is success with theTaskStatus(ExitStatus intact) returned so the caller can read/surface it. Default (false) keeps every existing caller strict (vzdump/restore/destroy warnings can be meaningful — relaxing them is a future per-call decision with evidence). Any non-WARNINGS non-OK exit is still a*TaskError.reconcile.RestoreTestResult.StartWarnings/.WarningsRecognized+ a version-free recognizer (benignWarningAnchor = "enable nesting", case-insensitive substring — contains no systemd version number, so it can't rot back into the bug at systemd 258+).extractWarningLinespullsWARN…lines from the start-task log.reconcile.GuestAPI.TaskLogTail— the engine fetches the start task's log to surface warnings.hub.RestoreTest.warnings/.warnings_recognizedwire fields (omitempty), populated byToHubRestoreTest. Additive: the deployed v0.7.4 hub ignores them; hub v0.7.5 consumes them (passed-with-warnings INFO, or WARN when not recognized). Cross-repo golden updated with the hub side.
Changed
- Restore-test start step (
reconcile/restoretest.go) now waits withAllowWarnings:true, surfaces any start warnings, and continues towaitRunningas the verdict — boot+running is the pass, exactly as before; a real (non-WARNINGS) start-task error still fails. The restore and scratch-teardown WaitTasks stay strict. - Restore-test scheduler logging distinguishes a clean pass, passed-with-recognized-warnings (INFO), and passed-with-unrecognized-warnings (WARN) — nothing silent.
Tests
WaitTask: AllowWarnings acceptsWARNINGS(status returned intact); AllowWarnings still fails a real error; default still fails onWARNINGS(existing callers unaffected).- Restore-test (engine, mock proxmox): start-with-warnings + running → pass with warnings surfaced+recognized; unrecognized warning + running → pass, not-recognized; not-running → fail regardless of warnings (verdict is liveness); teardown still runs.
- Regression guard: the
"enable nesting"recognizer matches the advisory for systemd 256–300, proving it's version-independent and can't silently rot back into the false-fail.
v0.6.0 — slice 6 Phase B: PBS offsite tier (verify + PBS-API client + reporting) (2026-06-09)
Completes slice 6. The PBS spike (felhom.eu phase5-pbs-spike-findings.md) proved backup-to-PBS and restore-from-PBS reuse Phase A UNCHANGED (PBS is just a storage target + a volid), and the operator token needs no widening. So the only new agent code is the verify capability + a small PBS-API client + PBSSnapshot reporting. Escrow + host-loss DR stay slices 7/10.
Added
internal/pbs— the PBS-API client (the agent's SECOND privileged external surface, slice-1 discipline): TLS fingerprint-pinned to the PBS leaf cert (a spoofed PBS → rejected, mirroring the PVE pin), token auth (PBSAPIToken=<id>:<secret>; id from the storageusername, secret read at runtime from/etc/pve/priv/storage/<id>.pw— referenced by location, never logged/committed), typed, no shell. Methods:Verify(POST/admin/datastore/<ds>/verify→ UPID),Snapshots(incl. theverificationfield),TaskStatus/WaitVerify(node extracted from the UPID —localhostreturns "unknown", the spike B4 gotcha),NodeFromUPID.- The verify maintenance loop (
pbs/verify.go) — the cheap, key-free, ciphertext-level integrity check (§8) on its OWN cadence (default 6h, the 5th daemon goroutine). It is a reporting/maintenance task like the slice-5 watchdog: it does NOT go through the reconcile gate/journal. Each cycle: trigger verify → poll task → re-list snapshots → record per-snapshotverify_state. A failed verify is logged loudly. PBSSnapshotreporting — filled the stub (namespace/backup_type/backup_id/backup_time(RFC3339)/size_bytes/owner/protected/encrypted(fromfiles[].crypt-mode) /verify_state(ok|failed|none until verified)/verify_upid). NewPBSReportercollector seam + an in-memorySnapshotStore. Cross-repo golden (both repos, byte-identical)- bidirectional key-set tests; hub
handler.goparsespbs_snapshotsand logs a failed verify[WARN](loudest offsite-DR signal).
- bidirectional key-set tests; hub
- Truthful backup mode (
backup/runner.go) —Backup.modenow reflects the ACTUAL vzdump mode read from the task log (backup mode: <x>), since PVE may downgrade snapshot→stop for a stopped guest (spike B1); falls back to the requested mode if unparseable. - proxmox:
Storage.Username(parsed from the pbs storage config — the token id). - config
BackupConfig.{PBSVerifyCadenceSeconds, PBSSecretDir}(cadence 0→6h, <0 disabled). --selftest=pbs-verify— discover pbs storages → verify each → print the PBSSnapshot records (covers the runbook's verify + list). Standalone on the host.
Notes
- Backup/restore-to-PBS reuse Phase A with no change (the restore-test runs with
source_tier="pbs"when fed a pbs volid). Zero-knowledge holds: verify is ciphertext-level, the encryption key is never read here, and the PBS server has no client key (spike B6). - Daemon runs cleanly with no pbs storage / verify disabled.
go test -racecovers the new goroutine. Slice-3/4/5/6A surfaces, goldens, and adversarial tests intact.
v0.6.0-rc1 — slice 6 Phase A: backup + the self-restore-test (local target) (2026-06-09)
Phase A of the backup/restore slice (doc 03 §8) — the agent's guest-level backup layer and the self-restore-test, which closes "a backup you haven't restored isn't a backup". Everything here is BENIGN (backup, restore-to-NEW, scratch teardown): reuses the slice-4 classifier/gate/journal — no new destructive class, no new crypto. Local target only; PBS is Phase B. Restore is to a NEW guest only (no overwrite). Backups are crash-consistent only (app-consistency needs the controller quiesce, slice 8) — marked so in the report.
Added
- proxmox (
mutate.go/query.go):DestroyLXC(DELETE …/lxc/{vmid}?purge=1&destroy- unreferenced-disks=1 → UPID; the scratch-teardown primitive);VzdumpOptions.Notes→notes-template(verified on PVE 9.2.2);LatestBackupVolID(resolve a produced archive from the backup-storage listing — the task status carries no result volid). - reconcile self-restore-test (
restoretest.go) —Engine.RunRestoreTest: pick a free scratch VMID (configured band, excludes 9999; full band → skip, never out-of-band) → journal a Scratch-owned entry BEFORE any mutation → restore-to-new → benign net link-down SetConfig (so the clone can't conflict with a running source's MAC/IP; this is test-safety, NOT slice-7 identity reset) → boot → verify reachesrunning→ ALWAYS teardown (defer; benignClassGuestDestroy+ agent-tagged-scratch provenance, gated). Runs on the scratch VMID's queue lane. Reuses the journal/gate; result feeds the report. - Crash-safe recovery (
recover.go): a Scratch journal entry is resolved by TEARDOWN, not by re-checking the restore sub-task's UPID — special-cased BEFORE the generic path (else the restore task's OK would mark it succeeded while the guest leaks).Recovernow destroys a leaked scratch guest (idempotent: already-gone → clean; list-unreadable → left in-flight for a later pass).JournalEntry.Scratchflag;RecoverResult.ScratchClean/ ScratchDestroyed. GuestAPI gainsRestoreLXC/DestroyLXC/GuestStatus. internal/backuppackage:BackupRunner.Backup(vzdump + archive/size resolve + bulk-volume gap — a mountpoint is UNCOVERED unless it carries an explicitbackup=1, so an unsetbackup=is reported uncovered too, the safe DR direction);PickRestoreCandidate(newest backup); an in-memoryStore(latest-backup-per-target + latest-restore-test) implementing the hubBackupReporter/RestoreTestReporterseams; a cadenceScheduler(default 24h; the fourth daemon goroutine; disabled cleanly when off/misconfigured).- hub report (
report.go): filled theBackup+RestoreTeststubs (PBSSnapshotstays a Phase-B stub); collectorBackupReporter/RestoreTestReporterseams. Cross-repo golden updated in BOTH repos (byte-identical) + bidirectional key-set tests forbackups[0]/restore_tests[0]. Hubhandler.goparses + persists them (report_json; no new columns) and logs a FAILED restore-test prominently (the loudest DR signal). - config
BackupConfig(local target, restore storage, restore-test cadence, scratch VMID band 990000–990009 default) + accessors + env overlay + cadence-gated validation. --selftest=backup -vmid N(one-shot backup → print the Backup record) and--selftest=restore-test [-archive volid](Recover-then restore→boot→verify→teardown, print the RestoreTest record). Standalone on the Proxmox host.
Notes
- The daemon runs cleanly with the cadence off or misconfigured (logs + disables, never
crashes); a leaked scratch guest from a mid-test crash is reaped by
engine.Recoveron restart.go test -racecovers the new scheduler goroutine. - Slice-3/4/5 exported surfaces, goldens, and adversarial tests intact. Version bumps to v0.6.0 when Phase B (PBS) lands.
v0.5.1 — slice 5 live-validation prep: durable_id mis-id fix + re-mount UUID memory (2026-06-09)
Two correctness fixes surfaced while preparing the live USB validation on demo-felhom
(a real 1TB USB HDD, sdb1, ext4). Both are DR-load-bearing — exactly the "false-id →
re-attach the wrong disk" failure mode the slice warned about.
Fixed
- Unmounted dir-storage no longer inherits the ROOT filesystem's UUID (
observe.go). Previously, when a removable dir-storage was unmounted, the observer fell through to the containing mount (root) for the backing device, so itsdurable_idbecameuuid:<root-uuid>— a catastrophic DR mis-id (the hub would re-attach the wrong disk). Now the backing device/UUID/durable_idare derived ONLY from the target's OWN mountpoint; an unmounted target reports no device and a stablestore:<name>durable_id, never another filesystem's UUID. (Removed thecontainingMountDeviceroot-fallthrough.) - Watchdog remembers the fs-UUID observed while attached (
watchdog.go) so a re-mount works even after the known-set cache refreshes mid-drop (an unmounted target can't resolve its own UUID). The re-mount key is backfilled from this memory — aligning with doc 03 §7's "sourced from the existing definition, no hub manifest needed": the agent learns the UUID while the target is attached, then re-mounts by it on return.
Tests
- Observer: an unmounted dir-storage asserts NO
uuid:durable_id and no backing device. - Watchdog: a drop where the cache lost the UUID still re-mounts using the remembered UUID.
v0.5.0 — slice 5 Phase B: the host-root surface (mounts + SMART + grow + destructive gate) (2026-06-09)
The write surface — the agent's first step outside its Proxmox API token into OS-root. Isolated behind a narrow, argument-validated, adversarially-tested seam, exactly like the slice-4 gate. Completes slice 5 (Phase A = read-only observe/report/watchdog at v0.5.0-rc1).
Added
HostOpsseam +SudoHostOps(internal/storage/hostops.go) — the one privileged host surface: persistent mounts via systemd.mountunits keyed by fs-UUID (enabled to survive reboot), detach (stop+disable), SMART, and thin-pool metadata. Shells out via the fenced Runner (sudo -n, fixed arg vectors, no shell); a fake backs the tests (no real root in the suite).NoopHostOpsis the safe fallback when the surface is unavailable.- The argument validator (
internal/storage/validate.go) — the security boundary:ValidateUUID(strict hex),ValidateMountPath(absolute, no traversal, no metacharacters),ValidateSMARTDevice(raw-disk whitelist),ValidateLVMName, and an in-processsystemdEscapePath(nosystemd-escapeshell-out). Every argument is validated BEFORE a command is constructed. Headline test (validate_test.go): an adversarial matrix of shell metacharacters /../traversal / malformed inputs is rejected with zero exec. - SMART (
internal/storage/smart.go) — parsessmartctl -a -jintoStorageTarget.smart: SATA (reallocated/pending/offline-uncorrectable, temp, power-on-hours) and NVMe (critical_warning, media_errors, percentage_used, temp), degrading toUNKNOWNfor devices with no SMART (USB-SATA bridges).lvsfills the lvmthin thin-pool metadata fill (the value Phase A left null). Wired into the Observer's enrichment (Observe only, not the watchdog's fast Known path). - Watchdog re-mount response (
internal/storage/watchdog.go) — on a known mount-backed target's device returning unmounted (a newDevicePresentliveness probe), the watchdog dispatches a benign by-UUID re-mount off the poll path (a goroutine, never under the lock), rate-limited per target to the debounce window. The mount is routed through the gate as benign (gateRemounterinmain.go, sostoragestays decoupled fromreconcile). - Disk-grow executor (
internal/reconcile) —ActionResize(benignClassResize), planned grow-only (desired DiskBytes > actual →pct resize rootfs +<n>M; a shrink is refused, never silently grown) + a defensive executor guard (size must start with+). Newproxmox.Client.ResizeLXC(API;VM.Config.Disk+Datastore.AllocateSpace; async→UPID). Built + fixture-tested; unfed live (no hub spec until slice 10). - Destructive storage ops through the slice-4 gate (
internal/reconcile/storage_ops.go) —IntentForStorageMount(benign) andIntentForStorageDestructive(ClassStorageWipe/ClassDecommission). Host/target-scoped: the op binds on the storage target identity (carried intarget.guest_id). Reuses the existing verifier/role-scoping/binding/audit — no new gate, no new crypto. Storage cases added to the adversarial matrix (storage_test.go): unsigned wipe →pending_signature; "wipe A" signature vs "wipe B" →binding_mismatch; valid → accepted. Inert live. --selftest=storage[-watch <dur>] — the live USB-runbook harness: an observe pass (full table incl. SMART + thin-pool data+metadata), and a bounded watchdog window with the re-mount response live. Runs standalone on the Proxmox host (no hub).configs/felhom-agent.sudoers— the documented narrow allowlist (install unit / systemctl manage / smartctl / lvs), with the agent-side fine validation noted.- Config:
privileged.{unit_dir,stage_dir,systemctl,install,smartctl,lvs}(paths must match the sudoers entries).
Notes
- Daemon still runs cleanly with no removable storage / no signers / no hub manifest, and a
missing/declined sudoers entry degrades with a warning (SMART→UNKNOWN, mount→logged error),
not a crash.
go test -racepasses (the watchdog re-mount dispatches off the poll path). - Slice-3/4 + Phase-A exported surfaces, goldens, and adversarial tests intact.
authzuntouched. The destructive-storage executor + grow are built/tested but unfed live until slice 10.
v0.5.0-rc1 — slice 5 Phase A: storage observe + report + watchdog (read-only, live) (2026-06-09)
Phase A of the storage slice (doc 03 §7). Read-only and live: the agent now observes every
host storage target, reports it into the host-report's storage_targets (previously an empty
stub), and runs a fast-poll watchdog that pushes a disconnect to the hub in seconds. No
host-root writes this phase (mounts/SMART/grow/destructive-gate are Phase B). The hub-owned
desired manifest (class/role/policy/creds) is not served until slice 10, so reconcile against
it is built-but-unfed — this phase ships only the genuinely-useful read-only footprint.
Added
internal/storagepackage (new):StorageTargetwire contract (internal/hub/report.go) — filled the slice-3 stub:name/type/durable_id/state/reachable, usage (total/used/avail/used_fraction),content,mount_path/backing_device,class_hint(rotational HINT — never authoritative; class is hub-owned),role(empty until slice 10), athin_poolsub-object (lvmthin data fill; metadata fill is Phase B/lvs), and asmartsub-object (UNKNOWNuntil Phase B). Cross-repo golden kept byte-identical withfelhom.eu/huband guarded by the bidirectional key-set test (contract_test.go).durable_idderivation (durableid.go) — deterministic per type (the DR-load-bearing re-attach key): fs-UUID (usb/local-dir),server:export(nfs/cifs),repo+fingerprint(pbs),vg/pool(lvmthin); never empty (falls back to a stable store id).HostReaderseam +ProcHostReader(hostread.go) — non-privileged/proc/mounts,/dev/disk/by-uuid,/sys/.../rotational+removablereads. Root-free by construction.Observer(observe.go) — builds[]hub.StorageTargetfromListStorage/NodeStoragejoined with host reads; surfaces the lvmthin thin-pool data fill prominently (warns ≥85%).- Storage watchdog (
watchdog.go) — a third daemon goroutine fast-polling the known target set (a defined Proxmox storage and/or a previously-seen one) forattached↔disconnectedtransitions; on a transition it triggers an immediate, debounced out-of-band host-report. Only flags a known target's change (never a never-attached device); coalesces flaps within the debounce window (leading + trailing edge).CachingKnownTargetsrate-limits the Proxmox-derived known set;HostLivenessprobes device/mount presence (local) + a reachability dial (network), all non-privileged.
- Proxmox
Storagetype (internal/proxmox/types.go) — additive parse-only config fields (server/export/share/datastore/fingerprint/vgname/thinpool) feeding durable_id. - Collector
StorageObserverseam (internal/hub/collect.go) — populatesstorage_targetsvia the observer; a nil observer or an observe error degrades to empty (never sinks the heartbeat). Hub does not import storage (storage imports hub for the wire type). - Out-of-band report trigger (
internal/hub/loop.go) —Loop.SetTrigger: a watchdog signal runs one extra collect→report immediately without disturbing the regular cadence. StorageConfig(internal/config) — watchdog interval / debounce / known-refresh knobs (all optional; package defaults otherwise).- Hub ingest (
felhom.eu/hub) —hostReportPayloadnow parsesstorage_targets(full mirror struct), persists them viareport_json, counts + warns on disconnected targets, and has its own half of the bidirectional golden key-set test.
Notes
- The daemon still runs cleanly with no removable storage, no signers, and no hub manifest — the watchdog finds nothing to flag; storage reporting is best-effort.
proxmox/hub/authz/reconcileexported surfaces + their golden/adversarial tests are intact. No host-root writes, no destructive paths, no SMART this phase (all Phase B).- Version: v0.5.0-rc1 at the Phase-A checkpoint; v0.5.0 when Phase B lands.
v0.4.0 — slice 4 Phase B: reversibility gate + signed-op consuming layer (2026-06-08)
The security core of slice 4: hub-supplied intent stops being trusted for destructive
change. Layered in front of the per-guest queue's executor — every mutation now
passes the gate. Reuses internal/authz for all crypto (untouched surface). Inert
this slice: no destructive deltas are served until slice 10, so the destructive path is
classified, gated, and adversarially tested but not wired to live execution.
Added
- Classifier (
classify.go, doc 03 §4) — benign vs destructive by provenance + data-bearing-ness, NOT by verb. TheOpClassvocabulary (seeded by the committed slice-2op_blob.json:guest_destroy) is the agent-side contract slice 10 matches. Destroy/overwrite of customer data is destructive UNLESS agent-internal provenance (same-journaled-transaction create → compensating rollback, or agent-tagged scratch) makes it benign.Provenanceis journal-recorded and never populated from the hub (its zero value is the only thing an external intent may carry). Unknown op class fails safe → destructive. - Reversibility gate (
gate.go) —Gate.Authorize(intent, signed): benign → allowed unsigned; destructive → requires a verified, role-authorized, action-bound operator signature, else refusedpending_signature, never executed. Every decision is written to anAuditSink(audit is a signal, never the guard). - Signed-op consuming layer over
authz— verifies viaauthz.Verifier.Verify(the locked pipeline, untouched), then enforces on theVerifiedOp:- Role-scoping (doc 04 §4) — recovery key authorizes key-rotation re-pins ONLY; operational key authorizes ordinary destructive ops + planned rotation.
- Op-to-action binding — verified
op+ host + guest +paramsmust match the gated action (a signature for guest X / op A can't authorize guest Y / op B); params compared semantically (key-order/whitespace independent).
- Signed-job orchestration (
job.go) —RunSignedJob: idempotency dedupe (the op nonce as the journal key — a redelivered completed op is skipped, not re-run), gate authorization, then journal-wrapped execution via an injectedDestructiveExecutor(nil this slice — authorized destructive ops are inert, no executor wired until 6/7). - Crash-recovery consumer (
recover.go, Note 1 / doc 03 §10) —Engine.Recoverconsumes the journal'sInFlight()at startup: an op that crashed AFTER the Proxmox POST and BEFORE its terminal record (OpTaskRunning, nonce already consumed) is NOT covered by idempotency dedupe — only this resume-or-rollback resolves it (re-read the task via the newTaskStatusOnce, record the real outcome; a no-task-id op is abandoned fail-safe). Landed together with the signed-op executor, as Note 1 required. - Daemon wiring —
runDaemonbuilds the verifier fromconfig.Authz.Signers(a bad key / missing nonce-store path is a fatal misconfig; no signers = nil verifier, the common slice-4 state), constructs the gate (+SlogAudit), runsRecoverbefore issuing any mutation, and routes every reconcile action through the gate.
Changed
- Memory comparison canonicalized (Note 2) —
desiredMemoryMiBmakes the desired↔actual memory compare in the same MiB unit that is then written, so a non-MiB-alignedMemoryBytesconverges in one pass instead of re-issuing SetConfig forever (the numeric cousin of the description-newline normalization). Test proves convergence. Slice 10 should still serve MiB-aligned specs at the source.
Tests (the security proof — each independently rejected)
- Adversarial matrix via the REAL
authz.Verifierwith in-test-minted SSHSIGs (framing replicated in reconcile's test binary; production authz untouched, no signing added to the verify-only package): unsigned destructive job → pending_signature; unsigned destructive desired-state delta → pending_signature (distrusts hub desired state, not just jobs); forged/unknown signer →ErrUnknownSigner; expired →ErrExpired; replayed nonce across an agent restart (durableFileNonceStore) →ErrReplay; wrong host →ErrTarget; wrong guest / wrong op / wrong params → binding_mismatch; recovery key on ordinary destructive → role_denied; hub-supplied "scratch" tag ignored → still destructive → refused; valid + role + target + fresh nonce → accepted, and a second presentation →ErrReplay(nonce consumed). - Classifier (benign/destructive/provenance/key-rotation/fail-safe), role-scoping, params binding, crash-recovery (resume OK / fail / still-running / no-task rollback / unreadable / one-shot key applied on resume), signed-job idempotency (execute once, dedupe redelivery, refused-not-executed, no-executor-inert, executor-error).
- Full module race-clean (
go test -race) + vet clean on the Linux build server.
v0.4.0-rc1 — slice 4 Phase A: reconcile engine (structural; runs live, unfed) (2026-06-08)
The agent-side control core's structural half. Checkpoint marker — -rc1 is the
Phase-A push; awaiting validation before Phase B (the reversibility gate + signed-op
consuming layer) lands the final v0.4.0. Runs LIVE but UNFED: with no desired-state
provider until slice 10, the live engine computes an empty action set and performs
zero mutations.
Added
internal/reconcilepackage — the engine, the per-guest serializer, the desired-state model, the normalization layer, and the durable op journal:- Per-guest serializer (
Queue, doc 03 §10) — the single choke point ALL mutation sources funnel through. Same-vmid jobs run strictly one-at-a-time in submit order; independent vmids run in parallel. Each vmid is a cond-var FIFO lane (unbounded, non-blocking, order-preserving); graceful drain onClose. - Desired-state model +
DesiredProviderseam —DesiredGuest(per-field optional: run-state /*hub.GuestSpec/*description),DesiredState. The only live provider isEmptyProvider(slice 4 has no source);StaticProviderfeeds fixtures. The seam is where slice 10's hub-serving plugs in — no hub/local source invented here. - Normalization layer (
FieldNormalizers) — reconcile compares normalized desired-vs-actual so Proxmox round-trip quirks don't read as drift.description's trailing newline is the first registered case; the registry takes more (boolean coercion, list ordering) as discovered.normDescpromoted out ofcmd/felhom-agent/main.gotoreconcile.NormDescription; the--selftest=taskdescription round-trip now uses that shared helper (one source of truth for the quirk). - Plan engine (
Plan, pure function) — computes the minimal benign action set (Start/Stop/SetConfig) for guests present in both desired and actual, with normalized comparison, deterministic vmid ordering, config-before-run-state. Skips provision (desired-absent-in-actual, slice 7) and destroy (actual-absent-in-desired, gated, slice 10); never writes a config it couldn't first read (SpecKnown). Disk (rootfs grow) intentionally not reconciled here. - Reconcile engine (
Engine) — reads desired+actual, plans, dispatches each action onto the shared queue. Every Proxmox op handled per the mutate.go contract: non-empty UPID →WaitTask+ assertexitstatus; empty UPID → clean synchronous success (slice-4 proven). Per-action failures are counted, not fatal (other guests still converge). - Operation journal (
Journal) — durable fsync'd append-only JSONL mirroringauthz.FileNonceStore: records each op's lifecycle (started → task_running → succeeded/failed) with its Proxmox task id (crash mid-op is detected and re-checkable on restart viaInFlight()), plus an idempotency-key store (AlreadyApplied) so a one-shot op never re-runs across retries/restarts. Reconcile actions carry no idempotency key (convergent — must re-run on real drift).
- Per-guest serializer (
- Daemon wiring (
runDaemon) — reconcile runs alongside the hub loop on the poll cadence, sharing the per-guest queue. Journal path is ajournal.logsibling of the nonce store. The daemon runs cleanly with no desired state and no signers (reconcile is a logged live no-op; a journal-open failure degrades to journal-less, never crashes).
Tests
- Serializer: same-guest serialized (max-concurrency 1, submit order preserved) and different-guests parallel (cross-waiting jobs both complete — would deadlock if not); error propagation; drain-pending-on-close; submit-after-close.
- Normalization: description round-trip; unknown-field identity; extensibility seam (synthetic boolean-coercion + list-ordering normalizers).
- Plan: run-state start/stop, spec drift (cores/memory), disk-not-reconciled, description-newline-not-drift, unmanaged fields, spec-unknown skips config keeps run-state, desired-absent skipped, combined ordering, empty-desired no-op, deterministic vmid order.
- Engine: empty-provider zero mutations; async start (WaitTask); synchronous SetConfig (no WaitTask); WaitTask failure + POST error counted failed; list error = pass failure.
- Journal: lifecycle latest-wins; in-flight survives restart; idempotency dedupe across restart; failed key not applied; torn-trailing-line skipped.
- Full module race-clean (
go test -race) on the Linux build server; vet clean.
Not in this phase (Phase B)
- The benign/destructive classifier, the reversibility gate, and the signed-op consuming
layer over
internal/authz(doc 03 §4 / doc 04) — added next, in front of the queue's executor, landing v0.4.0.
v0.3.2 — SetConfig selftest extension (slice-4 pre-check) (2026-06-08)
The gate before slice 4: prove SetConfig works live under the scoped token before
reconcile is built on it. Self-gated live run PASSED on demo-felhom/guest 9999.
Added
- Reversible
SetConfigstep appended to--selftest=task(cmd/felhom-agent/main.go,selftestSetConfig): readGuestConfig→ write adescriptionmarker (felhom-selftest <RFC3339>) → verify it landed → restore the original value (ordeletethe key if it was absent) → verify the restore. Handles PVE's dual-modeSetConfigreturn per themutate.gocontract: empty UPID = synchronous success (printedsynchronous); non-empty UPID =WaitTask+ assertexitstatus=OK. The existing snapshot → rollback → delete-snapshot steps are unchanged. First live exercise of theVM.Config.*privilege cluster. normDesc/extraStringhelpers —extraStringdecodes a string-valued key fromGuestConfig.Extra(raw JSON);normDescstrips the trailing newline PVE appends todescriptionon read, so a written value round-trips equal.
Finding (live)
- The LXC
descriptionwrite returned synchronous (empty UPID) — PVE applied it inline, no task. The agent's dual-modeSetConfigmodeling is correct: the empty-string path is real and must not be treated as an error. - PVE appends a trailing
\ntodescriptionon read (stored URL-encoded as%0A). A naive exact-match reconcile would see perpetual drift — slice-4 reconcile must normalizedescriptioncomparisons (hencenormDesc).
Ops
- Standing operator token (
felhom-agent@pve!agent, privsep) rotated during this run (the prior secret was not retrievable); role + both user/token ACL rows re-confirmed at/. New secret stored out-of-band, not persisted to the repo. Guest 9999 left pristine (stopped, nodescription, no leftover snapshot). Version → 0.3.2.
Docs + live validation — no version bump (2026-06-08)
Changed
- Reflowed
CLAUDE.md— removed hard mid-paragraph line wraps (prose, list items, blockquotes now single-line, soft-wrapped); code blocks and tables untouched; rendered output unchanged. - Unified the REPORT/CHANGELOG convention in
CLAUDE.md:CHANGELOG.mdis the cumulative log (newest on top);REPORT.mdis overwritten with the most-recent implementation/validation only. Added an explicit no-secrets rule (never write tokens/passwords/keys into committed files; reference them as stored out-of-band).
Added
REPORT.mdrewritten for the live--selftest=taskvalidation on the demo host (demo-felhom): snapshot → rollback → delete-snapshot on guest 9999, each polled toexitstatus=OKunder thefelhom-agent@pve!agentprivsep token (UPIDs name the token actor — privsep path genuinely exercised); 16-privilegeFelhomAgentrole + both user & token ACLs confirmed;--selftest=readclean. Closes the slice-1 "mutating ops unit-tested only" gap;WaitTaskasync foundation validated live → slice 4 unblocked. (Token secret stored out-of-band, not in the repo.)
v0.3.1 — slice-3 validation follow-ups (2026-06-08)
Changed
- Collector keeps the known run-status on a
GuestConfigfailure (internal/hub/collect.go): previously a per-guest config-read error forcedstatus="unknown"; now the run-status fromListLXCis preserved (only thespecis dropped). An empty status is still normalized tounknown(wire value is alwaysrunning|stopped|unknown). Test renamed toTestCollect_GuestConfigFailureKeepsStatusOmitsSpecand asserts the preservedrunning+ nil spec. --selftestusage error string now reads(want read|task|hub).
Added
- Cross-repo contract fixture
internal/hub/testdata/host-report.golden.json+TestHostReport_ContractMatchesGolden— compares the marshaledHostReportfield-name sets (top level +host+guests[0]) against the golden, failing on any json-tag drift. The file is kept byte-identical with felhom-hub's copy (duplicated contract until a shared types module; revisit when slices 5/6 populate the empty collections). Version → 0.3.1.
v0.3.0 — hub client + host-report + first daemon loop (slice 3) (2026-06-08)
The agent's first daemon: a periodic read-only host-report POSTed to the hub (the heartbeat). No Proxmox mutations, no desired-state/signed-op consumption, no storage/backup collection yet — those are slices 4/5/6.
Added
internal/hubpackage:HostReportwire contract (report.go) shared field-for-field with the hub ingest: host metrics, guests (vmid+ spec),cloudflaredstatus, and thestorage_targets/backups/restore_tests/pbs_snapshots/audit_tailcollections defined but emitted empty (typed[], slices 5/6 fill them).Collector(collect.go) builds the report from a read-onlyproxmoxReader(adapted to the realinternal/proxmoxsurface — node held by the client, value returns,proxmox.Guest) + aCloudflaredProber. Partial-failure policy: a failedNodeStatusis a hard error (skip the POST); a failed per-guestGuestConfigdegrades that guest tostatus="unknown"(spec omitted) but still sends; a cloudflared probe failure →"unknown", never fatal.CloudflaredProber+SystemctlProber(systemctl is-active cloudflared; read-only — NOT a Privileged/root op; tunnel management is a later slice).Client(client.go):POST /api/v1/host-reportwithAuthorization: Bearer <key>, standard TLS (system roots or optionalca_file; verification always on). Typed*TransportError/*HTTPError; the bearer token never appears in any error.Loop(loop.go): the daemon — immediate first report then tick; adopts the hub'spoll_interval_secondsclamped to [60,3600]; resilient (a collect/report error is logged and the loop continues); clean shutdown on context cancel.ControlEnvelope: onlypoll_interval_secondsis acted on;blocked/desired_generation/has_signed_opsare parsed-but-ignored (logged at most) pending reconcile (slice 4).
- Config:
HubConfig(url/host_id/api_key/poll_seconds/timeout_seconds/ca_file),FELHOM_AGENT_HUB_*env overlay,HubConfig.Validate()(mode-aware — proxmox-only--selftest=read|taskstill runs without hub config),WithDefaults(), andRedacted()now also blanks the hub key.configs/agent.example.jsongainshub(andauthz) blocks. cmd/felhom-agent: the no---selftestmode is now the daemon (poll loop); added--selftest=hub(one collect+report, prints the report + envelope). Version 0.2.0 → 0.3.0.
Tests
- Report serialization (field names; empty collections are
[]notnull; spec omitted when unknown); client (Bearer header, non-2xx→*HTTPError, transport→*TransportError, token never in error); collector (host mapping, guest spec, per-guest failure degrades-but-still-reports, NodeStatus hard error, cloudflared error→unknown); loop (immediate first report, continuation after an injected error, interval adoption + clamp); config (hub validate/redact/env).
Notes
internal/proxmoxandinternal/authzwere not touched — no new proxmox surface was needed (ListLXCalready exposes status/maxmem/maxdisk;GuestConfigexposes cores). The task'sproxmoxReadersketch (node-arg/pointer/LXC) was adapted to the real exports as instructed.- Defined-but-empty this slice:
storage_targets,backups,restore_tests,pbs_snapshots,audit_tail(slices 5/6). Parsed-but-ignored: the envelope'sblocked/desired_generation/has_signed_ops(slice 4).
v0.2.0 — authz signed-op verifier (slice 2) (2026-06-08)
Production form of the Phase-4 signing primitive: a key-type-agnostic SSHSIG verifier for operator-signed destructive ops, with the full anti-replay/ authorization pipeline and a durable, crash-safe nonce store. What slice 4 (reconcile) will call to gate destructive desired-state deltas. No hub, no signing CLI, no reconcile loop.
Added
internal/authz—Verifier:New(signers, store, hostID)+Verify(blob, sigArmored) (*VerifiedOp, error). Runs the LOCKED pipeline (order is load-bearing): parse armor → namespace → parse pubkey → allow-list (by key material,pub.Marshal()equality, not key_id) → crypto verify (over the raw received bytes, never re-canonicalized) → parse blob → target → time window → nonce recorded LAST. Each post-crypto stage rejects even with a valid signature.- SSHSIG framing (
sshsig.go) viagolang.org/x/crypto/ssh—pem.Decode→ strip 6-byte magic →ssh.Unmarshal→ssh.ParsePublicKey→ recompute signed data with the named hash →pub.Verify(dispatches on key algorithm). No hand-rolled crypto. Key-type-agnostic: ed25519 / sk-ssh-ed25519 (FIDO2) / rsa / ecdsa via the one path. - Fixed namespace
felhom-op-v1(package constant, never caller-supplied). OpBlob(correctedhost_id/guest_idjson tags) +VerifiedOp(op, host/guest, params, key_id, matched signer). key_id is advisory/audit only — never an authz input.- Typed errors:
ErrMalformed, ErrNamespace, ErrUnknownSigner, ErrBadSignature, ErrTarget, ErrExpired, ErrNotYetValid, ErrReplay(errors.Is-friendly). NonceStore+ two impls:MemoryNonceStore(tests) andFileNonceStore— durable, crash-safe (fsync'd append log, replayed into an index on open, periodic compaction, expiry-only pruning). A nonce is fsync'd to disk beforeSeenOrRecordreturns false; replay protection survives restart; I/O failure fails safe (reports seen=true). Target generalization: host_id matched strictly, guest_id surfaced for the caller to route.- Config:
AuthzConfig(nonce-store path + pinned operatorsignerstaggedoperational/recoverywith a key_id, as authorized_keys lines). - Version 0.2.0.
Tests
- Real OpenSSH interop via a committed
ssh-keygen -Y signvector (hermetic CI); per-stage rejection (each with an otherwise-valid sig); the headline invalid-sig-does-not-burn-the-nonce invariant; replay; persistence across restart; synthetic sk-ssh-ed25519 through the unchanged path; byte-exactness (a re-serialized blob fails crypto — not re-canonicalized).
Notes / corrections to the Phase-4 reference
- §7's
Targetlacked json tags (host_id/guest_id) — fixed. - The doc paired "Go 1.24.4 / x/crypto v0.52.0", but v0.52.0 declares
go 1.25.0and does not build on Go 1.24. Resolved by upgrading the build server to go1.26.0 (backward-compatible; felhom-controller/hub unaffected); the module isgo 1.25.0on x/crypto v0.52.0. - Free function → constructed
Verifier; returns the fullVerifiedOp; typed errors; clock-skew tolerance added; durable nonce store is the net-new work. - Shared-contract dependency flagged (not built): the hub and the
felhom-signCLI must emit byte-identical canonical JSON or signatures won't verify; a shared canonicalizer both import would be the right home.
v0.1.0 — Scaffold + proxmox interaction layer (slice 1) (2026-06-08)
First slice: stand up the host-agent project and its foundation — the typed Proxmox interaction layer every other module will call. No reconcile loop, hub client, signing, or storage/backup orchestration yet (later slices).
Added
- Project scaffold: module
gitea.dooplex.hu/admin/felhom-agent, binaryfelhom-agent(cmd/felhom-agent/), Go 1.24, zero external dependencies (pure stdlib).--versionflag;versionvar overridable via-ldflags "-X main.version=<v>". internal/proxmox— API backend (Client): hand-rolled REST client overhttps://<host>:8006/api2/jsonwithPVEAPITokenauth. Typed read ops (Version,Nodes,NodeStatus,ListLXC,GuestStatus,GuestConfig,ListStorage,NodeStorage,StorageContent) and async mutating ops returning a UPID (RestoreLXC— the primary create path,Vzdump,Snapshot,Rollback,DeleteSnapshot,SetConfig,Start,Stop).WaitTask: pollsGET /nodes/{node}/tasks/{upid}/statusuntil stopped, then assertsexitstatus == "OK"(authorization can surface at task execution, not the POST — phase1-2 §1.3). Exponential backoff (1s→5s cap), context cancellation + timeout.*APIErrorparses the offending privilege from a 403;*TaskErrorparses it from a failed task exitstatus + log tail.internal/proxmox— fenced root-CLI backend (Privileged): limited to the three proven OS-root exceptions only —CreateGoldenLXC(keyctlpct create),MountUSBByUUID,SMART,Sensors; each cites why it can't be the API. Fence is structural (Client never shells out, Privileged never makes an HTTP call) and asserted in tests.- TLS trust: SHA-256 leaf-cert pinning (the host serves a self-signed cert) or
a CA file; an explicitly-named
insecure_skip_verifythat is off by default. No blanket verification disable. internal/config: JSON config file +FELHOM_AGENT_*env overrides; the token secret is never logged (Redacted()).internal/log: slog setup (text, stderr, configurable level).cmd/felhom-agent --selftest: read-only health report against a live host (version/nodes/status/guests/storage);--selftest=task --vmid NexercisesWaitTaskon a reversible snapshot→rollback→delete op (gated; default selftest mutates nothing).- Tests: unit tests with a mock HTTP transport + mock runner (UPID parse,
WaitTaskrunning→OK / failed-403 / timeout / ctx-cancel, 403→privilege error, response decoding against shapes captured live fromdemo-felhom, config redaction, and the API-vs-root routing fence).
Notes
- Types are grounded in the spike findings
(
felhom.eu/documentation/proxmox-platform.md,tests/phase{0,1-2,3}-findings.md) and the exact JSON shapes captured live fromdemo-felhom(PVE 9.2.2). - Verified:
go build/vet/testgreen on Go 1.24.4 (build server) and a live read-only--selftestagainst the demo host with TLS fingerprint pinning. - The 16-privilege
FelhomAgentrole + privsep token (role on both user and token) is provisioned out-of-band; the agent only consumes the token.