staleLockController.Guests() = ListLXC ∩ GET /pools/felhom members (ownership PROVEN via the pool registry, never assumed from enumeration scope); pool-read failure fail-safes the whole recovery through the existing guest-list guard. New Client.Pool read (needs Pool.Audit — host-install v1.9.0; Pool.Allocate does NOT satisfy it, spike T2). Composed pve:pool-read capability (non-critical) + --selftest pool-read line. Red-proofed negative tests drive the REAL controller over a broad-token-shaped fake. Per SPIKE-a1-pool-membership-read-2026-07-03.md; audit A1 (AUDIT-blast-radius-hostroot-localapi-2026-07-02). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PSK5g6qYLknKj8u3QAFEr6
181 KiB
v0.62.0 — A1: pool-membership ownership check for the stale-lock reaper (2026-07-03)
Implements audit finding A1 (AUDIT-blast-radius-hostroot-localapi-2026-07-02 §A) per the spike
verdict (SPIKE-a1-pool-membership-read-2026-07-03 — enumeration is pool-filtered under the scoped
token, so this is defense-in-depth: a future broad-token deployment can no longer re-arm the reaper
against co-tenant guests). Companion: host-install v1.9.0 (Pool.Audit added to
FelhomAgentGuest) — rescope BEFORE deploying this agent, else the reaper fail-safes (skips)
until the ACL catches up.
Client.Pool(proxmox/query.go):GET /pools/{name}→PoolInfo{PoolID, Members[]{VMID,Type}}. RequiresPool.Auditat/pool/{name};Pool.Allocatedoes NOT satisfy the read (spike T2).staleLockController.Guests()(localapi/stalelock.go): now returnsListLXC ∩ pool members(nonzero-vmid, non-storage entries only). Ownership is PROVEN via the pool registry, never assumed from enumeration scope. A pool-read failure returns a wrapped error ("pool membership read (pool=felhom): …") that rides the existing "guest list unavailable — skipping recovery" guard — fail-safe: NO unlock/snapshot-delete/start on ANY guest, never a fallback to the unfiltered list. One new INFO line per scan:stale-lock: scanning pool guests(pool, listed, scanned) — emitted by the controller (the unchangedStaleLockControllerseam can't carry the pre-intersect count).NewStaleLockControllergains(pool string, logger *slog.Logger); main.go threadsreconcile.DefaultPool.- Capability surfacing: the hub-report prober is now a composed closure — the sudo manifest
probe + one
pve:pool-readstatus (non-critical; degraded ⇒ reaper is fail-safed, visible on the report, no operator page). Composed in main.go;internal/capability/untouched. --selftest: new "pool read" line (pool id + member count + guest members).- Tests (stalelock_pool_test.go, driving the REAL controller over a broad-token-shaped fake):
TestStaleLock_ForeignGuestNotReaped(red-proved: intersect removed ⇒ FAILS withpct unlock 5000recorded),TestStaleLock_PoolGuestStillReaped(anti-over-filter),TestStaleLock_PoolReadFails_SkipsAll(red-proved: fallback-to-unfiltered ⇒ FAILS with mutations recorded),TestStaleLockController_GuestsIntersect(storage-member + empty-pool edges). The 9 existing Server-level stalelock tests pass unmodified.
docs — CLAUDE.md refresh: stable orientation, complete layout (2026-07-03)
No code change, no version bump. Deleted the version-pinned "Current: v0.31.0" narrative and the
per-slice history (stale by 30 versions — current state lives in CONTEXT.md/CHANGELOG top); Layout
completed with the 8 missing packages (capability, desired, escrow, guesthook, lanresolver, localapi,
provision, signedjobs + cmd/felhom-opsign — verified against the tree); build/deploy compressed to a
summary table pointing at the felhom-build-deploy skill (commands verified live on felhom-pve:
non-root felhom-agent user, /usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json).
Load-bearing Proxmox-model rules kept verbatim. Standing rule: no version-pinned state in CLAUDE.md.
docs — REUSE.md introduced (2026-07-03)
Cross-repo reuse-map rollout (docs-only, no code change, no version bump). New REUSE.md at the
repo root: curated map of canonical helpers (48 rows — exec/sudoers surface, format-safety guards,
durable-id seams, stores, local-API plumbing), patterns, dangerous lookalikes (req.Device TOCTOU,
raw mkfs, uuid:-vs-byid: scheme confusion, MemoryNonceStore, pool-blind stale-lock scan…), test
seams, extension points, and observed duplication (7 clusters, NOT fixed). Every entry code-verified
at file+symbol; cited paths machine-checked by felhom.eu/scripts/reuse_refs_check.py (green).
CLAUDE.md gains the "See REUSE.md before writing new code" pointer + the same-commit maintenance rule.
v0.61.0 — blast-radius audit fixes B1 + D1 + D2 + D3 (2026-07-03)
Four LOW/INFO fixes from AUDIT-blast-radius-hostroot-localapi-2026-07-02.md — the "second gate must
mirror the first" batch. Each shipped with a non-hollow test AND a companion red-proof (test shown
failing on the pre-fix implementation). A1 (stale-lock pool-membership) is deliberately NOT here — it
needs a pool-read spike (the role lacks Pool.Audit); C1/C2/A2/B2–B5/E1/E2 deferred.
- B1 (LOW) — random temp staging for root-installed scripts.
internal/guesthook/install.go(InstallSnippet) andinternal/localapi/intermediary.go(installSharedParentUnit, new sharedstageTemp) staged root-executed scripts through FIXED, predictable/tmpnames viaos.WriteFile(no O_EXCL, follows symlinks) — a local TOCTOU into a root-run PVE hookscript / boot script. Both now use theos.CreateTemprandom-name patternlanresolveralready used.configs/felhom-agent.sudoers(FELHOM_GUESTHOOK/FELHOM_INTERMEDIARY) install-SOURCE grants became globs (/tmp/felhom-guest-hook-*.sh,/tmp/felhom-shared-parent-*.{sh,service}; destinations stay pinned);internal/capability/manifest.gorepresentative vectors updated to match. Tests:TestInstallSnippet_RandomTempName,TestInstallSharedParent_RandomTempName(fake runner records the install source: random pattern, two calls differ, content + cleanup asserted). - D1 (LOW) — the guarded mkfs wrapper now mirrors the classifier's member/RO classes.
configs/felhom-mkfs-guarded.shre-checked only system-disk / LVM-PV (PATH-dependentcommand -v pvs) / foreign-mount — a bypassed agent could mkfs a ZFS/mdraid/LUKS/swap member or a read-only disk. Added (additive; nothing removed/reordered):/sys/block/<disk>/ro== 1 → die; an lsblk-FSTYPE loop over the whole disk refusing exactlyclaim.go'smemberFSTypes(LVM2_member/zfs_member/linux_raid_member/crypto_LUKS/swap — the FSTYPE catch works with pvs absent); pvs resolved via absolute candidates (/usr/sbin/pvs, /sbin/pvs). Validated by the newscripts/mkfs-guarded-harness.shon felhom-pve: throwaway loop devices + PATH-shimmed lsblk + a recorder bind-mounted over mkfs.ext4 in a private mount namespace (no real mkfs possible) — fixed wrapper 8/8 (incl. plain-blank-disk still formats); pre-fix wrapper red-proof: 7/8 hostile fixtures reached mkfs. - D2 (INFO) —
classifyClaimempty-lsblk fail-safe. A successful-but-emptylsblk({"blockdevices":[]}) skipped the member/mount loop and returnedunclaimed.internal/storage/claim.gonow refuses when the node tree is empty OR the target whole-disk is absent from it (undeterminable topology ⇒ claimed). Tests:TestClassifyClaim_EmptyNodesRefused,TestClassifyClaim_TargetAbsentFromTree. - D3 (INFO) — blank-format anti-retarget (AGENT-001's benign-branch twin).
handleDiskFormat's blank branch formatted the mutable caller-suppliedreq.Devicewith no durable-id binding — a /dev re-enumeration between inspect and mkfs could format a data-bearing disk that inherited the node. The blank branch (internal/localapi/disks.go) now derives the device's durable id (no durable id ⇒ 409 refuse — path-only formats are not permitted), re-resolves it via the newantiRetargetResolveBlank(wipe_reresolve.go: sharedantiRetargetResolveExpectcore; the blank variant asserts the device is STILL !DataBearing), and formats the RE-RESOLVED device. The format job record (formatjob.go) carriesblank; restart recovery re-checks blank jobs with the blank variant (durable-id-bound, fail-safe refuse). Confirmed/data-bearing branch untouched. Tests:TestFormatBlankPath_AntiRetarget_{ReassignedDataBearingRefused,ReassignedDifferentDiskRefused,UnresolvableRefused,SameBlankProceeds},TestFormat_Blank_{FormatsReresolvedDeviceNotCallerPath,ReresolveRefusalNoMkfs,NoDurableIDRefused}. - Deploy note: the host's
/etc/sudoers.d/felhom-agentMUST be updated together with the v0.61.0 binary (the old fixed-name grants deny the new random-name installs, and vice versa).
v0.60.0 — proof-of-launch destroy gating + restore-test band-advance (campaign F1/F2) (2026-07-02)
Fixes the pool-effects campaign's HIGH finding (F1, CAMPAIGN-pool-effects-2026-07-01.md): the bring-up
compensating rollback and the restore-test teardown fired DestroyLXC on the target vmid even when
RestoreLXC failed synchronously WITHOUT creating anything (PVE refusing a pre-existing vmid the
pool-blind duplicate guard / band scan couldn't see) — destroying a guest the transaction never made.
Only the pool ACL's 403 saved the non-pool subset; an in-pool pre-existing guest would have been
destroyed, and any broad-token deployment re-arms the bug. Root cause: SameTxnCreated/scratch provenance
was ASSUMED, never verified against proof-of-launch. The fix makes a RestoreLXC UPID the sole destroy
authorization, in ALL THREE destroy paths — the pool ACL is defense-in-depth again, not the guard.
- F1a
internal/reconcile/bringup.gorunBringUp: the compensating-rollback defer is gated onlaunched(set only after the restore POST is accepted). A synchronous restore failure (no UPID) closes the owning entry terminal-failed WITHOUT any destroy. The pre-restoreOpStartedjournal append is kept (crash-safety);rollbackBringUpis now only ever called launch-proven. - F1b
internal/reconcile/restoretest.gorunScratchTest: samelaunchedgate onteardownScratch— a synchronous restore refusal never destroys the picked band vmid. - F1c
internal/reconcile/recover.goRecover: the no-UPID "POST never confirmed → abandon fail-safe" check now runs BEFORE the Scratch/Rollback dispatch — a no-UPID Scratch/Rollback entry is abandoned (marked failed, NO destroy) instead of destroy-by-vmid-existence. Recover is now safe by DESIGN, not by the pool-blind "already gone" accident the campaign observed. - F2
restoretest.goRunRestoreTestband-advance: a band vmid PVE refuses with "already exists" (an invisible squatter — newpveAlreadyExists, mirrorspveConfigLock, never misclassifies a real restore failure) is skipped and the next free band vmid tried (bounded by the band width;pickScratchVMIDgained an exclude set). A fully-occupied band →Skipped(scheduler raises no "backup unrestorable" alert), never FAIL — one squatter no longer permanently breaks the restore-test. - Accepted residual (by design): a crash in the one-statement window between obtaining the UPID and journaling it leaks a half-built guest Recover won't destroy — cleanable, and vastly preferable to destroying an innocent guest.
- Tests: red-proof companions verified (gates reverted →
TestRunBringUp_NoLaunchNoDestroy,TestRunRestoreTest_RestoreNoLaunchNoTeardown,TestRecover_{BringUp,Scratch}NoUPIDAbandonedall fail with the innocent-guest destroy); no-regression…LaunchedTaskFailureStillTearsDown+ rollback table now includes an explicit restore-task-failure case; F2 advance + squatter-full-band-skips tests. Live-validated on felhom-pve (provision onto existing 9001 → no destroy armed; restore-test advances past a 990000 decoy).go build/vet/test ./...clean.
v0.59.0 — report backing device + capacity for a registry-sourced drive in /disks (2026-07-01)
Completes the /disks representation for a registry-sourced (raw, no-PVE-storage) drive: the agent-view
showed "—" for the device and no size bar, because the union row never populated backing_device or
total_bytes/used_bytes (Observe drives get those from pvesm status, which a raw drive has none of).
internal/localapi/disks.gohandleDisks registry union: resolveBackingDevicefrom the fs-UUID (storage.ByUUIDDevicePath) and read capacity viastatfsCapacity(new build-taggedcapacity_linux.go=syscall.Statfson the mount;capacity_other.go= no-op for dev builds).go build/vet/test ./...clean (Linux + Windows dev). Live: the registry drive now shows its device + size in the agent-view, matching the Observe-sourced drives.
v0.58.0 — report GuestPath/BoundUnderParent for a registry-sourced drive in /disks (2026-07-01)
Last piece of first-class raw-drive support: the /disks union row for a registry-sourced drive (Impl-2a
— a drive with no PVE storage) omitted GuestPath + BoundUnderParent, so the controller read it as
"Leválasztva" (disconnected) even though it was mounted + bound + live in the guest.
internal/localapi/disks.gohandleDisks registry union: populateGuestPath(StablePathForRaw)BoundUnderParent(boundUnderParent) on the registry row, identical to the Observe path — so a registry-only drive reports its true bound/active state.go build/vet/test ./...clean.
v0.57.0 — re-assert a RAW drive's guest-bind (ReassertGuestBinds mount-table fallback) (2026-07-01)
Completes the raw-drive durability the v0.56.0 fix started. ReassertGuestBinds (the startup / drive-
returned reconcile that re-binds an enrolled drive's felhom-data under the shared parent so it's live in
the guest) built its durable-id→mount map from Observe() only — so a RAW enrolled drive was never found
("enrolled drive not present"), and its in-guest bind was not re-asserted after a reboot or a watchdog
re-mount (the drive would show "Leválasztva" in the controller).
internal/localapi/disks.goReassertGuestBinds: augment the durable-id→mount map from the mount table — each raw/mnt/<name>mount → its device fs-UUID (HostReader.Mounts+ResolveUUID), skipping the/mnt/felhom-drivesbind (AttachDrive wants the raw path). Observe entries still win. Also: an Observe failure is no longer fatal (fall through to the mount-table scan) so raw drives re-assert even if the PVE view is momentarily unavailable.- Wiring fix (latent):
buildLocalAPIServernever passedOptions.HostReader, so the local-API server'shostwas nil in production — the v0.56.0durableIDForMountraw fallback (and the role gate's host classification) silently no-op'd. Now wired tostorage.NewProcHostReader(). This is what makes the v0.56.0 + v0.57.0 raw-mount resolutions actually fire live. go build/vet/test ./...clean. With v0.56.0 (guest-bind now RECORDED for raw drives) this closes the reboot/reconnect guest-bind durability gap for raw drives end-to-end.
v0.56.0 — record intent/guest-bind for a RAW enrolled drive (durableIDForMount fallback) (2026-07-01)
Surfaced by the first live raw enrollment (Impl-2b): a raw drive is not a PVE storage, so
durableIDForMount (Observe-based) returned "" for it → the enroll's intent + guest-bind recording
logged "durable-id unresolved" and silently skipped. Result: the drive mounted + bound + usable, but was
NOT intent-tracked (so RegistryKnownTargets — which gates on intent ≠ new — didn't health-track it) and
its guest-bind wasn't persisted.
internal/localapi/disks.godurableIDForMount: after the Observe lookup, fall back to resolving the fs-UUID directly from the mount table — the device mounted atwhere(viaHostReader.Mounts) → its by-uuid identity (HostReader.ResolveUUID) →uuid:<fs-uuid>(the SAME scheme Observe derives, so intent keys stay consistent). Fixes intent recording (enroll/eject) AND guest-bind recording for raw drives; the PVE-storage path is unchanged.- Test
TestDurableIDForMount_RawFallback(+ red-proof: Observe-only → "").go build/vet/test ./...clean. - (Residual noted here fixed in v0.57.0:
ReassertGuestBindsraw-drive guest-bind re-assert.)
v0.55.0 — raw-device discovery + registry-sourced drive tracking (Impl-2a) (2026-07-01)
Agent backend for drive enrollment (SPIKE-drive-enrollment §SQ1/SQ4/SQ5). Makes raw (non-PVE-storage) drives (a) discoverable for enrollment and (b) health-tracked WITHOUT being a PVE storage — so a drive enrolled the new way isn't enrolled-but-untracked (the 3b-fix false-detach class). No mkfs here (Impl-1 owns it); the controller wizard rewiring is Impl-2b.
GET /disks/candidates(internal/storage/candidates.go+internal/localapi/disks.go): enumerates host whole-disks from/sys/block, runs the Impl-1 unclaimed filter, and returns the free ones with probe info (size/model/FS/data-bearing/durable-id preview), split intoinitialize(all unclaimed) andattach(the subset carrying a mountable ext4/xfs FS). Fail-safe carries through (a device not provably unclaimed is omitted).RegistryKnownTargets(internal/storage/registry_known.go): the watchdog's known-DRIVE set now comes from the intent registry + Felhom.mountunits, NOTObserve()(PVE storages). A unit is tracked iff its intent ≠new(enrolled/ejected/decommissioned; the watchdog's existing IntentReader gate still decides re-mount).main.goswaps the watchdogKnownTargetssource.Observe()is KEPT for real PVE storages (local/local-lvm/pbs) — reports + the/disksview.handleDisksunion: additive + deduped-by-mount-path — appends registry drives Observe doesn't surface (a registry-only drive now appears in the agent-view) without dropping any Observe row (can't regress the current view).- Existing-drive migration (
ReconcileExistingDrives, idempotent, at agent start): records each currently-mounted Felhom-unit drive asenrolledso the registry-sourcedKnown()tracks it without its legacy PVE dir-storage. Does NOT create/remove PVE storages. - Tests: RegistryKnownTargets (enrolled tracked / new excluded / ejected tracked) + the red-proof
(Observe-based
Known()misses a drive with no PVE storage; the registry provider tracks it); migration idempotency; candidates init/attach split.go build/vet/test ./...clean. Watchdog/ HostLiveness/Remounter unchanged (only the source injected).
v0.54.0 — format-safety foundation: unclaimed-disk guard + guarded-mkfs wrapper (2026-07-01)
Impl-1 (SPIKE-drive-enrollment-2026-07-01). Hardens the destructive Format/mkfs path BEFORE the
enrollment feature: today Format delegates authorization to its caller and only checks DataBearing
(has-data), which is insufficient — the OS disk is data-bearing yet catastrophic — and the sudoers
permits mkfs /dev/*. Two independent, layered guards:
- Part A — mandatory unclaimed-disk guard inside
Format(the primary safety). Newinternal/storage/claim.go:classifyClaim(pure) +gatherClaimFactsrefuse to format any device not provably UNCLAIMED — reusingSystemDisks(OS disk) + lsblk FSTYPE (LVM2_member/zfs_member/ linux_raid_member/crypto_LUKS/swap) + foreign mounts + read-only + authoritativepvs/zpool. A Felhom-owned mount under/mnt/felhom-drivesis NOT a foreign claim (re-init stays allowed; the DataBearing wipe-confirm still gates data loss). FAIL-SAFE: any read error / undeterminable topology → CLAIMED → refuse. The guard is inFormat(lowest layer), not the handler, so no caller can bypass it. Read-only sudoers additions:pvs,zpool status(inFELHOM_DISK). - Part B — guarded-mkfs wrapper below the agent.
configs/felhom-mkfs-guarded.sh(root, 0755) is now the ONLY mkfs path the sudoers allows (FELHOM_FORMATno longer allowlists rawmkfs.*). It re-checks the cheap catastrophic cases (system disk / LVM PV / foreign mount) and refuses — so even an agent bug/compromise can't mkfs the OS disk.Formatexecs the wrapper (<device> <fstype>) viaBinaries.MkfsGuarded. - Tests:
claim_test.go— table-drivenclassifyClaim(every claim signal + fail-safe + the two allow cases) incl. the red-proof (a claimed, non-data-bearing OS disk: removing the isSystem check flips it to allowed → test fails, proving the guard adds safety beyondDataBearing); Format-guard integration tests (refuses system disk / LVM member, allows unclaimed → wrapper invoked); capability manifest updated (mkfs sample → the wrapper).go build/vet/test ./...clean. - Deferred to Impl-3: a raw disk passed through to ANOTHER VM looks unused to the host — a host-level filter can't detect it; the operator gate (shared-box mode) closes that. Impl-1 closes everything host-visible (a strict improvement over today's no-guard state).
v0.53.0 — restore guests INTO the felhom pool (pool-scoped-ACL enabler) (2026-07-01)
Colleague-safety batch #4 phase b (agent half). Enables the agent token to be scoped from / to
/pool/felhom + /storage/<targets> (real blast-radius containment on a shared host) by making every
restore allocate the guest INTO the pool — the only way a fresh vmid authorizes under a pool-scoped
token. Grounded by felhom.eu/documentation/audits/SPIKE-pool-scoped-acl-2026-07-01.md (PASS).
internal/proxmox/mutate.go:RestoreLXCOptionsgainsPool string;RestoreLXCsendspool=<p>only when non-empty (pct restore --pool). Omit-when-empty (a broad-token restore needs no pool) — unit-tested + red-proofed.internal/reconcile: newconst DefaultPool = "felhom"(single source of truth);BringUpSpecgainsPool, threaded to the bring-up restore. Both restore sites now pool the guest: the provision/DR bring-up (Pool: spec.Pool, set toDefaultPoolby the CLI) AND the restore-test scratch guest (Pool: DefaultPool) — the latter closes SPIKE residual #2 (a pool-scoped token would otherwise 403 on the out-of-pool scratch guest).cmd/felhom-agent/main.go: bothBringUpSpecliterals (bring-up/DR + provision) setPool: reconcile.DefaultPool.- No ACL/priv change in the agent — that ships in the host-install script (v1.6.0). The pool param
is INERT until the token is granted
Pool.Allocateat/pool/felhomand the pool exists; publishing is therefore safe ahead of the coordinated ACL swap. - New tests:
proxmox.TestRestoreLXC_PoolParam(set →pool=felhom; empty → omitted),reconcile.TestRestoreSitesUsePool(both restore sites carryDefaultPool). Both red-proofed (unconditionalSet→ omit test fails; drop either site'sPool→ both-sites test fails).go build/vet/test ./...clean.
v0.52.0 — operator-opt-in CPU/RAM cap for the provisioned guest (-cores / -memory) (2026-07-01)
Colleague-safety batch #3. So a trial appliance guest on a colleague's SHARED production Proxmox does not pressure his existing guests, the operator can now cap the guest's CPU cores + RAM at provision time, before the guest's first boot (the peak container-pull moment). Pure CLI→spec plumbing — the reconcile engine already applied the cap; this only wires the flags to it.
cmd/felhom-agent/main.go: new-cores N/-memory M(MiB) flags for--selftest=bring-up|provision(0 = keep the golden's baked size). They flow throughbringUpSizing(now carriesCores/MemoryMB) into thereconcile.BringUpSpec{Cores,MemoryMB}built by BOTHrunSelftestBringUpandrunSelftestProvision. The-selftestusage string documents them.- No engine change.
internal/reconcile/bringup.goalready carriesBringUpSpec.Cores/MemoryMB(0 = leave as restored) andbuildBringUpConfigalready emitscores/memoryinto the SAME coalesced config PUT as the identity reset — which runs BEFOREe.api.Start, so the cap lands pre-boot. Rejected thepct set-post-provision alternative (runs after boot = an uncapped window; bypasses the token/audit; second config source of truth). - Omit-when-zero guarantee: an unset cap (0) emits NEITHER
coresNORmemory, so an uncapped provision keeps the golden defaults (no regression for the normal single-purpose box) and can never shrink the guest to 0 cores. New pure-function testTestBuildBringUpConfig_ResourceCapsasserts both the set (cores=2,memory=4096) and the absent-when-unset cases; a red-proof (unconditional emit) was run and confirmed to fail the omit assertion, then reverted. - Deploy dependency: a FRESH host-install
--cores/--memory(felhom.eu script v1.4.0) requires the hub artifact manifest to serve agent ≥ v0.52.0, else the old agent rejects the unknown flag. The flags are opt-in, so nobody hits this until they intentionally cap. go build/go vet/go test ./...clean.
v0.51.0 — local vzdump retention default (--prune-backups keep-last=3) (2026-06-30)
The PREVENTIVE counterpart to the hub's host_disk + storage_fill detectors: the agent's periodic local
whole-guest vzdump now prunes its own old archives, so a box can't refill its own root via its own backups
(the felhom-pve incident's root cause — that vzdump carried no retention, ~18 dumps piled under
/var/lib/vz/dump).
internal/proxmox/mutate.go:VzdumpOptions.PruneBackups→ passed as PVE's--prune-backupson the vzdump POST (vmid+storage scoped, so PVE prunes only THIS guest's archives on THIS storage).internal/backup/runner.go:NewBackupRunnergains aretentionarg; the backup applies it vialocalPruneSpecONLY when the target is a non-PBS storage (resolved viaListStorage) — PBS offsite retention is a separate lifecycle and is never pruned by the per-run flag. Fail-safe: if the target type can't be confirmed (lookup error / not found) the run SKIPS pruning rather than risk pruning PBS (the detectors remain the safety net). Only the periodic local-API runner sets retention; the restore-test / selftest runners pass "".internal/config/config.go:backup.local_backup_retention(keep-last N) withKeepLast()clamped to ≥1 (0/unset/negative → default 3) — a mis-config can NEVER prune the just-made backup —PruneBackupsSpec()→keep-last=N. Wired into the local-API backup runner (main.go).
- Seeding:
felhom.eu scripts/felhom-host-install.shseedslocal_backup_retention: 3in the agent config; the code default also protects any box where it is unset (KeepLast → 3) from day 0. - F2-b stale-vzdump-lock recovery untouched.
- Tests: the local vzdump carries
--prune-backups keep-last=3(+ companion: no-retention runner emits no prune); PBS is never pruned (+ companion: same retention on a local target IS applied); fail-safe-on-unknown-target; the keep-last≥1 clamp companion.go build/vet/test ./...green.
v0.50.0 — NAS network storage Part A1: NFS/SMB automount foundation (2026-06-30)
Agent foundation of the validated SPIKE-nas-storage-2026-06-29.md (verdict READY): a customer NAS can
serve bulk media to a media app. The agent mounts a NAS share host-side under
/mnt/felhom-drives/<name> via a systemd .automount (+ .mount) pair; it propagates into guest 9201 for
free through the existing shared mp8 bind (no new mountpoint, no restart). A NAS is a distinct storage
class — it carries no durable-id and never enters the drive enroll/eject/decommission/wipe/SMART/
watchdog machinery. Bulk-media class only; STOP before A2 (controller registry/UI) + B (restic-SFTP).
internal/storage/netmount.go(NEW).NetworkMountSpec+ the locked SPIKE recipe:- NFS (preferred):
What=server:/export,Type=nfs4,Options=vers=4.1,soft,timeo=50,retrans=2,noatime,_netdev.softis the failure-isolation knob (clean EIO, never adf/guest wedge); a defaulthardmount is never emitted. The+100000uid mapping is the export's job (anonuid=101000), so the client mount carries no uid. - SMB (fallback):
What=//server/share,Type=cifs,Options=vers=3.0,credentials=<0600 file>,uid=<+100000>,gid=<+100000>,forceuid,forcegid,file_mode=0664,dir_mode=0775,_netdev(plain octal modes, never setgid 2775). The +100000 rule (container uid/gid N = host N+100000): a container uid 1000 rendersuid=101000so the guest sees its native id and reads+writes; a naïve+0lands asnobody:nogroup(not writable) — the documented trap, asserted by a companion test. .automountwithTimeoutIdleSec(on-demand + idle-unmount): an idle NAS reboot is a non-event.- per-share liveness (
ListNetworkMounts): TCP-probes the NAS endpoint (2049/445) + reads/proc/mounts— it neverstats the (possibly EIO/D-state) mountpoint, so a black-holed NAS cannot wedge a list. Healthok | idle | unreachable, scoped to the affected share, never box-wide. - role gate
NetworkMountRole: network storage is bulk-userdata only — confined to the/mnt/felhom-drivesnamespace; any other target is refused (most-protected). - Full validation (
ValidateNetworkMountSpec) before any unit is rendered: share name (safe segment), server, NFS export (absolute, no traversal) / SMB share name, uid/gid range, creds path.
- NFS (preferred):
- Drive-machinery bypass (Scenario D).
parseFelhomMountUnit(the host-reboot drive re-assert's classifier) explicitly refuses any unit carrying the network marker, so a NAS mount is never given a durable-id, SMART-probed, or re-asserted as a drive. Companion red-proof: the same by-uuid-shaped unit with the drive marker DOES parse — the guard is the discriminator, not luck. internal/localapi/netstorage.go(NEW). Self-scoped endpointsPOST /netstorage/add,GET /netstorage,POST /netstorage/remove. SMB credentials are written out-of-band to a 0600 file the agent owns (never in git, never in a plaintext registry, never logged). Role-gated to the user-data namespace.- sudoers: new narrow
FELHOM_NETMOUNTalias (install/enable/disable/stop the.automount+ remove the felhom mount-unit files; the.mounthalf reusesFELHOM_MOUNT, the mountpoint mkdir reusesFELHOM_INTERMEDIARY).visudo -cfclean. - config:
privileged.smb_creds_dir(default/var/lib/felhom-agent/smb-creds). - Runtime deps:
mount.nfs(nfs-common) +mount.cifs(cifs-utils) present on the host (confirmed live). - Tests: exact NFS/SMB option-set string-asserts + the +100000 companion; validation matrix; role gate; unit round-trip + health; the drive-machinery guard + companion; Ensure/Remove command sequences.
v0.49.0 — reboot-during-backup stale-lock recovery (F2-b) + shared-parent script redeploy fix (F2-a) (2026-06-30)
Closes the two host-reboot findings from TESTRUN-fullstack-2026-06-29.md.
-
F2-b — startup stale-lock recovery (
internal/localapi/stalelock.go, NEW). A host reboot DURING a vzdump backup leaves the guest with asnapshot-delete/backuplock + a danglingvzdumpsnapshot;onboot:1then can't start the locked CT → the customer box stays DOWN until a human runspct unlock. The agent now self-heals at startup (Server.RecoverStaleLockedGuests, called alongsideReassertGuestBinds/RecoverFormatJob): for each guest carrying a backup lock, only when no vzdump is genuinely in-flight (the load-bearing invariant — at startup the agent's own backup loop hasn't run, so the lock is stale; the guard fails SAFE if it can't confirm), itpct unlocks → deletes the danglingvzdumpsnapshot (API + WaitTask) → starts the CT iffonbootand not already running. Scope is strictly the two vzdump locks;migrate/disk/create/… are left untouched. Idempotent.internal/proxmox: new readsGuestConfig.Lock()/OnBoot(),Client.ListSnapshots,Client.ListRunningTasks, and theSnapshottype. Reads + snapshot-delete + start go through the API token; onlypct unlockshells out (no API equivalent).- sudoers + capability manifest: new narrow grant
FELHOM_STALELOCK = /usr/sbin/pct unlock [0-9]*and Critical capabilitystalelock-unlock(a stuck-locked guest = customer box down).visudo -cfclean; covered by the manifest↔sudoers build gate. - Tests: recovery sequence + companions — no-lock touches nothing; non-backup lock left alone; onboot=0 unlocked-but-not-started; delsnapshot only when a snapshot exists; invariant guard (live backup → not cleared; unconfirmable → fail-safe); already-running → not restarted.
-
F2-a — shared-parent boot script never redeployed (
internal/localapi/intermediary.go). The host's/mnt/felhom-driveswas still in root'sshared:1peer group (so every drive bind DOUBLED) because the live boot script predated the v0.36.6make-privatefix. Root cause:EnsureSharedParentgated the (re)install on the unit file only, so a script-only change never deployed. Fixed: the newsharedParentInstallStalehelper compares both the script and the unit (missing or differing → reinstall). Boot-time-only — it rewrites the on-disk script; it does NOT churn the live mount (the live bind/make-private/make-shared stays guarded on!isHostMountpoint). Verified empirically on the host: the correctbind → make-private → make-sharedsequence gives the parent its own group + no doubling.- Tests: a stale-script/current-unit case triggers reinstall (the F2-a regression); both-current is a
no-op; missing files are stale; a content guard asserts the shipped script keeps
make-private.
- Tests: a stale-script/current-unit case triggers reinstall (the F2-a regression); both-current is a
no-op; missing files are stale; a content guard asserts the shipped script keeps
-
Live-caught fixes (same version, found during felhom-pve validation): PVE 9.x rejects
GET /nodes/{node}/tasks?running=1(HTTP 400 "property not defined in schema") — the invariant guard now uses?source=active. And the unprivileged-LXC start emits a benignWARNINGS: 1(systemd-nesting) advisory that false-failed the recovery's start —Startnow usesAllowWarnings(matching the restore-test's start step). -
§D supervised reboot — both findings live-validated. F2-a: after reboot
/mnt/felhom-drivescame up as its OWN peer group (shared:94, notshared:1) with no doubling; guest sees both drives, apps healthy. F2-b: a reboot with the exact stale state (inducedsnapshot-deletelock + a real danglingvzdumpsnapshot) reproduced the stuck symptom (pve-guests "CT is locked (snapshot-delete)" → start failed) and the agent auto-recovered — unlock → removed the real dangling snapshot → started the CT. Zero spurious operator pages on the reboots. Seefelhom.eu/documentation/audits/TESTRUN-fullstack-2026-06-29.md. -
Version
0.48.0 → 0.49.0.
v0.48.0 — report the served local-API leaf fingerprint (hub-side re-key detection, Part A) (2026-06-29)
The agent now rides its served leaf fingerprint on every host report so the hub can detect an
agent re-key fleet-wide (the last self-health leg — host_leaf_changed, hub v0.22.0).
internal/hub/report.go: newHostReport.LeafFingerprint string(leaf_fingerprint) — the SHA-256 of the leaf the agent currently serves. Empty when the local API is disabled (no leaf) → the hub treats "" as unknown, never an alert. Not a secret.internal/hub/collect.go+cmd/felhom-agent/main.go:Collector.SetLeafFingerprint(fp)threads thefpfromEnsureLeaf(the SAME value the loud LOADED/REGENERATED log reports) into every report, next toCapabilities.- Tests: the report includes the fp when set,
""when unset (local API disabled); golden + contract + field-names tests updated (cross-repo golden mirrorsleaf_fingerprint). Version0.47.0 → 0.48.0.
v0.47.0 — controller-swap verify hardening: reject a crash-looping no-healthcheck image (F1) (2026-06-29)
Closes F1 from the no-mercy testrun: a controller image with no HEALTHCHECK that crash-loops could
land a single "Running" inspect poll → the swap marked it healthy → no rollback (alpine tagged as
the controller passed in ~4 s, then Restarting (0)). The real controller image has a healthcheck so
the live severity is low, but the rollback safety net had a hole.
internal/localapi/controllerswap.go:controllerHealthynow also reads{{.RestartCount}}(a 4thdocker inspect -ffield) —running && RestartCount>0→ not-ok (a process that has already crash-restarted isn't stably up, regardless of healthcheck). It also signalsneedsDwellfor the no-healthcheck (none) case.verifyadds a stability dwell: a no-healthcheck image must report ok onverifyDwell(=3) consecutive polls before it's accepted; a realhealthyresult is trusted immediately (Docker already gated it). Any not-ok resets the dwell. Timeout → existing rollback path runs. No change to writeImage, the sudoers grants (the*indocker inspect -f *spans the extended template — confirmed live), or the state-file/rollback orchestration.- Tests: F1 red-proof (
RestartCount>0→ verify false; companion: rc=0+dwell=1 verifies → the rc check is what blocks it); the dwell (single ok then crash → verify false; companion dwell=1 accepts it); a realhealthyimage verifies promptly (no false rollback). ExistingRollbackOnUnhealthy/HealthyWithNoHealthcheckstay green. Version0.46.0 → 0.47.0.
v0.46.0 — leaf lifecycle: signal + loud-log a regenerated leaf (prevention, Part B.1) (2026-06-29)
Makes an accidental local-API leaf regeneration (the 2026-06-28 root→non-root migration class —
moving /var/lib/felhom-agent aside silently minted a new leaf → every controller's pin invalidated
for days) visible immediately instead of silent.
EnsureLeafnow returnsgenerated bool(internal/localapi/cert.go): false = an existing pair was LOADED (stable fingerprint), true = a fresh leaf was GENERATED.- Loud call-site (
cmd/felhom-agent/main.go): a load logsINFO local-api leaf LOADED; a regeneration logsWARN local-api leaf REGENERATED — any previously issued bootstrap pins are now INVALID; controllers will fail the pin check until re-bootstrapped(with the new fingerprint). - No new sudo/capability surface — pure return + log change. The companion install-script preservation
(
--preserve-state-from+ the populated-host guard) lives infelhom.eu/scripts/felhom-host-install.sh. - Tests:
EnsureLeaffirst callgenerated==true, secondgenerated==falseAND same fingerprint (persistence keeps the pin stable). Version0.45.0 → 0.46.0.
Changelog
All notable changes to felhom-agent are recorded here. Update on every code change that gets pushed.
v0.45.0 — controller-swap under non-root: stdin tee write + narrow sudoers grants (Option A) (2026-06-29)
Restores fleet controller-swap / managed auto-update under the non-root agent — the one capability
the 2026-06-29 sudoers audit deliberately left broken because the old write vector needed arbitrary
in-guest execution. Mechanics spike-proven
(felhom.eu/documentation/audits/SPIKE-controllerswap-narrow-grants-2026-06-29.md, GO). No controller
change — the swap endpoint contract is unchanged; only the agent's internal write mechanism + the
allowlist.
writeImageno longer shells out. WasGuestExec("bash","-c","printf '%s\n' '<img>' > <file>")(the swap's only interpolated/shell vector). NowGuestExecStdin(strings.NewReader(img+"\n"), "tee", "/etc/felhom-controller-image")— the image ref is piped on stdin into an in-guesttee; no shell, no interpolation. The trailing\nkeeps the on-disk bytes byte-identical to the golden'sprintf '%s\n', and the bootstrap readsIMAGE=$(cat …)(newline-stripping), so the write is consumed identically.ValidControllerImagestill gates upstream.- New stdin seam (no fenced-runner bypass):
proxmox.Runner.RunStdin/ExecRunner.RunStdin(Run withcmd.Stdin),GuestBinder.GuestExecStdin, andGuestExecutor.GuestExecStdin— the swap routes stdin through the SAMEsudo -nfenced runner as every other privileged op. FELHOM_CONTROLLERSWAPsudoers alias (5 narrow, auditable grants):cat <fixed file>,docker image inspect *,docker inspect -f *,systemctl restart <fixed unit>,tee <FIXED image file>. No generalpct exec, nobash -c— the spike's negative controls (arbitrary exec,teeto any other path,docker rm,rm -rf) stay denied. The 5 are added to the v0.44.0 capability manifest (Critical — a silently-broken fleet auto-update is operator-alert-worthy), so the self-probe watches them and the build-test asserts grant↔code coverage (companion red-proof: dropping theteegrant fails the gate — demonstrated red→green on the real file).- Existing swap tests (happy / rollback-on-unhealthy / image-absent / no-healthcheck / bad-image /
single-flight) pass over the new write path; a new test asserts the write is stdin-
teewith exactimage\nand no shell vector. Version0.44.0 → 0.45.0.
v0.44.0 — privileged-capability self-probe (build-time manifest test + runtime probe + hub snapshot) (2026-06-29)
The agent now self-checks the sudo -n grants it depends on, so a missing allowlist entry (the
2026-06-28 cutover class: lxc-info/make-private/…) is caught LOUD — in CI at build time and on the
host at runtime — instead of surfacing days later as user-visible breakage. First slice of agent
self-health; the controller↔agent channel check is a separate later task.
internal/capability(NEW): aManifest()of the required(binary, representative-arg)vectors (seeded from the 2026-06-29 audit — the OK + CLOSED rows; the SURFACED/DEFERRED rowspct exec */pct create/mount UUID/sensorsare deliberately excluded).Prober.Probelists each against the live policy withsudo -n -l -- <binary> <args>(a policy LIST — never executes, safe for mkfs/pct entries) via a DIRECT runner, plus anos.Statexistence check, mapping took/degraded("sudo policy denied" | "binary not found"). A total sudo failure (drop-in missing) collapses to ONE aggregate signal. Serve-degraded: the probe never blocks startup, panics, or errors.- Build-time gate (
manifest_test.go): parsesconfigs/felhom-agent.sudoers, translates each glob to a regex, and asserts every manifest vector is covered by a grant — exactly what would have caught the droppedlxc-info/make-privatelines in CI. Includes a red-proof: with thelxc-infoline removed from an in-memory copy, the check FAILS forguest-init-pid(and passes on the real file) — proving the gate is not hollow. - Runtime wiring:
Proberuns once at startup (INFOcapabilities self-check N/N ok, plus an ERROR per degraded capability naming the gated feature) and on every hub-report cycle; the snapshot rides the report as the new non-nilHostReport.Capabilities []capability.Status(golden + contract test updated; cross-repo hub copy mirrors it). - No allowlist change; the live host is post-audit complete, so the probe reports N/N ok — itself
a live proof the probe agrees with the fixed sudoers. Version
0.43.0 → 0.44.0.
(sudoers completeness audit, folded into v0.44.0) — close non-root allowlist gaps (2026-06-29)
A full audit of every privileged command the agent shells via sudo -n against
configs/felhom-agent.sudoers, closing the read-only/fixed-vector gaps left by the 2026-06-28
root→non-root cutover. Sudoers-only change — no Go change, no version bump (the file is fetched
canonically by the host-install script). Root cause of the multi-drive "attach one, the other drops"
symptom (audit felhom.eu/documentation/audits/SPIKE-multidrive-mutual-exclusion-2026-06-29.md): the
allowlist was incomplete, so several sudo -n calls were denied under the non-root user.
lxc-info -n [0-9]* -p -H→ FELHOM_INTERMEDIARY (THE root-cause fix).guestInitPID(intermediary.go:256) shells this to resolve the guest init PID forGuestSeesMount→bound_under_parent. It was absent from the allowlist →sudo -ndenied → empty PID → every external drive reported absent → the controller drive-gate stopped each drive's apps (flapping). With the grant,bound_under_parentreports truthfully and the gate quiesces.mount --make-private /mnt/felhom-drives→ FELHOM_INTERMEDIARY.EnsureSharedParent(intermediary.go:110) calls it to isolate the shared parent's peer group on first setup; the allowlist had only--make-shared, so the parent stayed in root's peer group and host submounts "doubled". Guarded by a mountpoint check (never re-churns a live parent).systemctl restart dnsmasq→ FELHOM_DNSMASQ. The v0.29.x LAN-DNS fix switchedreload→restart(lanresolver.go restartDnsmasq) but the allowlist still only permittedreload→ split-horizon DNS self-heal was silently denied under non-root. Added alongside the retainedreload.pct set [0-9]* -onboot 1→ FELHOM_PROVISION. The provision back-half (backhalf.go, F3 auto-start) sets onboot; only-mp[0-9]*was allowed → denied under non-root.pct reboot [0-9]*→ FELHOM_GUESTHOOK.RebootGuest(disks.go:448, the enroll "activate pending binds" fallback) was unmatched.
Surfaced for operator decision (NOT added — would require arbitrary root-in-guest): GuestExec's
general pct exec [0-9]* -- <…> (controller-swap self-update / Phase-2 managed updates) runs variable
vectors incl. bash -c "<interpolated>" — granting it = arbitrary execution. Controller-swap is
currently broken under the non-root agent until narrow per-vector grants are decided. Deferred:
sensors -j (defined-but-unwired AND lm-sensors not installed on the host — no live caller, path
unverifiable). Not added (no daemon caller): pct create … (CreateGoldenLXC, maintenance/broad),
mount UUID=… … (MountUSBByUUID, legacy/unreferenced). Full audit table in REPORT.md.
v0.43.0 — canonical systemd unit + binary published to Gitea (BUNDLE slice) (2026-06-28)
Day-0 no longer needs a hand-installed agent. The agent binary is now PUBLISHED to Gitea as a generic package and the host-bootstrap script fetches → verifies (sha256 vs the hub-vouched manifest) → installs it. This commit adds the canonical systemd unit (was hand-made per host) and the publish tooling; the binary itself is a version-only rebuild (no behavioural change).
configs/felhom-agent.service(NEW, canonical):User=felhom-agent/Group=felhom-agent(the documented non-root production model — README "Process model";privileged.mode: "sudo"+ the narrow sudoers allowlist),ExecStart=/usr/local/bin/felhom-agent --config /etc/felhom-agent/agent.json,After=network-online.target pve-cluster.service pveproxy.service,Restart=on-failure,StateDirectory=felhom-agent. Deliberately NO sandboxing, with the reasons documented inline:NoNewPrivilegesis NOT set — it would block the setuidsudothe agent needs for every host-root op (mount/format/pct/dnsmasq), silently killing all privileged capability.- NO mount-namespacing hardening (
ProtectHome/ProtectSystem/PrivateTmp/…) — any of those give the unit a PRIVATE mount namespace, and the intermediary-mount drive model relies onmount --make-shared/--bindpropagating into the running guest; in a private namespace every drive enrollment would silently break. The agent shares the host mount namespace; the sudoers allowlist is the security boundary.
scripts/publish-agent.sh(NEW): builds (optional) + PUTs the binary to/api/packages/admin/generic/felhom-agent/<ver>/felhom-agent(Gitea generic), printsAGENT_VERSION+AGENT_SHA256, and does a GET round-trip (re-fetch + sha256 re-check) to prove the artifact is fetchable + intact. Pinned to a version (never:latest); idempotent (delete-then-PUT); asserts the binary's--versionmatches the publish version. Creds viaGITEA_USER/GITEA_TOKEN(falls back toREGISTRY_USER/REGISTRY_TOKEN).configs/build-golden.sh: after the vzdump archive is produced, computes its sha256 and PUTs it to/api/packages/admin/generic/felhom-golden/<golden-version>/golden.tar.zst(<golden-version>= the baked controller version), printingGOLDEN_VERSION+GOLDEN_SHA256. Opt-in (only when the Gitea creds are set); the local-golden auto-discovery stays as a fallback.configs/felhom-agent.sudoers(latent bug fix): escaped the commas in thelvs -o lv_name\,data_percent\,metadata_percentandlsblk -o NAME\,FSTYPE\,PTTYPE\,MOUNTPOINTargument lists. Sudoers treats a bare comma as a command separator, sovisudo -cfREJECTED the file — it had never been visudo-validated live because the demo host ran the agent as root+direct(sudoers unused). The escaped commas still match the agent's real comma-bearing args. Surfaced by the BUNDLE live install (the host-install scriptvisudo -cf-validates before installing).cmd/felhom-agent/main.go:version0.42.0 → 0.43.0.- The operator records the printed agent + golden version+sha256 in the hub (Configs → "Day-0 artifacts"); the host-bootstrap script verifies fetched artifacts against those before installing.
go build/vet/test ./...green.
build-golden.sh — default controller image bumped to current; golden rebuilt at 0.85.1 (2026-06-27)
Operational + a default fix (no agent binary change — version stays v0.42.0).
configs/build-golden.sh: theCONTROLLER_IMAGEdefault (positional arg 6) was a stale…/felhom-controller:0.43.0— an argument-less golden build baked a wildly old controller, so fresh Day-0 boxes started old (the demo started at 0.77). Bumped the default to the current…/felhom-controller:0.85.1so the worst case (no explicit arg) is merely "current", not ancient.- Always pass the controller version explicitly at each rebuild — this default only bounds the
worst case. A future
make goldenthat resolves the latest pullable tag would remove the need for a hand-bumped default (Observation, not this task).
- Always pass the controller version explicitly at each rebuild — this default only bounds the
worst case. A future
- Golden rebuilt at 0.85.1 on
felhom-pvewith the image passed explicitly (build-golden.sh 9100 … gitea.dooplex.hu/admin/felhom-controller:0.85.1). New archive volid:local:backup/vzdump-lxc-9100-2026_06_27-11_42_51.tar.zst(rootfs 32G + Docker-data 16G + user-data 8G, all in the archive; mp0+mp1 inclusion confirmed in the vzdump log).- Baked-image verify (cheap, mandatory): in the build guest
/etc/felhom-controller-image=…:0.85.1anddocker imagesshowed it baked (379 MB). New Day-0 provisions now ship current. - The host-bootstrap script auto-discovers the newest golden, so it picks up this rebuild automatically. The real demo 9201 was not re-provisioned (it is the Phase-2 floor test box).
- Baked-image verify (cheap, mandatory): in the build guest
v0.42.0 — agentic controller update: in-guest image swap + rollback (Phase 1) (2026-06-26)
The host agent now owns the in-guest controller image swap — the new-architecture replacement for
the controller's dead in-container docker compose self-update. The controller pre-pulls the target
image (shared docker socket, its own registry token) then asks the agent to swap; the agent — external
to the controller container, so it survives the controller being killed mid-swap — does the rest and
rolls back if the new controller doesn't come up healthy.
- New local-API routes (
internal/localapi/controllerswap.go, token-scoped viawithGuest):POST /controller/swap {image}→ 202{status:"swapping", previous_image, target_image}, then async: record previous (crash-safety state file/var/lib/felhom-agent/controller-swap-<vmid>.json) → confirm the target image is present in the guest (else abort, no swap) → write/etc/felhom-controller-image→systemctl restart felhom-controller-bootstrap.service→ poll the new controller to healthy (docker inspect, ≤90s) → roll back to the previous image + restart if it doesn't (the guest is never left without a controller). Single-flight per guest (409 if busy). Image ref is strict-validated (gitea.dooplex.hu/admin/felhom-controller:<semver>) before any action.GET /controller/swap/status→{state: swapping|done|failed, current, previous, target, error}.
GuestBinder.GuestExec(internal/localapi/guestbind.go): the onepct execseam the swap composes over (cat/inspect/write/restart), reusing the fenced root runner.--selftest=controller-swap -vmid -image <ref>: exercise the primitive directly (the target image must already be pulled in the guest).- Wired
ControllerSwap: guestBinderinto the local-API server (cmd/felhom-agent/main.go). - Tests (
controllerswap_test.go): happy swap, rollback-on-unhealthy (+ companion red-proof: dropping the rollback leaves the guest on the bad image and fails the test), image-absent no-swap, no-healthcheck-running, bad-image 400, single-flight 409.
v0.41.0 — provisioned customer guests auto-start after a host reboot (onboot:1) (2026-06-24)
F3 fix. The provision back-half now sets onboot:1 on the customer guest, so after a host
reboot/power-cut the customer's whole home-server (controller + apps) comes back on its own —
previously every provisioned guest inherited the golden's --onboot 0 and stayed stopped until a
manual pct start (confirmed live in the stable-path/sys-drive restart campaign, Phase 4.1). The new
step is a fatal pct set <vmid> -onboot 1 placed right after the config-mount attach (backhalf.go),
mirroring the config-mount/parent-bind pct set ops. No startup/boot-order/delay — the v0.75
mountpoint-gate already covers the drive-bind race at boot (Phase 4.4), so the controller won't write
app data onto the rootfs while the agent re-binds drives.
The golden stays onboot:0 (build-golden.sh unchanged): a template must not auto-start, and
onboot is a per-guest property the back-half is the right place to set. Unit-tested
(TestProvision_SetsOnbootOne asserts the exact pct set … -onboot 1 invocation, with a red-proof
against removing the call). The pre-existing demo guest 9201 (provisioned pre-fix) was remediated
non-destructively with pct set 9201 -onboot 1. Capstone live-validated (2026-06-24): destroyed +
re-provisioned 9201 through the real provision chain with v0.41.0 → fresh pct config showed onboot: 1
with no manual set; a subsequent felhom-pve host reboot brought 9201 back running with no manual
pct start (the onboot:0 scratch guests correctly stayed stopped), controller + base infra healthy,
drives re-bound at stable, sys_drive separate — the exact Phase-4.1 failure now passes.
v0.40.0 — third CT volume: SSD user-data (/mnt/sys_drive, mp1) baked + -sysdata-grow (2026-06-23)
The third golden volume. Extends the OS/Docker-data split (v0.29.x) to a three-volume layout:
rootfs + Docker-data (mp0) + SSD user-data (mp1 @ /mnt/sys_drive, backup=1) — the
controller's system_data_path. Until now /mnt/sys_drive was a plain directory on the 32 GB OS
rootfs, so the controller correctly warned that SSD app data (<sys_drive>/felhom-data) lands on the
OS drive. Baking it as its own thin volume clears that warning with zero controller change (the
controller already auto-discovers <sys_drive>/felhom-data and warns via system.IsMountPoint); the
mp under the guest's /mnt reaches the controller container through the existing
-v /mnt:/mnt:rslave bind.
configs/build-golden.sh—pct creategains--mp1 ${ROOTFS_STORAGE}:${GOLDEN_SYSDATA_GB},mp=/mnt/sys_drive,backup=1(new envGOLDEN_SYSDATA_GB=8, near-empty; provision grows it). The resilience guards are mirrored formp1: afindmnt /mnt/sys_driveseparate-mount assertion, and the vzdump-inclusion guard now aborts if eithermp0ormp1is EXCLUDED (the B3 trap — extra mountpoints defaultbackup=0). The golden does NOT pre-createfelhom-data; the controller does once it's a real mountpoint.internal/reconcile/bringup.go—const DefaultSysDataMount = "mp1";BringUpSpecgainsSysDataGrowGB int+SysDataMount string; a new "4c" grow block (online, grow-onlyResizeLXC, its own task) mirrors the "4b" Docker-data grow.0 = skip(separateness comes from the golden, not the grow — the warning clears regardless of size).cmd/felhom-agent/main.go—-sysdata-grow/-sysdata-mountflags (mirror-datavol-grow/-datavol-mount);bringUpSizingcarries them into all three bring-up/provision call sites;--selftest=provisionhelp text updated.- Static volume, NOT an enrolled drive.
/mnt/sys_driveis part of the baked golden layout; it never enrolls/ejects/decommissions and is deliberately kept off the drive-intent machinery.freeMountSlotauto-skips the bakedmp0/mp1so enrolled drives never collide. - Tests:
TestRunBringUp_StorageSplit_SysDataGrow(assertsResizeLXC(vmid,"mp1","+42G")) +…_SysDataGrowZeroNoResize(0 → no mp1 resize). RUNBOOK-provisioning-storage.md extended to the three-volume layout (default ~512 GB SSD: 32 rootfs + 200 docker-data + 50 user-data).
v0.39.0 — DR recipe completion: live PBS coord + drop the two unfillable drive fields (2026-06-16)
DR-recipe agent-half completion. A live eyeball of the demo recipe (v0.38.0) found three host-half problems; all three are resolved here. No behavior change outside the recipe path.
- PBS coord now resolved LIVE each collect. New
internal/pbs/live_reporter.go—LiveSnapshotReporterimplementshub.PBSReporterby doing the cheapClient.Snapshots()list itself, with last-known-good fallback, instead of reading only the verify-loop'sSnapshotStore. Previously the recipe'spbsblock was omitted whenever the store was empty — which a one-shot collect (--selftest=hub) and the first ~6 h window of every daemon after a restart always saw (the verify loop populates the store on its own 6 h cadence). The restore SOURCE must not depend on a maintenance cadence. Per-datastore: a live error/timeout → that datastore's last-known-good; a successful (even empty) response is authoritative and updates the shared store. Targets-resolution failure → the full LKG aggregate. Bounded byDefaultLiveSnapshotTimeout(8 s) so a hung PBS never stalls the heartbeat. List only — it never triggers aVerify. The verify loop keeps Recording into the SAME store (shared last-known-good); both use one hoistedpbsTargetsclosure.SnapshotStore.Get(datastore)added (per-datastore LKG copy) — the onlySnapshotStorechange.- Wired into the collector in BOTH
runDaemonandrunSelftestHub(the selftest built its own collector with anilreporter — that is why the live--selftest=hubshowedpbs_snapshots:[]). - Intended side effect:
report.pbs_snapshotsis now live too (fresher hub PBS view).
drives[].roleDROPPED from the v1 host-half shape. A drive's purpose is a hub/operator-owned manifest concept, not cleanly derivable host-side (both demo externals arecontent=backup, yet one is the primary data drive and the other holds no apps). Deferred until the hub/operator stamps it.drives[].restic_repo_coordDROPPED from the v1 host-half shape. It named a backup tier that does not exist — cross-drive backup is rsync to the SAME internal SSD; there is no offsite/second-failure- domain bulk copy. RESERVED for a future tier (see the BACKLOG note in REPORT). v1 drive shape is now{durable_id, mount_path, intent, fs_type?, total_bytes}— identifiers/intent/size only.- The hub reads drives as
json.RawMessage, so dropping fields needs NO hub struct change — only golden + test sync. Cross-repo golden (host-report.golden.jsonhere + the hub's copy) re-pinned and verified byte-identical (sha25657f2a5e7…18b2f2b5— manual checksum-diff discipline): the hub copy previously lacked thedr_recipesection entirely; it is now a verbatim copy of the agent golden. - Tests: new
internal/pbs/live_reporter_test.go(T1 coord-present-without-prior-verify [load-bearing] + inline bare-store companion, T2 error→LKG fallback, T3 success-warms-store, T4 targets-error→aggregate, T5 bounded-by-timeout, T6 empty-success-authoritative);TestDRRecipeHostHalf_V1DriveShape(drive object carries neitherrolenorrestic_repo_coord);TestBuildDRRecipeHostHalf/TestHostReport_ContractMatchesGoldenupdated to the v1 drive shape. Each companion was demonstrated to FAIL on the pre-fix/mutated code, then reverted (see REPORT).
v0.38.0 — DR recipe: emit the secret-free storage/guest/PBS half in the host-report (2026-06-16)
DR recipe slice (agent half). Additive dr_recipe section on the host-report — the agent half of the
secret-free reconstruction recipe (SPIKE-dr-recipe-2026-06-16.md) that complements escrow (keys) +
PBS/restic (bytes): the non-secret SCAFFOLDING an operator must rebuild before the PBS bytes can land.
The hub assembles it with the controller's app half into one customer recipe.
internal/hub/dr_recipe.go—DRRecipeHostHalf{recipe_version, guests[], pbs, drives[], pve_storage[]}built by the pureBuildDRRecipeHostHalf(guests, targets, pbs)from facts the report ALREADY collects (no new privileged reads):guests[]= each guest's sizing (GuestSpec, skip status-unknown);drives[]= the user-data external drives (usb/local-dir with auuid:durable-id + mount path) with{durable_id, role, mount_path, intent, total_bytes};pve_storage[]= every storage target{name, type, content}(thestorage.cfgscaffolding);pbs= the latest snapshot's coordinates{repo_id (the pbs storage id), namespace, latest_snapshot_id}. Wired intoCollect()after the facts are gathered;HostReport.DRRecipe(always set, never null).- BOUNDARY (the Phase-1 lesson): every field is an identifier / intent / size / coordinate — NEVER a
key, password, token, hash, or
ENC:value. The PBS encryption key stays in escrow; the access token in identity-escrow; the restic password in escrow — the recipe names only therepo_id/namespace/durable_id/restic_repo_coordthe restore TARGETS.recipe_version=1; read is ignore-unknown (forward-compat). The wire shape is pinned in the cross-repo golden (host-report.golden.jsonhere + the hub's copy — keep them byte-identical; manual checksum-diff on any change). - Tests:
TestBuildDRRecipeHostHalf(drives = only user-data; pve_storage = all; pbs = latest; guests skip nil-spec),..._NoPBS(omitted, non-nil slices),TestDRRecipeHostHalf_NoSecrets(the lighter boundary mirror — serialized half carries NO credential-shaped key; the load-bearing version is on the controller emitter), and thedr_recipekey-set added toTestHostReport_ContractMatchesGolden.
v0.37.0 — host-reboot remount re-resolves enrolled drives by filesystem UUID (2026-06-16)
TASK A — close out the reboot story (agent half). On a host reboot the kernel can re-enumerate block
devices and move a drive's node (felhom-usb /dev/sdb→/dev/sdc), and a .mount unit left disabled
by a prior detach never auto-mounts at boot — so an enrolled drive could stay unmounted (or, with any
node-trusting remount, mount the WRONG device). Root cause pinned LIVE: felhom-usb's systemd mount unit
was disabled (no multi-user.target.wants symlink) while felhom-flash's was enabled; What= was
already correct (by-UUID), but nothing re-asserted the unit at startup.
storage.ResolveStorageDevice(durableID)— resolves the enrolleduuid:<fs-uuid>storage scheme to its CURRENT backing/devnode by re-scanning/dev/disk/by-uuid(never a cached node); errors if the UUID is genuinely absent so a caller skips a gone drive instead of fail-mounting a stale node.storage.parseFelhomMountUnit— pure inverse ofrenderMountUnit(Name/UUID/Where/Type/Options) keyed on aManaged by felhom-agentmarker; ignores any foreign.mountunit.(*SudoHostOps).ReassertEnrolledMounts(ctx)— at startup (BEFORE binding into the guest) and on the periodic 20s tick: for each enrolled.mountunit, re-resolve by UUID and re-runEnsureMount(idempotentsystemctl enable --now) — re-enables a disabled unit AND mounts the CURRENT device by UUID, so a/dev/sdXreshuffle is a no-op. Skips ONLY the durable steady state (mounted AND enabled), via the pureshouldReassertMount; a mounted-but-DISABLED unit (the exact live felhom-usb bug — it serves now but a reboot would not auto-mount it) is still re-asserted to re-create the wants-symlink. Enabled-state is read with a privilege-freeos.Lstatof themulti-user.target.wantssymlink (unitEnabled) — nosystemctl is-enabledsubprocess, no new sudoers entry. An absent UUID is skipped (re-asserts on a later tick).- Wired in
main.goahead ofReassertGuestBindsso mounts are live before the guest binds re-assert. - Tests (Linux, seam the device-resolution):
TestResolveStorageDevice_ToleratesDeviceLetterMove(UUID symlink moved sdb→sdc → resolves sdc; companion asserts the cached enroll-time node differs from the freshly-resolved one — a node-based remount would target the wrong device),..._AbsentAndScheme(absent UUID errors; only theuuid:scheme resolvable),TestParseFelhomMountUnit(render→parse round-trip + rejects a foreign unit),TestShouldReassertMount(the four mounted/enabled combos — pins the mounted-but-disabled re-assert),TestUnitEnabled(wants-symlink detection).
TASK A2 — verdict: enrolling a NEW drive does NOT need an LXC restart. The enroll path lands on the
live intermediary-mount AttachDrive (/disks/guest-attach → handleDiskGuestAttach → AttachDrive,
"no pct, no reboot") under the single shared parent — unbounded named live slots — NOT the legacy
RebootGuest branch. The operator's pre-created-slot-pool idea is therefore unnecessary.
v0.36.7 — isolate the shared parent only on CREATE (no peer-group churn) (2026-06-15)
Follow-up to v0.36.6: make-private+make-shared must run ONLY when the self-bind is first created, not on every reconcile — re-doing it churns the peer-group id and ORPHANS the guest`s already-established slave (propagation silently dies, guest sees empty). Guarded on the mountpoint check; on a fresh boot it runs once before pve-guests so the guest slaves the right group.
v0.36.6 — shared parent gets its OWN peer group (make-private first) — ROOT CAUSE of double-bind (2026-06-15)
The shared-parent self-bind INHERITED the root mounts shared peer group (/mnt/felhom-driveswasshared:1same as/), so every drive bind under it propagated back via the root peer and DOUBLED (2 stacked binds per drive — the real cause behind v0.36.3-.5). EnsureSharedParent + the boot script now make-private(detach from the root group) BEFOREmake-shared` (own group whose only slave is the
guest), so a drive bind propagates to the guest exactly once.
v0.36.5 — AttachDrive normalizes to exactly one bind (2026-06-15)
AttachDrive now COUNTS the binds at a stable path (countHostMounts) and normalizes to exactly one: it is a no-op only when there is exactly ONE bind the guest sees; otherwise it strips ALL existing binds (bounded loop) and lays down one fresh bind. This converges a stacked double-bind to one — the old umount-one+mount-one force-rebind never did. Caught when a double-bind survived a guest reboot.
v0.36.4 — serialize AttachDrive/DetachDrive (no double-bind race) (2026-06-15)
A mutex on GuestBinder serializes AttachDrive/DetachDrive so a controller-triggered reconnect and the agent`s periodic reconcile can no longer both pass the isHostMountpoint check and double-bind the same stable path (a TOCTOU race observed live as 2 stacked binds during rapid eject/reconnect).
v0.36.3 — DetachDrive loop-umounts stacked binds (2026-06-15)
DetachDrive now umounts ALL stacked binds at a stable path (bounded loop), not just one layer — so an eject/detach fully detaches even if more than one bind accumulated (operator bind on top, or a rare attach race), keeping the fail-close intact. Caught in the E13 rapid eject/reconnect sweep.
v0.36.2 — eject also keeps the raw mounted (reconnectable) (2026-06-15)
Extends v0.36.1 to EJECT: eject now DetachDrive`s the bind under the parent but LEAVES the raw /mnt/ mounted (consistent with decommission), so the H1 disconnect→reconnect roundtrip re-binds cleanly on a non-removable drive. Physical removal is the separate "remove from system" action. Test: eject calls DetachDrive + does NOT unmount the raw.
v0.36.1 — decommission keeps the raw mounted (re-enrollable) (2026-06-15)
Fix caught in the E10 acceptance test: the self-serve decommission unmounted the RAW /mnt/ host mount, which orphaned a non-removable drive (no re-plug) so a one-click re-enroll bound an empty dir. On the intermediary model decommission is now a LOGICAL retire — it DetachDrive`s the bind under the parent (drive invisible to the guest) but LEAVES the raw mounted, so re-enroll re-binds cleanly. Physical removal stays the separate "remove from system" action. Test updated.
v0.36.0 — guest boot-id on /disks (deterministic guest-reboot recreate) (2026-06-15)
The agent now emits guest_boot_id on GET /disks: <host-btime>-<guest-init-starttime> — changes on
every guest boot (host reboot OR guest reboot) but is STABLE across a controller-only restart. The
controller persists the last-seen value and DETERMINISTICALLY recreates drive-backed apps when it
changes (replacing the fragile timed state-sample that could miss an app stopped at the sample instant).
GuestBootID reads /proc/stat btime + field 22 of /proc/<init-pid>/stat (parsed after the last
) so a comm with spaces/parens does not break it).
v0.35.1 — shared-parent unit: run before pve-guests on host boot (2026-06-15)
Fix for the host-reboot ordering (the shared-parent oneshot never ran before pve-guests on the live
host, so the guest bound a not-yet-shared parent → private bind → propagation broken). The unit now uses
WantedBy=pve-guests.service (pve-guests PULLS IT IN + Before= orders it first) instead of the
unreliable WantedBy=multi-user.target, and drops DefaultDependencies=no. EnsureSharedParent
reinstalls the unit when its content differs (so the fix deploys on the next agent start/reconcile).
v0.35.0 — intermediary mount: guest-reboot re-propagation (load-bearing) (2026-06-15)
Fix for the guest-reboot gap (caught in the live demo migration). A guest's parent bind is NON-RECURSIVE, so on a guest reboot it does NOT carry the pre-existing drive submount, and mount propagation only delivers events created AFTER the bind exists — so an enrolled drive is bound on the HOST but INVISIBLE in the fresh guest namespace until re-bound. Without this, every guest reboot left the apps on empty dirs.
AttachDrivenow takesvmidand checks GUEST visibility (GuestSeesMount, reading/proc/<guest-init-pid>/mountinfo): if the host has the bind but the guest doesn't see it (post-reboot), it FORCE re-binds (umount + mount) to re-fire propagation into the current guest ns.- A periodic reconcile (20s ticker in main) re-runs
ReassertGuestBinds, so a guest reboot self-heals without an agent restart.EnsureSharedParentskips the unit re-install when already present (cheap on repeat). /disksBoundUnderParentnow reflects GUEST visibility (not the host mount) — the accurate signal the controller's drive-absent gate keys on to stop/restart apps across a guest reboot.
v0.34.0 — intermediary mount model: shared-parent + host-side attach/detach + reconcile (2026-06-15)
The drive hot-swap re-architecture (SPIKE-intermediary-mount). Replaces the per-drive pct set -mpN
bind (which needed a guest reboot to activate and bricked the guest when a drive was absent at boot)
with a SINGLE permanent parent bind /mnt/felhom-drives plus host-side swaps underneath it.
internal/localapi/intermediary.go—GuestBinder.EnsureSharedParent(mkdir + self-bind +--make-shared+ installs/enables afelhom-shared-parent.serviceordered Before=pve-guests so the guest's parent bind inherits the shared peer group asslave);AttachDrive(mount --bind /mnt/<name>/felhom-data /mnt/felhom-drives/<name>— propagates into the RUNNING guest live, no pct, no reboot; confined to felhom-data; the stable dir stays host-root-owned = fail-closed);DetachDrive(umount, leaving the bare fail-closed dir);StablePathForRaw/DriveNameFromRaw;isHostMountpoint.ReassertGuestBindsis now a pure HOST-SIDE reconcile: for each enrolled+present drive ensure its felhom-data is bound under the parent (no guest-config read, no slot, no reboot) — fixes F9 and drive-reconnect for free. Runs at startup (ensures the shared parent first).handleDiskGuestAttachusesAttachDrive(returns the stableguest_path); eject + decommission callDetachDrive. LegacyAttachBind/DetachBindretained for the transition (decommission still--deletes any lingering legacy mp)./disksreporting addsGuestPath(the stable/mnt/felhom-drives/<name>the controller repoints HDD_PATH to) andBoundUnderParent(live-in-guest signal for the controller's drive-absent gate).- Provision adds the one permanent parent bind (
-mp8 /mnt/felhom-drives,mp=/mnt/felhom-drives).
Tests (non-hollow + companions): TestGuestAttach_BindsUnderParent (uses AttachDrive not legacy pct),
TestReassertGuestBinds_RestoresMissingBind (host-side reconcile, legacy AttachBind never called),
TestStablePathForRaw_DriveName, TestDisks_GuestPathAndBoundUnderParent. Sudoers: new
FELHOM_INTERMEDIARY alias (mount/umount under /mnt/felhom-drives, the unit install, the parent bind).
v0.33.0 — C1 net: pre-start self-heal hook + decommission mp-delete (2026-06-15)
The transitional defense for the C1 brick (B3 critical bug) ahead of the intermediary-mount re-architecture (which makes C1 structural). Two independent nets:
- Pre-start self-heal hook (
internal/guesthook): a PVEpre-starthookscript runsfelhom-agent guest-hook <vmid> <phase>which, for every BIND mountpoint whose source path is missing, creates an empty host-root-owned placeholder dir so the bind succeeds and the guest always boots — fail-closed (host uid 0 is unmapped in the unprivileged-LXC userns, so the guest can't write to the placeholder; a returning drive shadows it). It CREATES rather than DELETEs becausepct set --deletein pre-start would take the config lock the start task already holds (dead-times-out → still bricks); the heal logic is in unit-tested Go, the wrapper just delegates. Installed + registered per-guest by the provision back-half (InstallSnippet/Register). - Decommission mp-delete (
GuestBinder.DetachBind+handleDiskDecommission): decommission now runspct set <vmid> --delete mpNon the slot binding the drive (lock-safe on the running guest), so its now-missing source can't brick the next reboot. The old handler unmounted but left the deadmpNin config — the exact B3 C1 bug. Eject keeps its mp (temporary; the hook covers a reboot-while-ejected).
Tests (non-hollow, each with a companion that fails the pre-fix/trivial impl):
internal/guesthook/heal_test.go (selector ignores storage volumes + present binds, heals only the
absent one; "return nothing"/"return all" both fail) and TestDecommission_DeletesGuestMount
(asserts the correct slot is --deleted; pre-fix never calls DetachBind → fails).
Sudoers: new FELHOM_GUESTHOOK alias (snippet install, pct set --hookscript, pct set --delete mpN).
v0.32.0 — self-serve decommission + intent-aware re-assert (B2a) (2026-06-14)
Customer-self-serve storage decommission (no operator signature; non-destructive — never formats), plus the load-bearing fix that keeps a decommissioned drive from auto-rebinding into the guest.
POST /disks/decommission(internal/localapi/disks.gohandleDiskDecommission, route inserver.go) — mirrorshandleDiskEjectexactly:withGuestself-scoping,scopedFromBody, and the same user-data role gate (roleForMountPathmust beRoleUserData, else 403; fail-safe-to- protected on ambiguity) so a compromised controller can't decommission system/backup storage. It records a PERMANENTIntentDecommissioned, prunes theGuestBindStoreentry (hygiene), and unmounts (so the drive is physically removable). It NEVER calls any format/mkfs path — the data stays on the drive. The operator-signedDecommissionExecutor+reconcile.Classifyclassification are untouched (the absent-drive/DR route).ReassertGuestBindsis now intent-aware (THE correctness fix): the startup re-assert skips any durable-id whose intent is notenrolled, so a decommissioned- (or ejected-) but-still-present drive is never auto-rebound into the guest on agent restart. A nil intent store falls back to legacy bind-all (matching the watchdog's nil-intent rule). Covers both the self-serve and the operator- signed decommission paths (both land onIntentDecommissioned).GuestBindStore.Remove(vmid, durableID)(internal/localapi/guestbindstore.go) — idempotent (absent = no-op), atomic tmp+rename likeRecord; drops the vmid key when its set empties. Re-enroll re-Records via the existingrecordGuestBindon guest-attach, so Remove doesn't break re-commission.IntentRecorderextended withSetDecommissioned+Get(both already on*storage.IntentStore).- Non-hollow tests (
internal/localapi/decommission_test.go): role-gate refuses system/backup (403, no unmount); decommission sets intent + removes the bind + unmounts + never formats; intent-aware re-assert does NOT rebind a decommissioned-but-present drive (companion: enrolled DOES rebind; the intent-blind pre-fix code fails this); re-commission re-records;Removeidempotency + persistence.
v0.31.0 — live-drive F9 + F20-BUG2 + F20-BUG3 (disk bind/wipe) (2026-06-14)
The last live-drive findings, all disk/localapi-side, implemented + deployed on felhom-pve and
validated live on guest 9201 (approach: attach-to-existing, no re-provision — see the audit fixspec).
- F9 — guest data-drive bind survives a re-provision (
4cd1d02). The in-guest bind (pct set -mpN) is config state a destroy+re-provision drops, and nothing restored it → a re-provisioned guest came up with its enrolled HDD unattached. NewGuestBindStore(durable-id-keyed, per guest, recorded at guest-attach) +ReassertGuestBindson agent startup re-adds any bind a guest is missing — only when the durable-id still resolves to a present drive (a swapped/absent disk is never auto-bound), idempotent. PlusDiskInfo.GuestAttached— the missing "bound into THIS guest" signal (vs mere host presence; resolves the F2hdd_configureddisagreement). Live-proven: dropped the bind, restarted the agent (real trigger) → re-attached with no manual call; reboot activated it; an HDD app then deployed onto the drive with data on/dev/sdb1. - F20-BUG2 — one wipe durable-id scheme (
a2a76e7)./disksadvertised onlydurable_id(uuid:, used for assign), but the wipe gate resolvesbyid:/byuuid:→ confirming a wipe with the advertised id was abinding_mismatch. NewDiskInfo.WipeDurableIDvia a shareds.deviceDurableIDseam used by BOTH the list and the gate, so the id the customer copies is the id the gate accepts. Live-proven: a confirmed wipe using/api/disks'swipe_durable_idis accepted (no mismatch). - F20-BUG3 — format runs detached; survives a request deadline AND an agent restart (
4777f8a). mkfs ran under the HTTP request context, so a client deadline SIGKILLed it mid-write → corrupt disk. Now mkfs runs offs.baseCtxvia a persistedformatJobrecord; the handler still returns the synchronous result (backward-compatible) but a dropped request no longer kills it. NewGET /disks/format/status;RecoverFormatJobon startup re-runs an interrupted durable-id-bound format (re-resolved; anti-retarget — a blank/path-bound or unresolvable job is not auto-re-run). Live-proven on the 916 GB felhom-usb: a 2 s client timeout left a ~30 s mkfs running to a clean ext4 (the live-drive corruption is gone); an agent restart mid-format was recovered + completed to a clean fs.
Security fix (from the 2026-06-13 deep-sweep audit). The inline customer-confirmed wipe in
internal/localapi/disks.go handleDiskFormat inspected and gate-bound the device by its durable id
but then ran mkfs on the caller-supplied mutable /dev path (req.Device). A USB re-enumeration
reassigning that /dev node to a different physical disk between inspection and mkfs (a
classify→mkfs TOCTOU) could wipe the wrong drive.
- New
internal/localapi/wipe_reresolve.go:antiRetargetResolve(injected-deps, unit-tested) mirrorssignedjobs.WipeExecutor.Execute— resolve the confirmed durable id → current device, re-derive the device's durable id and require an exact match, re-inspect (still data-bearing), and return the re-resolved device.(*Server).reresolveDurableForWipewires the real storage funcs. handleDiskFormatnow formats the re-resolved device, neverreq.Device; any refusal →409 Conflict, nomkfs. InjectablereresolveWipeseam onServer(defaults to the real path).- Tests:
wipe_reresolve_test.gocovers happy-path, empty/gone/blank, re-inspect-error, and the coreretarget-mismatch-refusedcase. Round-trip safe for legitimate wipes (DeviceDurableID↔ResolveDurableDeviceschemes match). Agent-only deploy; no golden rebake. SeeAGENT-001-FIX-NOTES.md.
v0.29.1 — lanresolver: RESTART dnsmasq on change (not reload) — fixes stale split-horizon IP (2026-06-13)
Bug: after a guest's DHCP IP moved (e.g. the v0.29.0 9201 re-provision: .151 → .141), the LAN
split-horizon resolver kept answering the OLD IP, so LAN clients (via Pi-hole's conditional forward to
the host dnsmasq) resolved *.demo-felhom.eu to the dead IP. Root cause: lanresolver.Manager updated
the per-customer drop-in (address=/<domain>/<ip>) correctly but then ran systemctl reload dnsmasq
(SIGHUP) — and dnsmasq's SIGHUP does NOT re-read its config files (/etc/dnsmasq.d/*.conf); it only
clears the cache + re-reads /etc/hosts/addn-hosts. So the changed address= directive never took
effect until a restart. Fix: reload() → restartDnsmasq() (systemctl restart dnsmasq) for every
config-drop-in change (ReconcileGuest IP change, EnsureDnsmasq base change, Remove/decommission). Restart
is sub-second and the records carry local-ttl 0, so downstream forwarders don't cache a stale answer.
(Live: after the fix + a one-time host dnsmasq restart + a Pi-hole cache flush, *.demo-felhom.eu
resolves to the live guest IP again; future IP moves now self-heal on the loop's next tick.)
v0.29.0 — OS / Docker-data storage split: golden + provision (2026-06-13)
Phase 1 of the storage-split slice (Phase 2 = felhom-controller v0.58.0 prevention layer). The
controller guest's OS rootfs and Docker data are carved onto separate local-lvm volumes for
RESILIENCE — an isolated OS rootfs stays bootable + agent-recoverable if the Docker volume fills.
configs/build-golden.sh— split baked in:--rootfs ${ROOTFS_STORAGE}:${OS_SIZE_GB}(default 32, was hardcoded 8) plus--mp0 ${ROOTFS_STORAGE}:${GOLDEN_DOCKER_GB},mp=/var/lib/docker,backup=1(default 16). The baked controller + infra images land on the data volume and travel inside the golden archive (no empty-volume shadowing, no deploy-time pull).backup=1is MANDATORY — extra LXC mountpoints default tobackup=0= EXCLUDED from vzdump (spike B3), which would drop the images from the archive entirely. The script now also bakes Docker log rotation intodaemon.json(max-size 10m,max-file 3— prevention layer 2D), asserts/var/lib/dockeris a separate mount, and aborts if vzdump excludes mp0.internal/reconcile/bringup.go— sized provision:GuestMountgainsBackup(emits,backup=1— closes the spike-B3/B5 silent-DB-loss trap at the mount builder).BringUpSpecgainsDataVolGrowGB+DataVolMount(defaultmp0): provision GROWS the golden-carried Docker-data volume online to the per-customer target (grow-only, spike B4) rather than attaching a fresh empty volume that would shadow the baked images. PlusRootfsGrowGBfor the OS rootfs.- CLI seam:
--selftest=bring-up|provisiongain-rootfs-grow/-datavol-grow/-datavol-mountflags. Per-customer sizing source = flags now, the slice-10 hub storage manifest later. RUNBOOK-provisioning-storage.md(new): the split provisioning procedure + fresh-PVE-install thin-pool carving knobs (hdsize/maxroot/maxvz, spike B4) + the per-customer sizing seam.- Tests:
buildBringUpConfigbackup=1 emission; bring-up issues rootfs + data-volume resizes.
(no version) — storage OS/data-split spike findings (2026-06-13)
Investigation only — no code changed. Findings report: REPORT-storage-split-spike.md (gates the
provisioning spec for splitting the controller guest's OS rootfs from its Docker/data onto separate
local-lvm volumes). Proven on a throwaway unprivileged LXC (9300, since destroyed): Docker data-root
on a second local-lvm mountpoint works (overlayfs/ext4, no idmap issue, reboot-survives); the
move-then-verify migration is safe (copy-not-move). Key finding: additional LXC mountpoints are
excluded from vzdump by default — they need backup=1 set and a CT restart — so the docker-data
mount must be attached with ,backup=1 or named-volume DBs silently fall out of PBS. The exact seam is
internal/reconcile/bringup.go:313 (buildConfigParams), which today builds mpN without a backup=
flag; GuestMount should carry the flag. Per-customer sizes belong in the slice-10 hub storage manifest
(marked at bringup.go:49-50); the golden rootfs is hardcoded 8 at configs/build-golden.sh:40.
v0.28.0 — backup re-target → felhom-pbs (offsite DR) + operator-signed decommission (2026-06-12)
Whole-guest backup now defaults to the offsite PBS tier (real DR). BackupConfig.BackupTarget()
returns the configured backup.local_backup_target or, when empty, the new default felhom-pbs — a
PBS datastore on SEPARATE HARDWARE (the DooPlex box), so a host disk/hardware failure no longer takes
the backups with it. The target stays fully configurable (set local_backup_target to local/other
to override); no call site hardcodes it. All NewBackupRunner sites (restore-test scheduler, local-API,
--selftest=backup/restore-test) route through BackupTarget().
Proven live on demo-felhom before the re-point (PHASE 0 gate):
- snapshot-mode
vzdump → felhom-pbsstill fires thecreate storage snapshot 'vzdump'marker, so the 8B.2 early-resume/quiesce signal survives a PBS target (the marker is mode-driven, not target-driven); - the restore-test enumerates PBS backups through the SAME generic
StorageContent(/nodes/<node>/storage/felhom-pbs/contentreturnscontent:"backup"+ ctime/vmid/volid), soPickRestoreCandidate/latestArchiveneed NO PBS-client change; pct restorefrom a PBS volid round-trips cleanly (storage.cfg encryption key applied transparently);- PBS gotchas (
ignore-verified, node-from-UPID, privsep) touch only the verify-API path, not vzdump/restore.
Operator-signed decommission now reachable (slice 10 P3 completion). The previously-unreachable
IntentDecommissioned state (no production caller) is now reached ONLY via a gate-VERIFIED operator
signature — never customer-confirmable, distinct from a safe eject. New internal/signedjobs
DecommissionExecutor (op decommission, classified destructive in reconcile.Classify) calls
IntentStore.SetDecommissioned, keyed by the drive's STORAGE durable-id (the watchdog's key, e.g.
uuid:<fs-uuid> — NOT the device-level byid:/byuuid: scheme storage_wipe uses), so the recorded
intent actually gates future remounts. New ExecutorChain lets the signed-jobs runner serve both
storage_wipe and decommission; the runner wiring moved below the intent-store open in main.go.
felhom-opsign builds decommission params from -durable-id. No controller/customer UI — the operator
path is hub jobs-queue → signed-jobs runner.
Restore-test now boot-verifies slice-10 enrolled guests (bind-mount mountpoints). A guest whose
data drive is a host BIND mount (slice-10 P2 mp0) could not be vzrestore'd by the privsep token
("restoring 'mpN' to bind mount is only possible for root") — so the restore-test failed for every
enrolled guest, regardless of backup tier (surfaced during the felhom-pbs live validation). The
restore-test now reads the SOURCE guest config (vmid parsed from the archive volid — PBS ct/<vmid>/
and vzdump vzdump-lxc-<vmid>- forms) and passes RestoreLXCOptions.MountOverrides that neutralize
each bind-mount mpN to a throwaway 1G volume on the restore storage (needs no root; the boot-verify
doesn't need the drive's data, and the host paths would otherwise collide). Storage-backed mountpoints
are restored normally; best-effort (an unreadable source config restores as-is). proxmox.RestoreLXC
gained MountOverrides. Verified live: restore-test from felhom-pbs of bind-mounted guest 9201 →
boot+running PASS.
v0.27.0 — slice 10 P3: self-heal watchdog reconcile + 4-state intent model (2026-06-12)
The storage watchdog goes from detect-only → detect-and-reconcile: the agent autonomously re-mounts an enrolled external drive that dropped out-of-band (the colleague's Proxmox unmount), gated by a persisted INTENT model so it never auto-adopts an unknown drive or fights an official eject.
internal/storage/intent.go—IntentStore— durable, durable-id-keyed (UUID/WWN, never sdX/path), atomic-write 4-state model:new(not recorded → never auto-mount),enrolled(desired mounted → reconcile drift),ejected(intentional unmount → leave alone),decommissioned(permanent).OnAbsentclearsejected→enrolledso a replug auto-mounts (the replug rule). Records intent ONLY through the official enroll/eject paths — an out-of-band unmount records nothing and is healed. Tests cover the states, persistence, the replug rule, and the reconcile gate.watchdog.go— intent-gated reconcile + flapping guard (3C) — the re-mount candidate (device present, not mounted) now fires ONLY for anenrolleddrive (viaIntentReader); a present→absent transition (device gone) callsOnAbsent. Exponential backoff (debounce·2^fails) + an alert after 4 failed cycles + a hard stop after 8 (no infinite loop). Failure = "still not present a full backoff window after we dispatched" (a slow async re-mount isn't miscounted). Tests: colleague-unmount→ reconciled; ejected/new/decommissioned→left alone; ejected→absent→replug→auto-mount; flapping→caps.internal/localapi—POST /disks/guest-attachrecordsenrolled;POST /disks/ejectrecordsejected(BEFORE unmount, while the durable-id still resolves) via the newIntentRecorder.main.goopens oneIntentStore(<StateDir>/drive-intents.json) shared by the watchdog + local API; open failure degrades to ungated legacy remount (logged).
v0.26.0 — slice 10 P2 activation: guest-reboot endpoint (user-triggered drive activation) (2026-06-12)
A drive enrolled into a RUNNING unprivileged guest can't be live-activated (proven: pct set won't
hot-apply; /proc/<pid>/root bind → mount-locking refusal; nsenter -m loses the host source). So the
bind activates at the next guest boot. This adds the user-triggered restart path.
POST /guest/reboot(internal/localapi) — self-scoped (vmid from token). Runspct reboot <vmid>detached (it blocks ~30s until the guest is back) and returns 202 immediately, so the calling controller gets a clean response before the reboot takes it down (the agent is host-side and survives).GuestBinder.RebootGuestover the fenced runner. Tests:TestGuestReboot_Accepted(202 + RebootGuest invoked for the token's vmid),TestGuestReboot_CrossGuest403(body vmid mismatch refused, no reboot). Pairs with controller v0.49.0 (pending-activation detection + "Újraindítás most").
v0.25.0 — slice 10 P2: bind enrolled user-data drives into the guest (passthrough) (2026-06-12)
External user-data drives are mounted on the HOST but were never passed INTO the guest (diagnosed
Branch A), so apps silently wrote to the rootfs and the controller couldn't see them. This adds the
guest passthrough. Spike-proven on 9201 first (see REPORT / the usb-passthrough-spike findings):
pct set bind form (host path, never storage:size), chown to the guest base (idmap not clean
for mixed-ownership data), shared:49 propagation host↔guest automatic.
POST /disks/guest-attach(internal/localapi) — self-scoped (vmid from token). Binds an enrolled drive's felhom-data namespace into the guest at/mnt/<name>(Model A: the felhom-data dir is the bind source mounted AT/mnt/<name>, so only Felhom's namespace crosses into the guest — the customer's other data on the drive never does). Idempotent (returns the existing slot if already bound); picks the lowest freempN; validateswhereis/mnt/<name>(no traversal).GuestBinder(internal/localapi/guestbind.go) — the host-root steps over the fencedproxmox.Runner(same pattern as the provision back-half's bind):mkdir -p <drive>/felhom-data→chown 100000:100000the namespace ROOT (not -R; per-app subdirs are chowned at deploy) →pct set <vmid> -mpN <drive>/felhom-data,mp=/mnt/<name>(RW bind). The namespace is created fresh + uniformly owned, which sidesteps the drive's pre-existing mixed-ownership data entirely.- Tests —
TestGuestAttach_*: free-slot selection (mp0 when mp9 taken), idempotency (no re-bind +already:true), bad-path rejection (traversal/non-/mnt/multi-component), not-configured 503.
Pairs with felhom-controller P2C (enroll triggers attach) + the golden's /mnt:rslave controller bind
(P2B). Self-heal reconcile (P3) and dual-role (P4) follow.
v0.24.0 — role-gate the eject path (system/backup mounts are unmount-protected at the agent) (2026-06-12)
Closes the eject gap in the storage-authorization redesign: POST /disks/eject now refuses to
unmount a system or backup storage, enforced at the agent — not just hidden in the controller UI.
A direct API call (or a compromised controller) trying to eject {where:"/var/lib/vz"} or the PBS
mount is refused 403; only user-data mounts are ejectable.
handleDiskEject(internal/localapi/disks.go) — beforeUnmount, resolves the AUTHORITATIVE protection role of the storage mounted atwhere(the agent's own storage-view + host-topology classification, never the caller's claim) via the newroleForMountPath. Refuses (403, noUnmount) unless the role isuser-data. Fails SAFE: an unresolvable mount (view error or no storage target at that path) → treated as protected → refused (the same most-protected-on-ambiguity default the wipe gate uses). Mirrors the wipe path's "protected — eject refused by role" logging.roleForMountPath+hostReaderseam —roleForMountPathkeysRoleForStorageon the mount path (the eject input), mirroringdeviceRole.Options.HostReader(optional; defaults to the production*storage.ProcHostReader) injects the root-free topology reader so the role-gate is unit- testable.handleDisks/deviceRolenow share the same seam.- Tests —
TestEject_RoleGatedasserts asystemand abackupmount are refused with noUnmount, auser-datamount ejects, and an unresolvable mount fails safe to refused (the same non-hollowness the wipe tests use).TestEject_UnmountAndDependentsupdated to a user-data target.
v0.23.0 — device-ROLE classification + tiered storage-wipe gate (system/backup operator-only, user-data customer-confirmable) (2026-06-11)
The storage-authorization redesign (agent half). The gate's destructive-wipe path is now tiered by the device's protection ROLE, which the agent classifies from its OWN inspection — never the caller's claim (the storage analog of classify.go's data-bearing verdict).
internal/storage/role.go—DeviceRole(system|backup|user-data) + the authoritative classifier.RoleForStorage(storage-view targets) andRoleForRawDevice(a raw device, e.g. a fresh disk in the init flow) map a device to its tier viaSystemDisks(the whole-disks backing/,/boot,/boot/efi, root-free reads). Rules:pbs→ backup;lvmthin/ builtinlocal/ nfs / cifs / unknown → system;usb/local-diron a non-system external device → user-data. Fail-safe: any ambiguity (system disks unknown, or an unrecognizable device topology) → system (most-protected) — never silently user-data.GET /disks— eachDiskInfonow carriesrole. The controller drives the UI from it (system/backup get a lock + no destructive controls; user-data is customer-manageable).- Gate tier (
reconcile) — newCustomerConfirmabledisposition +Gate.AuthorizeStorageWipe:- role=user-data → customer-confirmable: allowed iff the request carries an explicit
customer confirmation bound to the device's durable id (the agent re-resolves the durable id
and matches; a confirmation for one disk can't wipe another). No operator signature. A
user-data drive is already within the in-guest controller's blast radius (it bind-mounts
/mnt), so customer-confirmation adds no new reach. Recorded in the audit log with the durable id (AuditRecord.DurableID). - role=system/backup → unchanged operator-signature (
pending_signature). Theconfirmedflag is IGNORED — a compromised controller assertingconfirmed:trueon a protected device is refused by role. Every other destructive class (guest_destroy,decommission,restore_overwrite,key_rotation) keeps operator-signature exactly as before.
- role=user-data → customer-confirmable: allowed iff the request carries an explicit
customer confirmation bound to the device's durable id (the agent re-resolves the durable id
and matches; a confirmation for one disk can't wipe another). No operator signature. A
user-data drive is already within the in-guest controller's blast radius (it bind-mounts
POST /disks/format— acceptsconfirmed+durable_id(inert for system/backup). The data-bearing path tiers by role: user-data customer-confirmed →mkfs; user-data unconfirmed → 403needs_confirmation(+ the durable id to confirm against, NOT an opsign command); system/backup → 403 with the operator-signature pending op (as before). Blank devices stay benignmkfs.- Tests —
role_test.go(demo-storage mapping + fail-safe),storage_wipe_test.go(the gate refuses aconfirmedwipe on system/backup → no exec; durable-id mismatch / missing-durable refused; unknown role fails safe), and the localapi format-handler branches (user-data confirmed → mkfs; user-data unconfirmed → needs_confirmation, no opsign; confirmed-but-protected → still refused).
Pairs with the controller's lockout + type-to-confirm UX + drive-list restyle.
v0.22.0 — expose durable_id in GET /disks (enable controller-side guided storage) (2026-06-11)
One-line, read-only addition: localapi.DiskInfo gains durable_id (mapped from
StorageTarget.DurableID, e.g. "uuid:<fs-uuid>" for usb/local-dir). The de-privileged controller
cannot read a device's fs UUID itself, yet POST /disks/assign mounts strictly by UUID — so without
this it could not complete the guided init/attach flows. The controller strips the uuid: prefix to
get the assign key. No new privilege, no behaviour change to format/assign/eject or the data-bearing
gate. Pairs with felhom-controller v0.43.0 (the storage-management UI rebuild).
v0.21.0 — agent-managed split-horizon LAN resolver (internal/lanresolver) (2026-06-11)
LAN clients can now reach their guest directly at the same public hostname with the same real wildcard cert (no Cloudflare hairpin), via a host-side dnsmasq the agent manages. The host is the stable anchor (static LAN IP); the guest stays DHCP/ephemeral and the agent tracks its live IP.
internal/lanresolver— renders a dnsmasq base drop-in (bind to the host LAN IP, no-resolv, upstreams) + a per-customer drop-inlocal=/<domain>/+address=/<domain>/<guest-ip>. The proven two-line shape:local=makes dnsmasq authoritative for the zone so AAAA returns NODATA (no Cloudflare-AAAA split-brain — the guest has only link-local v6),address=is the wildcard A; all other names (and their AAAA) forward upstream unchanged.Managerensures dnsmasq present (apt) + the base config + enabled, discovers the guest's live IPv4 (pct exec <vmid> -- ip -4 -o addr show dev eth0) and domain (read from the guest controller's pulledcontroller.yaml— the v2 bootstrap omits it), writes drop-ins write-if-changed, and reloads (not restarts) dnsmasq. Tolerates the early-boot pre-lease window (empty IP → skip+retry, never a blank record). Logs IP transitions.Loop— a 7th daemon goroutine: every interval (default 300s) it enumerates provisioned guests (/var/lib/felhom-agent/guests/<vmid>/) and reconciles each, so the resolver follows DHCP IP changes. Configlan_resolver.{enable,host_ip,upstreams,interval_seconds}(host_ip defaults to the local-API bridge IP).--selftest=lanresolver -vmid N.configs/felhom-agent.sudoers— newFELHOM_DNSMASQalias (apt install dnsmasq; install felhom-.conf drop-ins; systemctl enable/reload dnsmasq; rm felhom-.conf; the two FIXEDpct execreads). The agent never touches/etc/resolv.conf(host's own resolution unaffected).- Box-down robustness is a documented router config (DNS = [host-IP primary, upstream secondary]) so a box reboot degrades to the Cloudflare path, not total DNS loss — see REPORT install step.
- Spiked live on felhom-pve first (
:53free, host IP static192.168.0.162, host DNS intact, full loop from a real LAN client returned the guest IP + AAAA NODATA + the real wildcard cert200 0).
v0.20.0 — golden: stacks-dir bind + per-guest hostname/CT name + bake base-infra images (2026-06-11)
Lockstep with felhom-controller v0.41.0 + a golden rebake. Changes in configs/build-golden.sh and
the provision path; no change to the proxmox/authz/token fences.
- Section-G mount fix (the load-bearing one): the in-guest controller writes app/infra compose
stacks under
/opt/docker/stacksinside its container, but the baked controller-bootstrapdocker runnever bind-mounted that path. Sodocker compose up(run by the GUEST daemon over the shared socket) resolved every relative bind source on the guest filesystem — silently creating empty dirs — which broke every bind-mounted stack (base infra AND customer apps like immich/nextcloud). The bootstrap unit nowmkdir -p /opt/docker/stacksand adds a same-path host bind-v /opt/docker/stacks:/opt/docker/stacks(a named volume would NOT fix this). Empirically confirmed on guest 9201 before writing the fix. - Per-guest container hostname (3A): the bootstrap unit derives
customer.idfrom/etc/felhom-bootstrap/bootstrap.jsonwith a portablesedparse (NO jq in the golden) and passes--hostname <customer-id>todocker run, so the controller'sos.Hostname()(its hub-reported hostname) is the customer id, not the Docker container ID. Fail-safe: no parse → no--hostname. - Per-guest CT/LXC name (3B):
--selftest=provisionnow defaults-hostnameto the (DNS-safe sanitized)-customer-idwhen not given, so the bring-up's existingSetConfig hostnamestep (bringup.go) names the CT meaningfully (e.g.demo-felhom) instead of inheriting the golden'sfelhom-golden. NewsanitizeHostname(lowercase, collapse invalid →-, trim, ≤63). - Bake base-infra images: the golden now also pulls the three PINNED, PUBLIC base-infra images
(
traefik:v3.6.7,cloudflare/cloudflared:2026.6.0,gtstef/filebrowser:1.3.3-stable) into its Docker storage so the controller's first-boot bring-up is OFFLINE-capable. A hard gate (docker manifest inspect) fails the bake early on a bad pin. Tags MUST match the controller'sinternal/infraconstants.
v0.19.0 — bootstrap contract v2: agent relays the hub retrieval passphrase (no host key in the guest) (2026-06-11)
Lockstep with felhom-controller v0.40.0. Fixes the onboarding 401: a freshly provisioned guest's
controller used to come up with the agent's host hub key baked in, which the hub's /api/v1/report
(customer-scoped auth) rejects. The agent now bakes a v2 bootstrap carrying only what the controller
needs to pull its own config from the hub — the agent never touches the customer-scoped key or CF
tokens.
Changed — bootstrap contract v1 → v2 (internal/provision)
SchemaV1 → SchemaV2 = "felhom.bootstrap/v2".DocCustomerdropsname/domain/email(keepsid).DocHubdropsapi_key/host_id, addsretrieval_password(the customer's hub retrieval passphrase — SECRET).DocLocalAPIunchanged. The contract is byte-compatible with the controller'sinternal/bootstrap.Bootstrap(cross-repo round-trip verified).backhalf.go: renders the v2 Doc; validation now requirescustomer.id+hub.url+hub.retrieval_password(wascustomer.id+customer.domain). Write/0600/chown/pct setunchanged.cmd/felhom-agent/main.go--selftest=provision: new required-hub-passwordflag (the customer's hub retrieval passphrase; the customer must already exist in the hub). Stops bakingcfg.Hub.APIKey/cfg.Hub.HostID.-customer-domain/-name/-emailstill accepted (bring-up may use them) but NOT baked.
Changed — configs/build-golden.sh
- Default
CONTROLLER_IMAGEbumped off the stale:v0.35.0→:0.40.0(matches the registry's no-vtag convention; latent footgun fixed).
Tests
doc_test.go/backhalf_test.goupdated to the v2 shape (assert noapi_key/host_id,retrieval_passwordpresent,customercarries onlyid).go build ./... && go test ./...green.
v0.18.0 — slice 10D: DR capstone — identity escrow + restore-mode consumption (agent side) (2026-06-10)
The agent half of the slice-10 DR capstone (closes slice 10). Grounded by both 10-series spikes (escrow-consumption + identity-restore). The hub half (recovery-mode toggle, re-enroll + credential rotation, directive serving) is hub v0.11.0. Operator-side rotation model (locked): the hub holds no Cloudflare write-power; the destructive tunnel/PBS rotation is the operator's step from a trusted environment (same spirit as 10B).
Added (internal/escrow)
- Identity escrow (
identity.go):WrapIdentity/UnwrapIdentity(+…Bundle) wrap the{tunnel_token, pbs_token}bundle under the SAME recovery codeRviaage(scrypt + ChaCha20-Poly1305 — a vetted passphrase-AEAD, not hand-rolled), reusing the K-escrow pty mechanism (passphrase via the tty, data via files;R/tokens never logged). Same two-factor, zero-knowledge shape as the K-escrow. A wrong R fails closed (no bundle).ageis a runtime dep for the identity path (analogous to proxmox-backup-client for K). escrow.Creategains an optionalIdentityBundle→ also emits anIdentityBlobunder the same R (additive; the K-escrow + 10CConsumepaths are byte-unchanged). Self-verifies the identity round-trip before shipping.--selftest=escrow-create -identity-bundle <file> -directive <file>— also wrap + upload the identity blob + the non-secret DR directive (pbs repo/ns, expected key fingerprint, tunnel id).--selftest=identity-consume -blob <file> -keydest <file>(R viaFELHOM_RECOVERY_CODE) — recover the identity bundle through the real code; tokens written 0600, never logged.
Tests
- identity bundle round-trips (wrap→unwrap byte-identical; blob is opaque ciphertext); wrong R fails
closed + the blob stays retryable; input validation. K-escrow/10C tests byte-unchanged (additive).
(age integration tests gated to a host with the
ageCLI.)
v0.17.0 — slice 10C: escrow consumption (productionize the spike) (2026-06-10)
Turns the throwaway 10C spike harness into a real, tested Consume path: recover the PBS key
K from an R-wrapped escrow blob, gate it on the expected fingerprint, and install it for the
restore. The spike already proved the crypto + real-data restore; this bakes its findings into
production code. Agent-only — 10C reads the four inputs as parameters (so it stays
standalone-testable); 10D sources blob/fingerprint/PBS-connection from the hub and prompts for R.
Zero-knowledge holds: the hub serves everything except R (by hand from the customer), so a
hub compromise alone still can't decrypt.
Added
escrow.Consume(ctx, blob, R, expectedFingerprint, keyDest)— the consumption contract:- Unwrap the blob (a copy — F-C6: the input blob is read-only → a failed Consume is
retryable) with
R; a wrong R fails closed at the scrypt KDF (F-C3) → a clear, R-free error, nothing written. - Fingerprint gate (F-C4) —
KeyFingerprint(recovered)must equal the expected (the hub knows it); a mismatch fails fast + loud, no install, no restore attempted. - Atomic install (F-C2) at
keyDest(0600, write-temp-sibling→rename); any failure leaves no partial install. The recovered key lives only in a0700tempdir that is always removed. Secret discipline:Rand key bytes are never logged/persisted (only fingerprint prefixes);Kis never mutated.
- Unwrap the blob (a copy — F-C6: the input blob is read-only → a failed Consume is
retryable) with
--selftest=escrow-consume(-blob -fingerprint -keydest, R via envFELHOM_RECOVERY_CODEto keep it off the command line) — invokes the realConsumelive (the spike's S3 via the production path, not a harness).
Tests (non-hollow)
- valid → key installed +
KeyFingerprint(dest) == expected+0600+ blob byte-unchanged; wrong R → error, no file at dest, blob unchanged; fingerprint mismatch → fail fast, no install (the gate runs before any restore); input validation; format-tolerant fingerprint compare (no empty-fingerprint gate-bypass); atomic-install permissions (integration tests gated to a host withproxmox-backup-client).
v0.16.0 — slice 10B: operator-signed destructive completion (offline key + signing CLI) (2026-06-10)
The security centerpiece: a destructive op runs ONLY on a verified, operator-signed authorization
— signature valid against a pinned operator pubkey (never the hub's or the blob's), nonce
unseen + durably burned, in-window, host-bound, and resource-bound to a DURABLE device id that
execution re-resolves + re-inspects. Decision (a): offline operator key + signing CLI,
hardware-key-ready (sk-/YubiKey via ssh-keygen). The key floor holds: the signing key is NOT in
the hub and NOT in the agent. Concrete consumer: this closes the 8C data-bearing-wipe
pending_signature gap. Pairs with hub v0.10.0.
Added
cmd/felhom-opsign— the operator's offline signing CLI. Builds the canonicalOpBlobby reusingauthz.CanonicalBlob(the exact production path the verifier authenticates over — so signer + verifier can never drift) and signs it withssh-keygen -Y sign -n felhom-op-v1(hardware-ready). Output: a{op_blob_b64, sig_armored}envelope to hand to the hub jobs queue (optional--upload). Touches ONLY the operator's signing key.authz.CanonicalBlob— promoted to production (was test-only) so the CLI + verifier share one canonical-bytes source; params canonicalized (sorted keys, compact).internal/storagedurable device identity (durable_device.go):DeviceDurableID(derive a stablebyid:(wwn/serial)/byuuid:id from the world-readable udev symlinks — no privilege, no subprocess) +ResolveDurableDevice(re-resolve to the current/devpath; a path-only/unknown scheme is REFUSED). The resource-level anti-retarget.internal/signedjobs(new): the queue consumer.Runnerfetches each opaque job → runs it through the gate (the LOCKED authz pipeline) → on all-pass hands the verified op to anExecutor; the order is verify → nonce-burn (durable, in Verify) → execute → clear job. TheWipeExecutoris the 8C consumer: resolve the signed durable id → re-derive + match (anti-retarget) → re-inspect (8C classifier) the device is still the data-bearing target →mkfs. A vanished/changed/non-data-bearing device or a path-only binding is refused even with a valid signature. Wired as a secondEnvelopeObserver(runs onHasSignedOps).hub.Client.Jobs/CompleteJob+hub.MultiObserver; the 8C format refusal now surfaces the bound op (op + durable id + host) in its 403pending_op+ afelhom-opsign …hint.
Pinning / rotation
- Operator pubkeys are pinned via
authz.signers(config, trusted path — provision/agent config, NEVER hub-alone), multiple keys (KeyID selects; role-scoped), so a backup/rotation key exists without a flag-day. Unchanged from the slice-4 verifier wiring; 10B activates the execute path.
Tests (real crypto, non-hollow)
signedjobsrunner over the real gate+verifier (in-Go minted SSHSIGs): valid → executor runs once + job cleared; replay (nonce burned) / non-pinned signer / expired / retarget (other host) / forged sig / no pinned signer → all rejected, executor never called; malformed envelope cleared.WipeExecutor: valid →mkfsruns; path-only, durable-id mismatch, device gone, re-inspect non-data-bearing, not-probed → all refused,Formatnot called.storagedurable: wwn-preference, uuid-fallback, path-only/traversal refusal, round-trip, missing-device error (symlink tests gated to Linux — the agent's OS).
v0.15.0 — slice 10A: hub desired-state serving — the "Down" channel (2026-06-10)
The agent half of slice 10A. The control envelope (hub.ControlEnvelope) stops being "reserved — ignored" and becomes the live Down channel: a cheap change-notification on every heartbeat. The agent caches the hub's desired-state + its generation; only when DesiredGeneration advances does it fetch the full state (the heartbeat stays light, the heavy state moves on change). The engine then reconciles benign deltas and the gate marks an explicit destructive delta pending_signature (no signer in 10A → never executed; signed execution is 10B). Pairs with hub v0.9.0.
Added / changed
internal/reconcile:DesiredGuest.Decommission— the canonical destructive desired-state delta (an EXPLICIT flag, not "absent from the list", so a partial hub list can never mass-destroy). The planner emitsActionDecommission→ClassDecommission→ Destructive → the gate refuses itpending_signature.Reconcilenow counts apending_signaturerefusal asResult.Pending(expected, logged INFO) rather than a failure; any other refusal stays a real failure.ActionDecommissionhas no executor (slice 10B) — a defensive guard refuses to run it. NewCachingProvider(thread-safe DesiredState + generation cache;Desired/Update/Generation) — the productionDesiredProvider, replacingEmptyProviderin the daemon engine (empty until the hub serves intent → cold-start is a live no-op, unchanged).internal/hub: theControlEnvelopefields are now active (DesiredGeneration drives the fetch, HasSignedOps noted). New wire typesDesiredStateResponse+WireDesiredState(guests + forward-compatrestore_directive(10D) /pbs_namespace/ opaquestorage_manifest+backup_policy) +WireDesiredGuest(vmid/run/spec/description/decommission). NewClient.FetchDesiredState(GET/api/v1/hosts/{host_id}/desired-state, self-scoped to the client's own host). NewEnvelopeObserverloop seam +SetEnvelopeObserver— the loop hands the envelope to the sync layer each cycle (hub does not import reconcile/desired).internal/desired(new): theSyncer— implementshub.EnvelopeObserver, fetches desired-state on a generation advance, maps the wire shape to the reconcile domain, and updates theCachingProvider. Caches the fetched generation (robust to a generation that advanced mid-fetch); a fetch failure keeps the last-known state.restore_directiveis carried + logged, not acted on (10D). Wired incmd/felhom-agent(daemon): provider → engine, syncer → loop.
Tests
- reconcile: a desired-state with one benign + one decommission delta → benign applied, destructive gated pending (not executed);
Planemits decommission-only for a decommissioned guest + classifies Destructive;CachingProviderupdate/isolation. - desired: fetch-once-on-advance (no re-fetch on an unchanged generation), fetch-failure-keeps-cache, caches-the-fetched-generation.
- hub client:
FetchDesiredStatehits the self-scoped path with the bearer + decodes (incl.restore_directive); a 403 is a typedHTTPError. - loop: the cycle notifies the observer + adopts
PollIntervalSeconds; a report error skips the observer. - cross-repo golden:
testdata/desired-state.golden.json+control-envelope.golden.jsondecode + key-set guard, byte-identical with felhom.eu/hub.
v0.14.0 — slice 9: host metrics to the controller (GET /host/metrics + CPU-temp collector) (2026-06-10)
The de-privileged controller (slice 8C) sees only its own cgroup, so it can't read host health itself. Slice 9 re-serves the slice-4 collector's host + per-storage view to the customer over the local API, plus the one missing collector — CPU/chassis temperature — so the customer sees their box's health in the controller. Host-wide, token-authed, fresh (a live collect, not the 15-min hub snapshot). Assumption: one customer per host (the home-server model); if a host ever serves multiple customers, host-wide CPU/mem would leak cross-customer load → revisit then.
Added / changed
- CPU/chassis-temp collector (
internal/hub/cputemp.go):SysfsTempReaderreads the CPU package temperature straight from sysfs — hwmon (coretemp/k10temp/zenpower/cpu_thermal, preferring thePackage id 0input) then the thermal zones (preferringx86_pkg_temp/coretemp/cpu-thermal, falling back toacpitz). No external binary, no privilege (sysfs nodes are world-readable), so the root-CLI fence is untouched. Graceful-null: a missing sensor, an unsupported board, an implausible reading (outside 5–150 °C), or any read error all degrade tonull("n/a") — a missing sensor never fails the report. Wired into the collector via the newTempReaderseam (nil-safe). HostMetrics.CPUTempC *int(cpu_temp_c) — new nullable wire field on the sharedHostMetricsstruct (same nullable contract as the diskSmartSummary.TemperatureC). It rides the hub report too (operator freebie) → cross-repo host-report golden updated.Collector.HostMetricsNow(ctx)— a freshNodeStatus+ CPU-temp read returning just the host block, the source for the local API (current cpu%/temp, not the 15-min snapshot).Collect()now also populatescpu_temp_con the hub report.Collector.SetTempReaderinjects a fake in tests.GET /host/metrics(internal/localapi/host_metrics.go): host-wide health (cpu%/mem/load/uptime/cpu_temp_c) + per-storage capacity targets (total/used/fraction, thin-pool, SMART temp+wear). Token-authed viawithGuest(host-wide data; cross-guest?vmid=still 403). Best-effort on storage (a view error still returns the host block). Served only when theHostMetricsprovider (the shared collector) is wired — else 503 "not configured". Wired inbuildLocalAPIServer.
Tests
cputemp_test.go: a fake/syslayout proves hwmon package-preference, hwmon first-input fallback, thermal-zone-by-type selection over a non-CPU hwmon, graceful-null on a sensorless host (no error), and rejection of implausible (0 m°C) readings.hostmetrics_test.go:HostMetricsNowpopulates the temp, gracefully nulls it, hard-errors onNodeStatusfailure;Collect()carries the temp.host_metrics_test.go(localapi): populated host+storage with a valid token;cpu_temp_c:nullserializes; 401 without a token (collector never invoked); 403 on a cross-guest?vmid=; 503 when not configured.
v0.13.0 — slice 8B.2: quiesce downtime optimization (snapshotted phase) (2026-06-10)
The agent half of slice 8B.2. In snapshot mode, vzdump only needs the app-stopped state captured at
the storage-snapshot moment; after that it reads from the snapshot and the app can resume. The
agent now emits a snapshotted phase on GET /backup/status when the snapshot is taken, so the
controller (v0.38.0) resumes its app early — app downtime drops from whole-backup to
until-snapshot with no loss of app-consistency. Validated Phase-0 first on PVE 9.2.2: the marker is
INFO: create storage snapshot 'vzdump'; downtime ~24s→~1s for a 934 MB guest.
Added / changed (internal/backup + internal/localapi)
BackupRunner.BackupWithSnapshotHook(ctx, vmid, onSnapshot)— while the vzdump runs, a watcher tails the task log (TaskLogTail) for thecreate storage snapshotmarker and firesonSnapshotonce. The marker only appears in snapshot mode (stop/downgraded takes no storage snapshot), and the watcher also bails onbackup mode: stop— so it never fires in stop mode. (Backupkeeps its signature for the scheduler/selftest; both share one body.)/backup/statusphasesnapshotted(betweenrunninganddone):handleBackuppasses the hook →markSnapshottedflips the running job tosnapshotted.done/failedsemantics unchanged.
Tests
- localapi: snapshot mode → phase reaches
snapshottedbeforedone(gated fake holds the backup open); stop mode →snapshottednever emitted (stays running → done). runner: the watcher firesonSnapshoton the marker; in stop-mode log it never fires.snapshotWatchIntervalis a package var so tests run fast.
v0.12.0 — slice 8C Phase A: disk endpoints + data-bearing classifier gate + mkfs executor (2026-06-10)
The agent half of slice 8C, Phase A (additive). Adds the host disk-management endpoints the
controller's disk UI drives — with the 8C security invariant: the agent decides
data-bearing-ness by inspecting the actual device (agent-internal evidence), NEVER from the
caller's claim. A compromised controller asserting "this drive is blank" cannot wipe a data-bearing
drive. (Controller rewire + disk-subsystem retirement + de-privilege are Phases B/C, felhom-controller.)
Added
internal/storage—mkfsexecutor + data-bearing inspection.SudoHostOps.Format(device, fstype)(device-pinned,ValidateBlockDevice+ValidateFSType, narrowFELHOM_FORMATsudoers —mkfs.ext4 -F/mkfs.xfs -fon a/dev/*path the agent fine-validates first).SudoHostOps.InspectDevice(device)→DeviceProbe(filesystem signature viablkid -p, partition table / partitions / mount vialsblk -J).DeviceProbe.DataBearing()is conservative: any signature / partition table / partition / mount — OR a probe that did not read cleanly — is data-bearing (fail-safe; an unreadable device is never called blank).internal/localapi— the §6 disk endpoints, all self-scoped (token→guest; cross-guest 403):GET /disks— host drives + a data-bearing flag (UI hint). Read-only/benign.POST /disks/assign— attach a drive as a mount (benign, additive →EnsureMount). Self-serve.POST /disks/eject— safe-unmount (benign, data preserved) + the dependent guests that mount it (so the controller can warn which apps lose that storage).POST /disks/format— the security centerpiece: the agent inspects the device itself; blank → benign →mkfs; data-bearing → ClassStorageWipe → the slice-4 gate → refusedpending_signature(the operator-signed completion is slice 10). The caller's claim is ignored — only a device the agent reads as blank is formatted.
storageGateAdapterbridges the format path to the slice-4 reversibility gate (no new gate/crypto).
Tests
- localapi (security matrix): blank device → mkfs called, gate not consulted; a data-bearing
device → 403, mkfs NEVER called, gate consulted (
pending_signature); an ambiguous/unprobed device → treated destructive (fail-safe); even a gate that allows does not format data-bearing in 8C; assign →EnsureMount; eject →Unmount+ dependent guests; cross-guest → 403; bad device/fstype → 400; unconfigured → 503. - storage:
ValidateBlockDevice/ValidateFSType(whitelist + injection rejection);InspectDeviceblank/filesystem/partition-table/mounted/failed-probe-fail-safe;Formatinvokes the rightmkfs.*.
v0.11.0 — slice 8B: app-consistent backup — /backup/due policy + /backup/status phases (2026-06-10)
The agent half of slice 8B (doc 03 §8). Turns the 8A thin backup stubs into the real policy the
in-guest controller's quiesce loop drives (controller half: felhom-controller v0.36.0). No hub
change. The downtime optimization (vzdump --mode snapshot + a snapshotted phase) is the 8B.2
fast-follow; the hub-served per-guest policy is slice 10.
Changed (internal/localapi)
GET /backup/due— real cadence policy (replaces the 8A "never backed up" stub): a guest is due when no successful backup is recorded OR the newest one is older than the agent-local cadence (backup.backup_cadence_seconds, default 24h). A successfulPOST /backupflips due to false for the window, so the controller won't re-quiesce in a loop. A failed backup does not satisfy the cadence. Returnsage_secondsfor diagnosis.GET /backup/status— real phasesidle | running | done | failed+ the job id, so the controller can poll a backup to completion (was: just the latest stored backup).POST /backup— returns a job id +runningphase; tracks the in-flight job and is single-flight per guest (a second POST while one runs returns the same job — no concurrent vzdump). On completion the job transitions done/failed and the result is recorded to the store.- Config:
backup.backup_cadence_seconds+BackupCadence(); the local-API server takes the cadence.
Tests
/backup/due: due when stale / no backup, not due within the window after a success, due again past the cadence, a failed backup does not count./backup/status: running→done and running→failed (gated fake to observe the running phase).POST /backupsingle-flight (one vzdump for concurrent POSTs). All still self-scoped (token→guest).
v0.10.0 — slice 8A: agent local-API server + provisioning back-half (2026-06-10)
The host-agent half of slice 8A (doc 03 §6). Adds the per-guest local API the in-guest
controller calls over the bridge, and the provisioning back-half that follows the slice-7
bring-up front half. Grounded by felhom.eu/documentation/tests/slice8a-channel-deploy-spike-findings.md
(commit 4a81a96 — channel + deploy plumbing proven; the 5 gotchas resolved here). Controller half
is felhom-controller v0.35.0. No hub change.
Added
internal/localapi— the HTTPS local-API server (doc 03 §6), the per-guest authorization gate. Serves a persisted self-signed leaf with a stable SHA-256 fingerprint (generated once; a fresh cert each boot would invalidate every baked bootstrap pin). The 7 §6 endpoints, all self-scoped to the caller's own guest:GET /storage(this guest's mpN mounts + fast/slow class from the slice-5/7 storage view),POST /snapshot,POST /rollback,POST /backup(enqueued, crash-consistent — the app-consistent quiesce loop is 8B),GET /backup/due(thin in 8A),GET /backup/status,GET /restore-test/status.- Token store (
tokenstore.go): durable, crash-safe per-guest token→guest map that persists only a SHA-256 hash of each token (the plaintext exists transiently at mint→write-to-mount, then is discarded), last-write-wins per guest, fsync'd append-only JSONL (mirrors the nonce store). - Self-scoping: the VMID is resolved ONLY from the token; an explicit
vmid(query/body) that disagrees → 403 and the proxmox op is never issued for the other guest; absent/unknown → 401.
- Token store (
internal/provision— the back-half: mint the per-guest token → render the stablebootstrap.jsoncontract (schemafelhom.bootstrap/v1; no registry credential — the controller image is baked into the golden) → write it0600→chown 100000:100000(the unprivileged-LXC mapped guest-root, spike gotcha 1) → attach a read-only bind mount viapct set. Host-side only (F3 — the agent never enters the guest; nopct exec). The token plaintext is never logged and never returned.--selftest=provision— the full chain on-demand: bring-up (provision) front half + the back half; keeps the guest for the golden's baked controller-bootstrap unit to deploy.config.LocalAPIConfig(local_api) — enable + bridgelisten_addr+ cert/key paths + token store path. The server is an optional 6th daemon goroutine, disabled cleanly when unconfigured or on a token-store/cert failure (the daemon still reports/reconciles).configs/build-golden.shnow bakes the controller image (pulled once on the trusted build host, thendocker logout— no cred baked) + a controller-bootstrap unit that deploys the baked image from the config mount on boot (no login/pull at deploy).configs/felhom-localapi-firewall.example— host firewall narrowing of the local-API port to the guest bridge subnet (nft/iptables/PVE variants; defense-in-depth — the token stays the gate).configs/felhom-agent.sudoers— a narrowFELHOM_PROVISIONalias (chown 100000:100000+pct setbind-mount, both confined to the agent-owned/var/lib/felhom-agent/guests/*path) for the non-root least-privilege deployment.
Security / design notes
- The local-API leaf is pinned by leaf-cert SHA-256 (decision: consistency with the agent's PVE/PBS pinning); the fingerprint is baked into each guest's bootstrap.
- The back-half's host-root ops (chown + bind-mount attach) are NOT added to
proxmox.Privileged(which is fenced to its 3 exceptions) — they live ininternal/provisionand run through the sharedRunner(direct as root, orsudo -nwith the new sudoers alias). This is the per-guest provisioning host-root surface, host-side and F3-compliant.
Tests
- localapi: self-scoping (cross-guest snapshot/rollback/backup → 403, op never issued for the other
guest; own-guest uses the token's VMID), 401 paths,
/storageclass mapping,/backupenqueue, the thin/backup/due, status scoping; the token store persists only the hash (plaintext never on disk), last-write-wins, survives reopen, uniqueness; the leaf fingerprint is stable across reload. - provision: writes
0600+ chowns + attaches the bind mount with the right args; the token never appears in the Result; the cross-repobootstrap.jsoncontract key-set is pinned.
v0.9.0 — slice 7 close-out: PBS recovery-code escrow creation (2026-06-10)
The first code that touches the PBS client encryption key K and introduces the customer recovery
code R. Default posture is zero-knowledge: Felhom holds an opaque R-wrapped blob (cannot
open it), the customer holds R. Grounded by felhom.eu/documentation/tests/slice7-escrow-spike-findings.md
(round-trip proven on a throwaway: the R-recovered key restores a real encrypted snapshot). Hub
opaque storage is the felhom.eu half (hub v0.8.0); consumption/serving is slice 10.
Secret discipline (overriding)
R is crypto/rand, ≥128 bits, surfaced exactly once and never logged/persisted/committed;
the wrap pty's echo is discarded so R can't leak. K is read by location, never modified (the
live key file is byte-unchanged — Wrap operates on a copy), never logged.
Added
internal/escrow—CreategeneratesR(10 EFF-wordlist words ≈ 129 bits), wrapsKunderRvia the PBS-nativeproxmox-backup-client key change-passphrase --kdf scrypt, and self-verifies the blob recoversK(fingerprint match) before shipping. The wrap is driven over a stdlib pty (x/sys/unix; spike F-A1 — the command is TTY-only) with output discarded (F-A2 — the pty echoes the passphrase). Opt-in outputs: (b)R-wrapped offline copy (two-factor, no extra trust) and (a) raw paperkey (single-factor, unrevocable — loud caveat).--selftest=escrow-create(-storage,-paperkey,-offline,-upload): surfacesRonce to stdout (never the logger), prints the opaque blob's size/fingerprint/posture, and with-uploadPUTs the blob to the hub (/api/v1/hosts/{host_id}/escrow, per-host key).- Config:
escrowsection (posturedefaultzero_knowledge,pbs_storage_id);PBSEncKeyPathhelper (the<id>.enckey K). - Runtime dependency on the
proxmox-backup-clientCLI (the PBS key+passphrase KDF).
Tests
Rentropy ≥128 / 10-word format / uniqueness; integration round-trip (wrap→unwrap fingerprint match, wrong-Rfails, liveKbyte-unchanged, blob ≠ plaintext key) guarded to linux+proxmox-backup-client; the agent→hub wire-contract key-set (mirrors the hub's).- Live-validated (demo):
escrow-create→R(10 words) surfaced once, blob 383 B opaque, self-verify ok, liveKsha256 unchanged, exactRabsent from stderr/journal.
v0.8.0 — slice 7 Phase 1: unified bring-up reconcile job (provision + guest-loss DR) (2026-06-09)
The shared FRONT HALF of provision and guest-loss DR, as a journaled reconcile job mirroring the
slice-6 restore-test's crash-safety — but it KEEPS the guest on success and applies a
scenario-specific identity policy. Agent-only; no hub/wire change (the new guest auto-appears in
the host-report via ListLXC). Grounded by the slice-7 bring-up spike findings (commit 3342993):
F1 (restore preserves the archived MAC → provision reset is unconditional), F3 (SSH host keys do
not auto-regenerate → a baked golden first-boot unit, not an agent guest-internal op), F4 (the
transient PVE config-lock 500 → bounded retry).
Added
reconcile.RunBringUp(bringup.go) —BringUpSpec(Modeprovision|dr_guest_loss, Archive, VMID, RestoreStorage, Hostname, Cores/MemoryMB, RootfsGrowGB, Mounts, KeepMAC, BootTimeout) →BringUpResult(VMID, AssignedMAC, Pass, Verified, StartWarnings/Recognized). Sequence (each mutation preceded by journaling the owning entry): restore → identity reset → size → attach mounts → start LINK-UP. Verdict is liveness (waitRunning), never the start exitstatus (reuses the v0.7.0 WARNINGS surface). Success KEEPS the guest (no teardown).- Scenario-specific identity reset (doc 03 §9): provision → fresh MAC unconditionally
(
PUT net0withhwaddromitted → PVE regenerates, F1) + hostname; machine-id + SSH host keys regenerate guest-side on first boot (golden bake + the new unit) — the agent does NOT touch guest internals. dr_guest_loss → preserve continuity (keep hostname; keep MAC unlessKeepMAC=false); never resets restic/tunnel/hub identity. - Compensating rollback — any mid-flight failure destroys the just-created guest
(
ClassGuestDestroy, benign viaProvenance{SameTxnCreated:true}, gated); on teardown failure the entry is left in-flight forRecover. New journal flagRollback+Recover'srecoverBringUpreap a half-built guest left by a mid-job crash (idempotent, viaListLXC). - F4 config-lock retry — steps 3+5 coalesced into ONE
PUT config(net0+hostname+cores+ memory+mpN); rootfs grow stays its own call.setConfigWithLockRetryretries ONLY the transient PVE config-lock 500 (pveConfigLock: 500 + "can't lock file"/"got timeout"); any other error fails immediately — never retried. --selftest=bring-up(-mode provision|dr -archive -vmid -hostname [-keep]) — runs the real journaled job (after aRecover), then tears the guest down unless-keep.configs/build-golden.sh— the validated golden recipe as a script, incl. the F3 first-bootfelhom-regen-hostkeys.serviceunit (Condition-gated: fires on provision, no-ops on DR). The slice-7 spike archive (which lacks the unit) is superseded.
Deferred (stated, not built)
- Provisioning BACK HALF (controller deploy, bootstrap, per-guest token mint) → slice 8.
- Host-loss DR + PBS escrow consumption → slice 10.
- The SOURCE of a
BringUpSpec(hub desired-state: which archive/VMID/mounts) → slice 10; this job takes the spec as input.GuestMountis defined minimally (no hub coupling).
Tests
- provision happy path (fresh MAC = net0 without hwaddr, hostname, coalesced sizing+mount, rootfs
grow separate, started, guest NOT destroyed); compensating rollback at each step (restore /
config / start-task / waitRunning — asserts the guest WAS destroyed); DR continuity (MAC kept,
hostname not reset) + DR
KeepMAC=falseresets MAC; liveness verdict (warnings+running pass / not-running fail); F4 (lock-500→retry→proceed; non-lock-500→fail without retry); owning entry journaled BEFORE restore; reserved/existing VMID refused;Recoverrolls back / clean.
Live-validated (demo-felhom)
- provision: fresh MAC + hostname; SSH host keys regenerated by the baked golden unit (agent
issued no
ssh-keygen), machine-id unique, Docker runs, clean DHCP lease → torn down. - dr: continuity preserved (hostname + host keys kept). Recover: a killed mid-restore left an
orphan; the re-run's
Recoverrolled it back (idempotent). - Live caught a bug, then fixed: the host-key unit's
ExecStartwas/usr/sbin/ssh-keygen(203/EXEC); on Debian 13 it is/usr/bin/ssh-keygen— corrected inbuild-golden.sh, golden rebuilt, re-validated. (Mocked unit tests couldn't surface this; the live run did.)
v0.7.0 — restore-test: verdict is liveness, not start-task exitstatus (2026-06-09)
Fixes a correctness bug found by the live hub-enrollment runbook: the self-restore-test reported
pass:false on every modern-distro guest. PVE's guest-start task exits "WARNINGS: 1" for the
benign systemd-nesting advisory (WARN: Systemd 257 detected. You may need to enable nesting.), and
WaitTask treated any non-"OK" exitstatus as a hard failure — so the verdict was decided by an
advisory exit code instead of by observed liveness, before the real boot check ran. A crying-wolf
test got it disabled on the demo host; this re-enables it. Single bump (0.6.0→0.7.0) covering the
agent's part of both task phases; the wire fields below are consumed by hub from v0.7.5.
Design invariant (in code): warning classification affects visibility only; pass/fail is liveness-only. A wrong/stale recognizer can at worst over-notice a benign warning — it can never false-fail and never hide a real warning.
Added
proxmox.WaitOptions.AllowWarnings— opt-in per call. When set, a task that completes"WARNINGS: N"is success with theTaskStatus(ExitStatus intact) returned so the caller can read/surface it. Default (false) keeps every existing caller strict (vzdump/restore/destroy warnings can be meaningful — relaxing them is a future per-call decision with evidence). Any non-WARNINGS non-OK exit is still a*TaskError.reconcile.RestoreTestResult.StartWarnings/.WarningsRecognized+ a version-free recognizer (benignWarningAnchor = "enable nesting", case-insensitive substring — contains no systemd version number, so it can't rot back into the bug at systemd 258+).extractWarningLinespullsWARN…lines from the start-task log.reconcile.GuestAPI.TaskLogTail— the engine fetches the start task's log to surface warnings.hub.RestoreTest.warnings/.warnings_recognizedwire fields (omitempty), populated byToHubRestoreTest. Additive: the deployed v0.7.4 hub ignores them; hub v0.7.5 consumes them (passed-with-warnings INFO, or WARN when not recognized). Cross-repo golden updated with the hub side.
Changed
- Restore-test start step (
reconcile/restoretest.go) now waits withAllowWarnings:true, surfaces any start warnings, and continues towaitRunningas the verdict — boot+running is the pass, exactly as before; a real (non-WARNINGS) start-task error still fails. The restore and scratch-teardown WaitTasks stay strict. - Restore-test scheduler logging distinguishes a clean pass, passed-with-recognized-warnings (INFO), and passed-with-unrecognized-warnings (WARN) — nothing silent.
Tests
WaitTask: AllowWarnings acceptsWARNINGS(status returned intact); AllowWarnings still fails a real error; default still fails onWARNINGS(existing callers unaffected).- Restore-test (engine, mock proxmox): start-with-warnings + running → pass with warnings surfaced+recognized; unrecognized warning + running → pass, not-recognized; not-running → fail regardless of warnings (verdict is liveness); teardown still runs.
- Regression guard: the
"enable nesting"recognizer matches the advisory for systemd 256–300, proving it's version-independent and can't silently rot back into the false-fail.
v0.6.0 — slice 6 Phase B: PBS offsite tier (verify + PBS-API client + reporting) (2026-06-09)
Completes slice 6. The PBS spike (felhom.eu phase5-pbs-spike-findings.md) proved backup-to-PBS and restore-from-PBS reuse Phase A UNCHANGED (PBS is just a storage target + a volid), and the operator token needs no widening. So the only new agent code is the verify capability + a small PBS-API client + PBSSnapshot reporting. Escrow + host-loss DR stay slices 7/10.
Added
internal/pbs— the PBS-API client (the agent's SECOND privileged external surface, slice-1 discipline): TLS fingerprint-pinned to the PBS leaf cert (a spoofed PBS → rejected, mirroring the PVE pin), token auth (PBSAPIToken=<id>:<secret>; id from the storageusername, secret read at runtime from/etc/pve/priv/storage/<id>.pw— referenced by location, never logged/committed), typed, no shell. Methods:Verify(POST/admin/datastore/<ds>/verify→ UPID),Snapshots(incl. theverificationfield),TaskStatus/WaitVerify(node extracted from the UPID —localhostreturns "unknown", the spike B4 gotcha),NodeFromUPID.- The verify maintenance loop (
pbs/verify.go) — the cheap, key-free, ciphertext-level integrity check (§8) on its OWN cadence (default 6h, the 5th daemon goroutine). It is a reporting/maintenance task like the slice-5 watchdog: it does NOT go through the reconcile gate/journal. Each cycle: trigger verify → poll task → re-list snapshots → record per-snapshotverify_state. A failed verify is logged loudly. PBSSnapshotreporting — filled the stub (namespace/backup_type/backup_id/backup_time(RFC3339)/size_bytes/owner/protected/encrypted(fromfiles[].crypt-mode) /verify_state(ok|failed|none until verified)/verify_upid). NewPBSReportercollector seam + an in-memorySnapshotStore. Cross-repo golden (both repos, byte-identical)- bidirectional key-set tests; hub
handler.goparsespbs_snapshotsand logs a failed verify[WARN](loudest offsite-DR signal).
- bidirectional key-set tests; hub
- Truthful backup mode (
backup/runner.go) —Backup.modenow reflects the ACTUAL vzdump mode read from the task log (backup mode: <x>), since PVE may downgrade snapshot→stop for a stopped guest (spike B1); falls back to the requested mode if unparseable. - proxmox:
Storage.Username(parsed from the pbs storage config — the token id). - config
BackupConfig.{PBSVerifyCadenceSeconds, PBSSecretDir}(cadence 0→6h, <0 disabled). --selftest=pbs-verify— discover pbs storages → verify each → print the PBSSnapshot records (covers the runbook's verify + list). Standalone on the host.
Notes
- Backup/restore-to-PBS reuse Phase A with no change (the restore-test runs with
source_tier="pbs"when fed a pbs volid). Zero-knowledge holds: verify is ciphertext-level, the encryption key is never read here, and the PBS server has no client key (spike B6). - Daemon runs cleanly with no pbs storage / verify disabled.
go test -racecovers the new goroutine. Slice-3/4/5/6A surfaces, goldens, and adversarial tests intact.
v0.6.0-rc1 — slice 6 Phase A: backup + the self-restore-test (local target) (2026-06-09)
Phase A of the backup/restore slice (doc 03 §8) — the agent's guest-level backup layer and the self-restore-test, which closes "a backup you haven't restored isn't a backup". Everything here is BENIGN (backup, restore-to-NEW, scratch teardown): reuses the slice-4 classifier/gate/journal — no new destructive class, no new crypto. Local target only; PBS is Phase B. Restore is to a NEW guest only (no overwrite). Backups are crash-consistent only (app-consistency needs the controller quiesce, slice 8) — marked so in the report.
Added
- proxmox (
mutate.go/query.go):DestroyLXC(DELETE …/lxc/{vmid}?purge=1&destroy- unreferenced-disks=1 → UPID; the scratch-teardown primitive);VzdumpOptions.Notes→notes-template(verified on PVE 9.2.2);LatestBackupVolID(resolve a produced archive from the backup-storage listing — the task status carries no result volid). - reconcile self-restore-test (
restoretest.go) —Engine.RunRestoreTest: pick a free scratch VMID (configured band, excludes 9999; full band → skip, never out-of-band) → journal a Scratch-owned entry BEFORE any mutation → restore-to-new → benign net link-down SetConfig (so the clone can't conflict with a running source's MAC/IP; this is test-safety, NOT slice-7 identity reset) → boot → verify reachesrunning→ ALWAYS teardown (defer; benignClassGuestDestroy+ agent-tagged-scratch provenance, gated). Runs on the scratch VMID's queue lane. Reuses the journal/gate; result feeds the report. - Crash-safe recovery (
recover.go): a Scratch journal entry is resolved by TEARDOWN, not by re-checking the restore sub-task's UPID — special-cased BEFORE the generic path (else the restore task's OK would mark it succeeded while the guest leaks).Recovernow destroys a leaked scratch guest (idempotent: already-gone → clean; list-unreadable → left in-flight for a later pass).JournalEntry.Scratchflag;RecoverResult.ScratchClean/ ScratchDestroyed. GuestAPI gainsRestoreLXC/DestroyLXC/GuestStatus. internal/backuppackage:BackupRunner.Backup(vzdump + archive/size resolve + bulk-volume gap — a mountpoint is UNCOVERED unless it carries an explicitbackup=1, so an unsetbackup=is reported uncovered too, the safe DR direction);PickRestoreCandidate(newest backup); an in-memoryStore(latest-backup-per-target + latest-restore-test) implementing the hubBackupReporter/RestoreTestReporterseams; a cadenceScheduler(default 24h; the fourth daemon goroutine; disabled cleanly when off/misconfigured).- hub report (
report.go): filled theBackup+RestoreTeststubs (PBSSnapshotstays a Phase-B stub); collectorBackupReporter/RestoreTestReporterseams. Cross-repo golden updated in BOTH repos (byte-identical) + bidirectional key-set tests forbackups[0]/restore_tests[0]. Hubhandler.goparses + persists them (report_json; no new columns) and logs a FAILED restore-test prominently (the loudest DR signal). - config
BackupConfig(local target, restore storage, restore-test cadence, scratch VMID band 990000–990009 default) + accessors + env overlay + cadence-gated validation. --selftest=backup -vmid N(one-shot backup → print the Backup record) and--selftest=restore-test [-archive volid](Recover-then restore→boot→verify→teardown, print the RestoreTest record). Standalone on the Proxmox host.
Notes
- The daemon runs cleanly with the cadence off or misconfigured (logs + disables, never
crashes); a leaked scratch guest from a mid-test crash is reaped by
engine.Recoveron restart.go test -racecovers the new scheduler goroutine. - Slice-3/4/5 exported surfaces, goldens, and adversarial tests intact. Version bumps to v0.6.0 when Phase B (PBS) lands.
v0.5.1 — slice 5 live-validation prep: durable_id mis-id fix + re-mount UUID memory (2026-06-09)
Two correctness fixes surfaced while preparing the live USB validation on demo-felhom
(a real 1TB USB HDD, sdb1, ext4). Both are DR-load-bearing — exactly the "false-id →
re-attach the wrong disk" failure mode the slice warned about.
Fixed
- Unmounted dir-storage no longer inherits the ROOT filesystem's UUID (
observe.go). Previously, when a removable dir-storage was unmounted, the observer fell through to the containing mount (root) for the backing device, so itsdurable_idbecameuuid:<root-uuid>— a catastrophic DR mis-id (the hub would re-attach the wrong disk). Now the backing device/UUID/durable_idare derived ONLY from the target's OWN mountpoint; an unmounted target reports no device and a stablestore:<name>durable_id, never another filesystem's UUID. (Removed thecontainingMountDeviceroot-fallthrough.) - Watchdog remembers the fs-UUID observed while attached (
watchdog.go) so a re-mount works even after the known-set cache refreshes mid-drop (an unmounted target can't resolve its own UUID). The re-mount key is backfilled from this memory — aligning with doc 03 §7's "sourced from the existing definition, no hub manifest needed": the agent learns the UUID while the target is attached, then re-mounts by it on return.
Tests
- Observer: an unmounted dir-storage asserts NO
uuid:durable_id and no backing device. - Watchdog: a drop where the cache lost the UUID still re-mounts using the remembered UUID.
v0.5.0 — slice 5 Phase B: the host-root surface (mounts + SMART + grow + destructive gate) (2026-06-09)
The write surface — the agent's first step outside its Proxmox API token into OS-root. Isolated behind a narrow, argument-validated, adversarially-tested seam, exactly like the slice-4 gate. Completes slice 5 (Phase A = read-only observe/report/watchdog at v0.5.0-rc1).
Added
HostOpsseam +SudoHostOps(internal/storage/hostops.go) — the one privileged host surface: persistent mounts via systemd.mountunits keyed by fs-UUID (enabled to survive reboot), detach (stop+disable), SMART, and thin-pool metadata. Shells out via the fenced Runner (sudo -n, fixed arg vectors, no shell); a fake backs the tests (no real root in the suite).NoopHostOpsis the safe fallback when the surface is unavailable.- The argument validator (
internal/storage/validate.go) — the security boundary:ValidateUUID(strict hex),ValidateMountPath(absolute, no traversal, no metacharacters),ValidateSMARTDevice(raw-disk whitelist),ValidateLVMName, and an in-processsystemdEscapePath(nosystemd-escapeshell-out). Every argument is validated BEFORE a command is constructed. Headline test (validate_test.go): an adversarial matrix of shell metacharacters /../traversal / malformed inputs is rejected with zero exec. - SMART (
internal/storage/smart.go) — parsessmartctl -a -jintoStorageTarget.smart: SATA (reallocated/pending/offline-uncorrectable, temp, power-on-hours) and NVMe (critical_warning, media_errors, percentage_used, temp), degrading toUNKNOWNfor devices with no SMART (USB-SATA bridges).lvsfills the lvmthin thin-pool metadata fill (the value Phase A left null). Wired into the Observer's enrichment (Observe only, not the watchdog's fast Known path). - Watchdog re-mount response (
internal/storage/watchdog.go) — on a known mount-backed target's device returning unmounted (a newDevicePresentliveness probe), the watchdog dispatches a benign by-UUID re-mount off the poll path (a goroutine, never under the lock), rate-limited per target to the debounce window. The mount is routed through the gate as benign (gateRemounterinmain.go, sostoragestays decoupled fromreconcile). - Disk-grow executor (
internal/reconcile) —ActionResize(benignClassResize), planned grow-only (desired DiskBytes > actual →pct resize rootfs +<n>M; a shrink is refused, never silently grown) + a defensive executor guard (size must start with+). Newproxmox.Client.ResizeLXC(API;VM.Config.Disk+Datastore.AllocateSpace; async→UPID). Built + fixture-tested; unfed live (no hub spec until slice 10). - Destructive storage ops through the slice-4 gate (
internal/reconcile/storage_ops.go) —IntentForStorageMount(benign) andIntentForStorageDestructive(ClassStorageWipe/ClassDecommission). Host/target-scoped: the op binds on the storage target identity (carried intarget.guest_id). Reuses the existing verifier/role-scoping/binding/audit — no new gate, no new crypto. Storage cases added to the adversarial matrix (storage_test.go): unsigned wipe →pending_signature; "wipe A" signature vs "wipe B" →binding_mismatch; valid → accepted. Inert live. --selftest=storage[-watch <dur>] — the live USB-runbook harness: an observe pass (full table incl. SMART + thin-pool data+metadata), and a bounded watchdog window with the re-mount response live. Runs standalone on the Proxmox host (no hub).configs/felhom-agent.sudoers— the documented narrow allowlist (install unit / systemctl manage / smartctl / lvs), with the agent-side fine validation noted.- Config:
privileged.{unit_dir,stage_dir,systemctl,install,smartctl,lvs}(paths must match the sudoers entries).
Notes
- Daemon still runs cleanly with no removable storage / no signers / no hub manifest, and a
missing/declined sudoers entry degrades with a warning (SMART→UNKNOWN, mount→logged error),
not a crash.
go test -racepasses (the watchdog re-mount dispatches off the poll path). - Slice-3/4 + Phase-A exported surfaces, goldens, and adversarial tests intact.
authzuntouched. The destructive-storage executor + grow are built/tested but unfed live until slice 10.
v0.5.0-rc1 — slice 5 Phase A: storage observe + report + watchdog (read-only, live) (2026-06-09)
Phase A of the storage slice (doc 03 §7). Read-only and live: the agent now observes every
host storage target, reports it into the host-report's storage_targets (previously an empty
stub), and runs a fast-poll watchdog that pushes a disconnect to the hub in seconds. No
host-root writes this phase (mounts/SMART/grow/destructive-gate are Phase B). The hub-owned
desired manifest (class/role/policy/creds) is not served until slice 10, so reconcile against
it is built-but-unfed — this phase ships only the genuinely-useful read-only footprint.
Added
internal/storagepackage (new):StorageTargetwire contract (internal/hub/report.go) — filled the slice-3 stub:name/type/durable_id/state/reachable, usage (total/used/avail/used_fraction),content,mount_path/backing_device,class_hint(rotational HINT — never authoritative; class is hub-owned),role(empty until slice 10), athin_poolsub-object (lvmthin data fill; metadata fill is Phase B/lvs), and asmartsub-object (UNKNOWNuntil Phase B). Cross-repo golden kept byte-identical withfelhom.eu/huband guarded by the bidirectional key-set test (contract_test.go).durable_idderivation (durableid.go) — deterministic per type (the DR-load-bearing re-attach key): fs-UUID (usb/local-dir),server:export(nfs/cifs),repo+fingerprint(pbs),vg/pool(lvmthin); never empty (falls back to a stable store id).HostReaderseam +ProcHostReader(hostread.go) — non-privileged/proc/mounts,/dev/disk/by-uuid,/sys/.../rotational+removablereads. Root-free by construction.Observer(observe.go) — builds[]hub.StorageTargetfromListStorage/NodeStoragejoined with host reads; surfaces the lvmthin thin-pool data fill prominently (warns ≥85%).- Storage watchdog (
watchdog.go) — a third daemon goroutine fast-polling the known target set (a defined Proxmox storage and/or a previously-seen one) forattached↔disconnectedtransitions; on a transition it triggers an immediate, debounced out-of-band host-report. Only flags a known target's change (never a never-attached device); coalesces flaps within the debounce window (leading + trailing edge).CachingKnownTargetsrate-limits the Proxmox-derived known set;HostLivenessprobes device/mount presence (local) + a reachability dial (network), all non-privileged.
- Proxmox
Storagetype (internal/proxmox/types.go) — additive parse-only config fields (server/export/share/datastore/fingerprint/vgname/thinpool) feeding durable_id. - Collector
StorageObserverseam (internal/hub/collect.go) — populatesstorage_targetsvia the observer; a nil observer or an observe error degrades to empty (never sinks the heartbeat). Hub does not import storage (storage imports hub for the wire type). - Out-of-band report trigger (
internal/hub/loop.go) —Loop.SetTrigger: a watchdog signal runs one extra collect→report immediately without disturbing the regular cadence. StorageConfig(internal/config) — watchdog interval / debounce / known-refresh knobs (all optional; package defaults otherwise).- Hub ingest (
felhom.eu/hub) —hostReportPayloadnow parsesstorage_targets(full mirror struct), persists them viareport_json, counts + warns on disconnected targets, and has its own half of the bidirectional golden key-set test.
Notes
- The daemon still runs cleanly with no removable storage, no signers, and no hub manifest — the watchdog finds nothing to flag; storage reporting is best-effort.
proxmox/hub/authz/reconcileexported surfaces + their golden/adversarial tests are intact. No host-root writes, no destructive paths, no SMART this phase (all Phase B).- Version: v0.5.0-rc1 at the Phase-A checkpoint; v0.5.0 when Phase B lands.
v0.4.0 — slice 4 Phase B: reversibility gate + signed-op consuming layer (2026-06-08)
The security core of slice 4: hub-supplied intent stops being trusted for destructive
change. Layered in front of the per-guest queue's executor — every mutation now
passes the gate. Reuses internal/authz for all crypto (untouched surface). Inert
this slice: no destructive deltas are served until slice 10, so the destructive path is
classified, gated, and adversarially tested but not wired to live execution.
Added
- Classifier (
classify.go, doc 03 §4) — benign vs destructive by provenance + data-bearing-ness, NOT by verb. TheOpClassvocabulary (seeded by the committed slice-2op_blob.json:guest_destroy) is the agent-side contract slice 10 matches. Destroy/overwrite of customer data is destructive UNLESS agent-internal provenance (same-journaled-transaction create → compensating rollback, or agent-tagged scratch) makes it benign.Provenanceis journal-recorded and never populated from the hub (its zero value is the only thing an external intent may carry). Unknown op class fails safe → destructive. - Reversibility gate (
gate.go) —Gate.Authorize(intent, signed): benign → allowed unsigned; destructive → requires a verified, role-authorized, action-bound operator signature, else refusedpending_signature, never executed. Every decision is written to anAuditSink(audit is a signal, never the guard). - Signed-op consuming layer over
authz— verifies viaauthz.Verifier.Verify(the locked pipeline, untouched), then enforces on theVerifiedOp:- Role-scoping (doc 04 §4) — recovery key authorizes key-rotation re-pins ONLY; operational key authorizes ordinary destructive ops + planned rotation.
- Op-to-action binding — verified
op+ host + guest +paramsmust match the gated action (a signature for guest X / op A can't authorize guest Y / op B); params compared semantically (key-order/whitespace independent).
- Signed-job orchestration (
job.go) —RunSignedJob: idempotency dedupe (the op nonce as the journal key — a redelivered completed op is skipped, not re-run), gate authorization, then journal-wrapped execution via an injectedDestructiveExecutor(nil this slice — authorized destructive ops are inert, no executor wired until 6/7). - Crash-recovery consumer (
recover.go, Note 1 / doc 03 §10) —Engine.Recoverconsumes the journal'sInFlight()at startup: an op that crashed AFTER the Proxmox POST and BEFORE its terminal record (OpTaskRunning, nonce already consumed) is NOT covered by idempotency dedupe — only this resume-or-rollback resolves it (re-read the task via the newTaskStatusOnce, record the real outcome; a no-task-id op is abandoned fail-safe). Landed together with the signed-op executor, as Note 1 required. - Daemon wiring —
runDaemonbuilds the verifier fromconfig.Authz.Signers(a bad key / missing nonce-store path is a fatal misconfig; no signers = nil verifier, the common slice-4 state), constructs the gate (+SlogAudit), runsRecoverbefore issuing any mutation, and routes every reconcile action through the gate.
Changed
- Memory comparison canonicalized (Note 2) —
desiredMemoryMiBmakes the desired↔actual memory compare in the same MiB unit that is then written, so a non-MiB-alignedMemoryBytesconverges in one pass instead of re-issuing SetConfig forever (the numeric cousin of the description-newline normalization). Test proves convergence. Slice 10 should still serve MiB-aligned specs at the source.
Tests (the security proof — each independently rejected)
- Adversarial matrix via the REAL
authz.Verifierwith in-test-minted SSHSIGs (framing replicated in reconcile's test binary; production authz untouched, no signing added to the verify-only package): unsigned destructive job → pending_signature; unsigned destructive desired-state delta → pending_signature (distrusts hub desired state, not just jobs); forged/unknown signer →ErrUnknownSigner; expired →ErrExpired; replayed nonce across an agent restart (durableFileNonceStore) →ErrReplay; wrong host →ErrTarget; wrong guest / wrong op / wrong params → binding_mismatch; recovery key on ordinary destructive → role_denied; hub-supplied "scratch" tag ignored → still destructive → refused; valid + role + target + fresh nonce → accepted, and a second presentation →ErrReplay(nonce consumed). - Classifier (benign/destructive/provenance/key-rotation/fail-safe), role-scoping, params binding, crash-recovery (resume OK / fail / still-running / no-task rollback / unreadable / one-shot key applied on resume), signed-job idempotency (execute once, dedupe redelivery, refused-not-executed, no-executor-inert, executor-error).
- Full module race-clean (
go test -race) + vet clean on the Linux build server.
v0.4.0-rc1 — slice 4 Phase A: reconcile engine (structural; runs live, unfed) (2026-06-08)
The agent-side control core's structural half. Checkpoint marker — -rc1 is the
Phase-A push; awaiting validation before Phase B (the reversibility gate + signed-op
consuming layer) lands the final v0.4.0. Runs LIVE but UNFED: with no desired-state
provider until slice 10, the live engine computes an empty action set and performs
zero mutations.
Added
internal/reconcilepackage — the engine, the per-guest serializer, the desired-state model, the normalization layer, and the durable op journal:- Per-guest serializer (
Queue, doc 03 §10) — the single choke point ALL mutation sources funnel through. Same-vmid jobs run strictly one-at-a-time in submit order; independent vmids run in parallel. Each vmid is a cond-var FIFO lane (unbounded, non-blocking, order-preserving); graceful drain onClose. - Desired-state model +
DesiredProviderseam —DesiredGuest(per-field optional: run-state /*hub.GuestSpec/*description),DesiredState. The only live provider isEmptyProvider(slice 4 has no source);StaticProviderfeeds fixtures. The seam is where slice 10's hub-serving plugs in — no hub/local source invented here. - Normalization layer (
FieldNormalizers) — reconcile compares normalized desired-vs-actual so Proxmox round-trip quirks don't read as drift.description's trailing newline is the first registered case; the registry takes more (boolean coercion, list ordering) as discovered.normDescpromoted out ofcmd/felhom-agent/main.gotoreconcile.NormDescription; the--selftest=taskdescription round-trip now uses that shared helper (one source of truth for the quirk). - Plan engine (
Plan, pure function) — computes the minimal benign action set (Start/Stop/SetConfig) for guests present in both desired and actual, with normalized comparison, deterministic vmid ordering, config-before-run-state. Skips provision (desired-absent-in-actual, slice 7) and destroy (actual-absent-in-desired, gated, slice 10); never writes a config it couldn't first read (SpecKnown). Disk (rootfs grow) intentionally not reconciled here. - Reconcile engine (
Engine) — reads desired+actual, plans, dispatches each action onto the shared queue. Every Proxmox op handled per the mutate.go contract: non-empty UPID →WaitTask+ assertexitstatus; empty UPID → clean synchronous success (slice-4 proven). Per-action failures are counted, not fatal (other guests still converge). - Operation journal (
Journal) — durable fsync'd append-only JSONL mirroringauthz.FileNonceStore: records each op's lifecycle (started → task_running → succeeded/failed) with its Proxmox task id (crash mid-op is detected and re-checkable on restart viaInFlight()), plus an idempotency-key store (AlreadyApplied) so a one-shot op never re-runs across retries/restarts. Reconcile actions carry no idempotency key (convergent — must re-run on real drift).
- Per-guest serializer (
- Daemon wiring (
runDaemon) — reconcile runs alongside the hub loop on the poll cadence, sharing the per-guest queue. Journal path is ajournal.logsibling of the nonce store. The daemon runs cleanly with no desired state and no signers (reconcile is a logged live no-op; a journal-open failure degrades to journal-less, never crashes).
Tests
- Serializer: same-guest serialized (max-concurrency 1, submit order preserved) and different-guests parallel (cross-waiting jobs both complete — would deadlock if not); error propagation; drain-pending-on-close; submit-after-close.
- Normalization: description round-trip; unknown-field identity; extensibility seam (synthetic boolean-coercion + list-ordering normalizers).
- Plan: run-state start/stop, spec drift (cores/memory), disk-not-reconciled, description-newline-not-drift, unmanaged fields, spec-unknown skips config keeps run-state, desired-absent skipped, combined ordering, empty-desired no-op, deterministic vmid order.
- Engine: empty-provider zero mutations; async start (WaitTask); synchronous SetConfig (no WaitTask); WaitTask failure + POST error counted failed; list error = pass failure.
- Journal: lifecycle latest-wins; in-flight survives restart; idempotency dedupe across restart; failed key not applied; torn-trailing-line skipped.
- Full module race-clean (
go test -race) on the Linux build server; vet clean.
Not in this phase (Phase B)
- The benign/destructive classifier, the reversibility gate, and the signed-op consuming
layer over
internal/authz(doc 03 §4 / doc 04) — added next, in front of the queue's executor, landing v0.4.0.
v0.3.2 — SetConfig selftest extension (slice-4 pre-check) (2026-06-08)
The gate before slice 4: prove SetConfig works live under the scoped token before
reconcile is built on it. Self-gated live run PASSED on demo-felhom/guest 9999.
Added
- Reversible
SetConfigstep appended to--selftest=task(cmd/felhom-agent/main.go,selftestSetConfig): readGuestConfig→ write adescriptionmarker (felhom-selftest <RFC3339>) → verify it landed → restore the original value (ordeletethe key if it was absent) → verify the restore. Handles PVE's dual-modeSetConfigreturn per themutate.gocontract: empty UPID = synchronous success (printedsynchronous); non-empty UPID =WaitTask+ assertexitstatus=OK. The existing snapshot → rollback → delete-snapshot steps are unchanged. First live exercise of theVM.Config.*privilege cluster. normDesc/extraStringhelpers —extraStringdecodes a string-valued key fromGuestConfig.Extra(raw JSON);normDescstrips the trailing newline PVE appends todescriptionon read, so a written value round-trips equal.
Finding (live)
- The LXC
descriptionwrite returned synchronous (empty UPID) — PVE applied it inline, no task. The agent's dual-modeSetConfigmodeling is correct: the empty-string path is real and must not be treated as an error. - PVE appends a trailing
\ntodescriptionon read (stored URL-encoded as%0A). A naive exact-match reconcile would see perpetual drift — slice-4 reconcile must normalizedescriptioncomparisons (hencenormDesc).
Ops
- Standing operator token (
felhom-agent@pve!agent, privsep) rotated during this run (the prior secret was not retrievable); role + both user/token ACL rows re-confirmed at/. New secret stored out-of-band, not persisted to the repo. Guest 9999 left pristine (stopped, nodescription, no leftover snapshot). Version → 0.3.2.
Docs + live validation — no version bump (2026-06-08)
Changed
- Reflowed
CLAUDE.md— removed hard mid-paragraph line wraps (prose, list items, blockquotes now single-line, soft-wrapped); code blocks and tables untouched; rendered output unchanged. - Unified the REPORT/CHANGELOG convention in
CLAUDE.md:CHANGELOG.mdis the cumulative log (newest on top);REPORT.mdis overwritten with the most-recent implementation/validation only. Added an explicit no-secrets rule (never write tokens/passwords/keys into committed files; reference them as stored out-of-band).
Added
REPORT.mdrewritten for the live--selftest=taskvalidation on the demo host (demo-felhom): snapshot → rollback → delete-snapshot on guest 9999, each polled toexitstatus=OKunder thefelhom-agent@pve!agentprivsep token (UPIDs name the token actor — privsep path genuinely exercised); 16-privilegeFelhomAgentrole + both user & token ACLs confirmed;--selftest=readclean. Closes the slice-1 "mutating ops unit-tested only" gap;WaitTaskasync foundation validated live → slice 4 unblocked. (Token secret stored out-of-band, not in the repo.)
v0.3.1 — slice-3 validation follow-ups (2026-06-08)
Changed
- Collector keeps the known run-status on a
GuestConfigfailure (internal/hub/collect.go): previously a per-guest config-read error forcedstatus="unknown"; now the run-status fromListLXCis preserved (only thespecis dropped). An empty status is still normalized tounknown(wire value is alwaysrunning|stopped|unknown). Test renamed toTestCollect_GuestConfigFailureKeepsStatusOmitsSpecand asserts the preservedrunning+ nil spec. --selftestusage error string now reads(want read|task|hub).
Added
- Cross-repo contract fixture
internal/hub/testdata/host-report.golden.json+TestHostReport_ContractMatchesGolden— compares the marshaledHostReportfield-name sets (top level +host+guests[0]) against the golden, failing on any json-tag drift. The file is kept byte-identical with felhom-hub's copy (duplicated contract until a shared types module; revisit when slices 5/6 populate the empty collections). Version → 0.3.1.
v0.3.0 — hub client + host-report + first daemon loop (slice 3) (2026-06-08)
The agent's first daemon: a periodic read-only host-report POSTed to the hub (the heartbeat). No Proxmox mutations, no desired-state/signed-op consumption, no storage/backup collection yet — those are slices 4/5/6.
Added
internal/hubpackage:HostReportwire contract (report.go) shared field-for-field with the hub ingest: host metrics, guests (vmid+ spec),cloudflaredstatus, and thestorage_targets/backups/restore_tests/pbs_snapshots/audit_tailcollections defined but emitted empty (typed[], slices 5/6 fill them).Collector(collect.go) builds the report from a read-onlyproxmoxReader(adapted to the realinternal/proxmoxsurface — node held by the client, value returns,proxmox.Guest) + aCloudflaredProber. Partial-failure policy: a failedNodeStatusis a hard error (skip the POST); a failed per-guestGuestConfigdegrades that guest tostatus="unknown"(spec omitted) but still sends; a cloudflared probe failure →"unknown", never fatal.CloudflaredProber+SystemctlProber(systemctl is-active cloudflared; read-only — NOT a Privileged/root op; tunnel management is a later slice).Client(client.go):POST /api/v1/host-reportwithAuthorization: Bearer <key>, standard TLS (system roots or optionalca_file; verification always on). Typed*TransportError/*HTTPError; the bearer token never appears in any error.Loop(loop.go): the daemon — immediate first report then tick; adopts the hub'spoll_interval_secondsclamped to [60,3600]; resilient (a collect/report error is logged and the loop continues); clean shutdown on context cancel.ControlEnvelope: onlypoll_interval_secondsis acted on;blocked/desired_generation/has_signed_opsare parsed-but-ignored (logged at most) pending reconcile (slice 4).
- Config:
HubConfig(url/host_id/api_key/poll_seconds/timeout_seconds/ca_file),FELHOM_AGENT_HUB_*env overlay,HubConfig.Validate()(mode-aware — proxmox-only--selftest=read|taskstill runs without hub config),WithDefaults(), andRedacted()now also blanks the hub key.configs/agent.example.jsongainshub(andauthz) blocks. cmd/felhom-agent: the no---selftestmode is now the daemon (poll loop); added--selftest=hub(one collect+report, prints the report + envelope). Version 0.2.0 → 0.3.0.
Tests
- Report serialization (field names; empty collections are
[]notnull; spec omitted when unknown); client (Bearer header, non-2xx→*HTTPError, transport→*TransportError, token never in error); collector (host mapping, guest spec, per-guest failure degrades-but-still-reports, NodeStatus hard error, cloudflared error→unknown); loop (immediate first report, continuation after an injected error, interval adoption + clamp); config (hub validate/redact/env).
Notes
internal/proxmoxandinternal/authzwere not touched — no new proxmox surface was needed (ListLXCalready exposes status/maxmem/maxdisk;GuestConfigexposes cores). The task'sproxmoxReadersketch (node-arg/pointer/LXC) was adapted to the real exports as instructed.- Defined-but-empty this slice:
storage_targets,backups,restore_tests,pbs_snapshots,audit_tail(slices 5/6). Parsed-but-ignored: the envelope'sblocked/desired_generation/has_signed_ops(slice 4).
v0.2.0 — authz signed-op verifier (slice 2) (2026-06-08)
Production form of the Phase-4 signing primitive: a key-type-agnostic SSHSIG verifier for operator-signed destructive ops, with the full anti-replay/ authorization pipeline and a durable, crash-safe nonce store. What slice 4 (reconcile) will call to gate destructive desired-state deltas. No hub, no signing CLI, no reconcile loop.
Added
internal/authz—Verifier:New(signers, store, hostID)+Verify(blob, sigArmored) (*VerifiedOp, error). Runs the LOCKED pipeline (order is load-bearing): parse armor → namespace → parse pubkey → allow-list (by key material,pub.Marshal()equality, not key_id) → crypto verify (over the raw received bytes, never re-canonicalized) → parse blob → target → time window → nonce recorded LAST. Each post-crypto stage rejects even with a valid signature.- SSHSIG framing (
sshsig.go) viagolang.org/x/crypto/ssh—pem.Decode→ strip 6-byte magic →ssh.Unmarshal→ssh.ParsePublicKey→ recompute signed data with the named hash →pub.Verify(dispatches on key algorithm). No hand-rolled crypto. Key-type-agnostic: ed25519 / sk-ssh-ed25519 (FIDO2) / rsa / ecdsa via the one path. - Fixed namespace
felhom-op-v1(package constant, never caller-supplied). OpBlob(correctedhost_id/guest_idjson tags) +VerifiedOp(op, host/guest, params, key_id, matched signer). key_id is advisory/audit only — never an authz input.- Typed errors:
ErrMalformed, ErrNamespace, ErrUnknownSigner, ErrBadSignature, ErrTarget, ErrExpired, ErrNotYetValid, ErrReplay(errors.Is-friendly). NonceStore+ two impls:MemoryNonceStore(tests) andFileNonceStore— durable, crash-safe (fsync'd append log, replayed into an index on open, periodic compaction, expiry-only pruning). A nonce is fsync'd to disk beforeSeenOrRecordreturns false; replay protection survives restart; I/O failure fails safe (reports seen=true). Target generalization: host_id matched strictly, guest_id surfaced for the caller to route.- Config:
AuthzConfig(nonce-store path + pinned operatorsignerstaggedoperational/recoverywith a key_id, as authorized_keys lines). - Version 0.2.0.
Tests
- Real OpenSSH interop via a committed
ssh-keygen -Y signvector (hermetic CI); per-stage rejection (each with an otherwise-valid sig); the headline invalid-sig-does-not-burn-the-nonce invariant; replay; persistence across restart; synthetic sk-ssh-ed25519 through the unchanged path; byte-exactness (a re-serialized blob fails crypto — not re-canonicalized).
Notes / corrections to the Phase-4 reference
- §7's
Targetlacked json tags (host_id/guest_id) — fixed. - The doc paired "Go 1.24.4 / x/crypto v0.52.0", but v0.52.0 declares
go 1.25.0and does not build on Go 1.24. Resolved by upgrading the build server to go1.26.0 (backward-compatible; felhom-controller/hub unaffected); the module isgo 1.25.0on x/crypto v0.52.0. - Free function → constructed
Verifier; returns the fullVerifiedOp; typed errors; clock-skew tolerance added; durable nonce store is the net-new work. - Shared-contract dependency flagged (not built): the hub and the
felhom-signCLI must emit byte-identical canonical JSON or signatures won't verify; a shared canonicalizer both import would be the right home.
v0.1.0 — Scaffold + proxmox interaction layer (slice 1) (2026-06-08)
First slice: stand up the host-agent project and its foundation — the typed Proxmox interaction layer every other module will call. No reconcile loop, hub client, signing, or storage/backup orchestration yet (later slices).
Added
- Project scaffold: module
gitea.dooplex.hu/admin/felhom-agent, binaryfelhom-agent(cmd/felhom-agent/), Go 1.24, zero external dependencies (pure stdlib).--versionflag;versionvar overridable via-ldflags "-X main.version=<v>". internal/proxmox— API backend (Client): hand-rolled REST client overhttps://<host>:8006/api2/jsonwithPVEAPITokenauth. Typed read ops (Version,Nodes,NodeStatus,ListLXC,GuestStatus,GuestConfig,ListStorage,NodeStorage,StorageContent) and async mutating ops returning a UPID (RestoreLXC— the primary create path,Vzdump,Snapshot,Rollback,DeleteSnapshot,SetConfig,Start,Stop).WaitTask: pollsGET /nodes/{node}/tasks/{upid}/statusuntil stopped, then assertsexitstatus == "OK"(authorization can surface at task execution, not the POST — phase1-2 §1.3). Exponential backoff (1s→5s cap), context cancellation + timeout.*APIErrorparses the offending privilege from a 403;*TaskErrorparses it from a failed task exitstatus + log tail.internal/proxmox— fenced root-CLI backend (Privileged): limited to the three proven OS-root exceptions only —CreateGoldenLXC(keyctlpct create),MountUSBByUUID,SMART,Sensors; each cites why it can't be the API. Fence is structural (Client never shells out, Privileged never makes an HTTP call) and asserted in tests.- TLS trust: SHA-256 leaf-cert pinning (the host serves a self-signed cert) or
a CA file; an explicitly-named
insecure_skip_verifythat is off by default. No blanket verification disable. internal/config: JSON config file +FELHOM_AGENT_*env overrides; the token secret is never logged (Redacted()).internal/log: slog setup (text, stderr, configurable level).cmd/felhom-agent --selftest: read-only health report against a live host (version/nodes/status/guests/storage);--selftest=task --vmid NexercisesWaitTaskon a reversible snapshot→rollback→delete op (gated; default selftest mutates nothing).- Tests: unit tests with a mock HTTP transport + mock runner (UPID parse,
WaitTaskrunning→OK / failed-403 / timeout / ctx-cancel, 403→privilege error, response decoding against shapes captured live fromdemo-felhom, config redaction, and the API-vs-root routing fence).
Notes
- Types are grounded in the spike findings
(
felhom.eu/documentation/proxmox-platform.md,tests/phase{0,1-2,3}-findings.md) and the exact JSON shapes captured live fromdemo-felhom(PVE 9.2.2). - Verified:
go build/vet/testgreen on Go 1.24.4 (build server) and a live read-only--selftestagainst the demo host with TLS fingerprint pinning. - The 16-privilege
FelhomAgentrole + privsep token (role on both user and token) is provisioned out-of-band; the agent only consumes the token.